failed to reach significanceP = 0.0295
Though GPT descriptively gave worse responses following the three different “jailbreaking” prompts, compared to two contrasting prompts requesting ethical responses, this similarly failed to reach significance ( P = 0.0295, d = 0.588) after correction.