close to significancep = .069
However, post hoc test (Bonferroni) didn’t reveal any significant differences, with only one being seemingly close to significance - GPT-4 performing better than GPT-3.5 ( p = .069), meaning potentially that the sample size is not big enough to catch small effect size of this difference.