Barely Significant
← all excerpts

Comparing large Language models and human annotators in latent content analysis of sentiment, political leaning, emotional intensity and sarcasm.

Sci Rep · 2025 · PMC11968858 · PMID 40181141

1
hedged sentence
0.0690
closest p · 1.4× alpha
0.0690
boldest claim

The sentences

close to significancep = .069so close (0.05 < p ≤ 0.1)
However, post hoc test (Bonferroni) didn’t reveal any significant differences, with only one being seemingly close to significance - GPT-4 performing better than GPT-3.5 ( p = .069), meaning potentially that the sample size is not big enough to catch small effect size of this difference.

also in 692 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.