Barely Significant
← all excerpts

Evaluation of the Performance of Generative AI Large Language Models ChatGPT, Google Bard, and Microsoft Bing Chat in Supporting Evidence-Based Dentistry: Comparative Mixed Methods Study.

J Med Internet Res · 2023 · PMC10784979 · PMID 38009003

1
hedged sentence
0.0490
closest p · 1.0× alpha
0.0490
boldest claim

The sentences

marginally statistically significantP =.049actually significant
Corroborating evidence was provided by Wilcoxon test, which did not detect any statistically significant difference overall between the scores given by the 2 evaluators for the answers provided by the 4 LLMs ( Table 2 ), except for the scores given for the answers provided by ChatGPT-4, between which a marginally statistically significant difference was found ( P =.049).

also in 227 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.