Barely Significant
← all excerpts

Benchmarking large language models' performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google Bard.

EBioMedicine · 2023 · PMC10470220 · PMID 37625267

1
hedged sentence
closest p
boldest claim

The sentences

may not be significantno p-value reported
While the observed improvements in the transition of responses from ‘poor’ to ‘good’ (with one such example in each LLM-Chatbot) may not be significant, they underline the present capacity of LLMs to acknowledge potential inaccuracies when prompted and make attempts at self-correction ( Table 5 , Table 6 , Table 7 ).

also in 1,077 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.