Barely Significant
← all excerpts

Performance of ChatGPT-4o and Four Open-Source Large Language Models in Generating Diagnoses Based on China's Rare Disease Catalog: Comparative Study.

J Med Internet Res · 2025 · PMC12192912 · PMID 40532199

1
hedged sentence
0.0200
closest p · 0.4× alpha
0.0200
boldest claim

The sentences

marginally significantP =.02actually significant
Larger models displayed architectural divergence where qwen2.5:72b maintained cross-lingual consistency (82.6% vs 83.5%; OR 0.88, P =1.000), while Llama3.1:70b showed language-dependent performance (80.2% vs 90.1%; OR 0.29, P =.02, marginally significant).

also in 26,082 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.