Barely Significant
← all excerpts

Performance of o1 pro and GPT-4 in Self-Assessment Questions for Nephrology Board Renewal.

Front Med (Lausanne) · 2025 · PMC12685630 · PMID 41377810

1
hedged sentence
closest p
boldest claim

The sentences

In the remaining domains (AKI, hypertension/vascular diseases, and ADPKD/urology), o1 pro’s accuracy exceeded that of GPT-4, although these differences did not reach statistical significance.

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.