Barely Significant
← all excerpts

Assessing the performance of large language models (LLMs) in answering medical questions regarding breast cancer in the Chinese context.

Digit Health · 2024 · PMC11462564 · PMID 39386109

1
hedged sentence
closest p
boldest claim

The sentences

The differences for chosen accuracy metrics among the three LLMs did not reach statistical significance, but only ChatGPT demonstrated a sense of human compassion.

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.