Barely Significant
← all excerpts

Evaluating reasoning large language models on rumor generation, detection, and debunking tasks.

iScience · 2025 · PMC12554143 · PMID 41146718

1
hedged sentence
closest p
boldest claim

The sentences

a clear trendno p-value reported
Across all three datasets ( Tables 3 , 4 , and 5 ), a clear trend emerged: traditional machine learning models (SVM, CatBoost) and deep learning models (BERT) consistently outperformed both reasoning and non-reasoning LLMs in rumor detection.

also in 19,460 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.