Across all three datasets ( Tables 3 , 4 , and 5 ), a clear trend emerged: traditional machine learning models (SVM, CatBoost) and deep learning models (BERT) consistently outperformed both reasoning and non-reasoning LLMs in rumor detection.
← all excerpts
Evaluating reasoning large language models on rumor generation, detection, and debunking tasks.
1
—
—