Barely Significant
← all excerpts

Evaluating the Intelligence of large language models: A comparative study using verbal and visual IQ tests.

Comput Hum Behav Artif Hum · 2025 · PMC13189213 · PMID 42170347

2
hedged sentences
closest p
boldest claim

The sentences

an obvious trendno p-value reported
Overall, there is an obvious trend of larger models having higher IQ scores.

also in 1,006 other papers

a clear trendno p-value reported
The results show a clear trend of increasing IQ scores as model size scales up, even within the same model architecture.

also in 19,460 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.