This is only an overall trend, and ORs are not strictly increasing over time nor are individual differences between LLMs necessarily statistically significant, as is the case for LLMs released around a similar time: the difference in ORs in experiment 1 a between GPT-4 (2.51) and Mistral (2.60) is not statistically significant.
← all excerpts
Large language models can consistently generate high-quality content for election disinformation operations.
1
—
—