The model × prompt interaction showed a trend toward significance but did not reach the conventional threshold.
← all excerpts
How Reliably Do Large Language Models Reproduce Vital Pulp Therapy Guidelines? A Mixed-Effects Evaluation of Guideline-Concordance and Error Directionality
2
—
—
The sentences
The overall model × prompt interaction did not reach statistical significance, although individual interaction terms suggested model-dependent prompt responsiveness. 3.3.