Our post hoc evaluation using the McNemar test revealed that performance differences between the LLMs (GPT-4, LLaMa 3, and o3-mini) were only partially significant.
← all excerpts
Automated Safety Plan Scoring in Outpatient Mental Health Settings Using Large Language Models: Exploratory Study.
1
—
—