Barely Significant
← all excerpts

Accuracy and Reliability of Chatbot Responses to Physician Questions.

JAMA Netw Open · 2023 · PMC10546234 · PMID 37782499

2
hedged sentences
0.0460
closest p · 0.9× alpha
0.0460
boldest claim

The sentences

a significant trendP = .046actually significant
Among both descriptive and binary questions, the median accuracy scores for easy, medium, and hard answers were 6.0 (IQR, 6.0-6.0; mean [SD] score, 5.9 [0.3]), 5.5 (IQR, 3.9-6.0; mean [SD] score, 4.8 [1.7]), and 5.8 (IQR, 5.0-6.0; mean [SD] score, 5.3 [1.1]), respectively, with a significant trend ( P = .046 determined by the Kruskal-Wallis test).

also in 6,248 other papers

Subjectively more difficult questions seemed to have slightly less accurate scores (mean score, 4.2) than easier questions (mean score, 4.6), suggesting a potential limitation in handling complex medical queries, but this did not reach statistical significance.

also in 75,868 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.