Claude showed a trend toward lower scores for patient prompts (1.78 vs 1.56), although this did not reach statistical significance (U = 1156.00, P = .258).
← all excerpts
Large Language Model Hallucinations in Spine Surgery: A Comparative Analysis of Clinician vs Patient-Level Prompts.
1
—
—