Image-based questions demonstrated lower accuracy (44.7%, 17/38) compared with text-based questions (54%, 27/50), though this difference did not reach statistical significance (Fisher's exact test, no conventional test statistic; p=0.52) (Table 3 ).
← all excerpts
Accuracy Is Not Enough: Reasoning and Reference Reliability in Orthopaedic Large Language Model (LLM) Applications.
1
0.5200
0.5200