highly significantP =.006
The Kruskal-Wallis test revealed highly significant group differences ( H =11.24, df=2; P =.006).
The Kruskal-Wallis test revealed highly significant group differences ( H =11.24, df=2; P =.006).
IM-1 (80%) was not significantly different from Med-PaLM 2 (80% vs 92%, difference=12 percentage points, P =.16), LLaVA-Med (80% vs 76%; P =.50), BioGPT (80% vs 68%; P =.22), or ChatGPT 5.0 (80% vs 64%; P =.048—marginally significant).
Performance differences from LLaVA-Med ( P =.12), IM-2 ( P =.12), and BioGPT ( P =.38) did not reach statistical significance but consistently favored the comparators.