borderline significantP = .051
When accuracy differences across the 4 subtypes (radiograph, clinical photo, schematic illustration, and mixed) were assessed using chi-square tests, subtype-level variability within each model was significant (Claude-4 Opus, P = .040; Gemini 2.5 Pro, P = .005) or borderline significant (ChatGPT-4o, P = .051), confirming that image subtype substantially influenced model performance.