did not reach statistical significancep = 0.05664
One‐sample one‐sided Wilcoxon signed‐rank testing against the clinical routine standard of 4.0 showed that all native image categories and nearly all AI‐generated image categories were rated significantly above the clinical routine threshold (Table S1 ), with the exception of AI‐FLAIR posterior fossa ratings, which were higher than the clinical reference value of 4.0 but did not reach statistical significance ( p = 0.05664).