highly significantP <.001
The profound and consistent intermodel differences, confirmed by highly significant Friedman tests ( Table 3 ; all P <.001) and detailed in the pairwise comparison heatmap ( Figure 3 ), underscore that these performance characteristics are inherent to model architecture and training, not artifacts of query phrasing.