marginally significantp < 0.05
TF-IDF demonstrated the weakest correlation, with a marginally significant correlation only with expert-generated questions (ρ = 0.38, p < 0.05), while correlations with ACG-MCQ performance (ρ = 0.30, p = 0.13) and real-world questions ( ρ = 0.28, p = 0.16) were not statistically significant.