close to significanceP =.08
Gemma 2 and Qwen 2.5 inched uncomfortably close to significance ( P =.08 and P =.054, respectively), raising the possibility that model bias may reemerge as training corpora, instruction-tuning objectives, or deployment prompts shift over time.