For Gemini 3.0 Pro and clinical educator consensus score, mean scores were higher in the traditional style, but these differences did not reach statistical significance after Bonferroni correction ( P =.03 for both; corrected threshold P <.017).
← all excerpts
Agreement Between Reasoning-Oriented Generative AI Models and Clinical Educators in Evaluating Japanese Objective Structured Clinical Examination Transcripts: Preliminary Comparative Study.
1
0.0300
0.0300