Specifically, “Errors or inaccurate information” were low across configurations (mean>3.88), with a favorable trend toward the simulated-agents GPT, particularly in the “vignette” (mean 4.75, SD 0.46) and the “checklist” (mean 4.88, SD 0.35) components. “Missing information” was rare in the simulated-agents GPT (mean≥4.19 across all sections), but more common in the personalized GPT, especially in the “script” (
← all excerpts
AI-Driven Objective Structured Clinical Examination Generation in Digital Health Education: Comparative Analysis of Three GPT-4o Configurations.
1
—
—