did not reach statistical significancep = 0.1045
The five-point margin of OpenAI-o1 over ChatGPT-4o did not reach statistical significance ( p = 0.1045) but could represent one additional correct decision in twenty clinically relevant scenarios.
The five-point margin of OpenAI-o1 over ChatGPT-4o did not reach statistical significance ( p = 0.1045) but could represent one additional correct decision in twenty clinically relevant scenarios.
Our primary aim was to determine whether any observed performance gain is clinically meaningful, defined as a statistically and practically significant margin over resident performance that could justify integrating the model as a decision support aid during patient work-up and surgical counseling.