Comparisons with earlier versions of OpenAI models in other medical domains demonstrate a clear trend of improving accuracy in addressing medical knowledge tasks.
← all excerpts
Performance of the ChatGPT-5 Language Model in Solving a Specialty Examination in Balneology and Physical Medicine.
3
—
—
The sentences
While the Mann-Whitney U test (p = 0.07) did not confirm statistically significant differences in confidence between the two categories of questions, it suggested a possible trend.
The Mann-Whitney U test (p = 0.07) indicated that the difference in confidence levels between clinical and theoretical questions did not reach statistical significance (α = 0.05), although a trend toward potential differences was observed.