highly significantP <.001
With the prompt, ChatGPT-4.0 Turbo and Claude 2 exhibited highly significant differences ( P <.001) compared with most other models, indicating substantial performance enhancement when the prompt was used.
With the prompt, ChatGPT-4.0 Turbo and Claude 2 exhibited highly significant differences ( P <.001) compared with most other models, indicating substantial performance enhancement when the prompt was used.