Gemini-1.5-Pro-002 demonstrated optimal performance with the original and reflection prompts, but this did not reach statistical significance (P=0.635).
← all excerpts
Diagnostic performance of multimodal large language models in radiological quiz cases: the effects of prompt engineering and input conditions.
1
0.6350
0.6350