In the ChatGPT (OpenAI; GPT-5.2) group, few-shot prompting modestly improved accuracy (from 0.700 to 0.850), macro F1 score (from 0.693 to 0.848), and kappa (from 0.400 to 0.700), although the difference did not reach statistical significance ( p = 0.108).
← all excerpts
Evaluation of large language models for VI-RADS reports: a comparative analysis of zero-shot and few-shot prompting.
1
0.1080
0.1080