Notably, for GPT-4, its best prompt strategy (scene-definition) outperformed its weakest strategy (basic) by approximately 1.67 percentage points in average F1, but the difference did not reach statistical significance ( p = 0.094 ).
← all excerpts
Supervised Learning and Large Language Model Benchmarks on Mental Health Datasets: Cognitive Distortions and Suicidal Risks in Chinese Social Media.
1
0.0940
0.0940