The results from the Wilcoxon signed-rank test comparing the translation metrics for the two prompts combining the responses from all four LLMs revealed highly significant differences in all evaluation metrics, with prompt 2 faring significantly better in terms of BLEU, METEOR, and CHRF scores, whereas prompt 1 had significantly better TER scores ( p < 0.001 for all comparisons).
← all excerpts
Comparative Evaluation of Large Language Models for Translating Radiology Reports into Hindi.
1
—
—