2 , an interesting trend emerged: VideoLLaMA’s mean scores consistently fell between those of Researcher A and Researcher B, QwenVL tended to assign higher ratings than both experts, while InternVL produced consistently lower scores.
← all excerpts
Benchmark evaluation of video large language models in quality assessment of science popularization videos for dry eye.
3
—
—
The sentences
InternVL also showed statistical significance with Researcher B, while QwenVL did not reach statistical significance with either rater.
For VIQI II (information accuracy), both VideoLLaMA and InternVL presented significant agreement with both experts ( p < 0.05), whereas QwenVL again failed to achieve significance.