While the Qwen3-32B multiagent framework achieved a higher mean factual consistency score of 0.734 (SD 0.06) compared to the single-LLM Qwen3-32B approach 0.699 (SD 0.07), this numerical improvement did not reach statistical significance (adjusted P =.33).
← all excerpts
A Large Language Model-Powered Multiagent Framework Emulating Standardized Patients in Clinical Communication Skills Training: Development and Evaluation Study.
1
0.3300
0.3300