While overall differences in trustworthiness (mDISCERN) were observed, pairwise comparisons did not reach statistical significance, and inter-rater agreement for this metric was modest, indicating that conclusions related to trustworthiness should be interpreted cautiously.
← all excerpts
Domain-Specific vs. General-Purpose Large Language Models in Orthodontics: A Blinded Comparison of AlimGPT, GPT-4o, Gemini, and Llama.
1
—
—