Qwen3-Next-80B-A3B-Instruc t achieved the highest observed accuracy for exact/lenient biome agreement (86.37%), very slightly above GPT-5-mini (85.57%), GPT-4.1 (84.97%), GPT-3.5-turbo-1106 (84.77%), and Microsoft-Phi-4 (83.37%); however, these differences did not reach statistical significance after post hoc correction (letters “a” in Fig. 8A ).
← all excerpts
Enhanced semantic classification of microbiome sample origins using large language models (LLMs).
1
—
—