McNemar’s test confirms that this improvement is not due to random fluctuation: when comparing ORCH with the strongest single-model c = 1 (i.e., 151 questions answered correctly only by ORCH and 1 only by DeepSeek), yielding χ 2 ≈ 148.03 and p ≈ 8.99 × 10 − 34 <0.05, indicating a highly significant difference at the 95% confidence level. 4.4 Comparative experiment 2—ORCH-1 ablation: removing one agent (XAI) To verify that the performance gains of ORCH indeed stem from the cooperation of multiple agents, we conduct an ablation study in which one agent (XAI) is removed, yielding a two-agent variant denoted ORCH-1 ( Levy et al., 2024 ).
← all excerpts
ORCH: many analyses, one merge-a deterministic multi-agent orchestrator for discrete-choice reasoning with EMA-guided routing.
1
—
—