Most importantly, accuracy remains consistently high, though it is better when Omission errors are excluded compared to when all outputs are considered The other models were less resilient and robust to prompt size and complexity and there was a clear trend that emerged: as the prompt size increased—while the total number of tasks remained fixed—the models’ performance declined.
← all excerpts
A strategy for cost-effective large language model use at health system-scale.
1
—
—