OpenAI models follow a clear trend of improvement, progressing from GPT-4 to o1-preview, with o1-preview leading the APBench, albeit by a narrow margin.
← all excerpts
APBench and benchmarking large language model performance in fundamental astrodynamics problems for space engineering.
1
—
—