Paired t -tests on the individual classification outcomes revealed that the differences in recall (∆ = 0.008, p = 0.53), precision (∆ = 0.0252, p = 0.047), and accuracy (∆ = 0.008, p = 0.53) between GPT-4 and the supervised NLP did not reach statistical significance ( p < 0.05 across all tests) suggesting that the overall performance of GPT-4 and the supervised NLP model is statistically comparable (when GPT-4 is provided a curated master list).
← all excerpts
LLM enabled classification of patient self-reported symptoms and needs in health systems across the USA.
1
0.0500
0.0500