c2023-gpt3-ice-abstain-49-vs-llama-28

IN premise — summaries/2026/08/24/cohen-2023-ripple-effects-s6-conclusion-and-discussion.md

Created 2026-08-25T02:57:57+00:00

GPT-3 with ICE abstains on 49% of LG (logical generalization) queries, compared to 28% for LLAMA with ICE

Summary

GPT-3 declines to answer logical generalization questions at nearly twice the rate of LLAMA when the ICE abstention mechanism is active, meaning it flags itself as unsure far more often. This gap matters in practice because any system built on these models will get a noticeably lower usable-answer rate from GPT-3 on reasoning-heavy tasks, so choosing one over the other has a direct cost in coverage.