alpaca-30b-50pct-closebook-consistency

IN premise — summaries/2026/08/24/xu-2024-knowledge-conflicts-survey-s4-intra-memory-conflict.md

Created 2026-08-25T02:59:02+00:00

Alpaca-30B is consistent in only approximately 50% of close-book QA cases (Li et al., 2023d).

Summary

When asked the same question without external tools, the Alpaca-30B model matches its own previous answer only about half the time, which is barely better than a coin flip. Any workflow that relies on this model producing stable, repeatable outputs will need to build in retries, cross-checks, or fallback mechanisms to compensate for that level of inconsistency.