mistral-llama-weekday-month-accuracy
IN premise — summaries/2026/08/24/engels-2024-not-all-features-linear-s4-s-parse-autoencoders-find-multi-d-imensional-features.md
Created 2026-08-25T02:58:01+00:00
Mistral 7B and Llama 3 8B achieve approximately 29–31/49 on natural-language Weekdays tasks and 125–143/144 on Months tasks, yet trivial accuracy on raw modular-addition prompts.
Summary
These small open-source models can mostly get right which weekday or month follows a given one when the question is phrased like a normal English sentence, but they collapse to near-random when the same calculation is written as a raw modular-arithmetic expression. The takeaway is that they have pattern-matched the surface language rather than learned the underlying arithmetic, so their apparent calendar reasoning won't transfer to any setting that changes the phrasing.