hendel-llama-13b-architecture

IN premise — summaries/2026/08/24/hendel-2023-icl-task-vectors-sR-references.md

Created 2026-08-25T02:58:05+00:00

LLaMA 13B has 5120 hidden dimensions, 40 layers, and 40 attention heads.

Summary

This records the internal blueprint of the LLaMA 13B model: how wide its internal representations are, how many stacked processing stages it passes through, and how it splits its attention into parallel channels. It matters because every claim about the model's memory footprint, inference speed, or how it compares to other architectures rests on these specific structural numbers being correct.