code-llama-500b-code-tokens

IN premiseentries/2026/06/21/wiki-LLaMA-chunk-1.md

Created 2026-06-21T09:50:09+00:00

Code Llama is fine-tuned Llama 2 on 500B code tokens + 20B long-context tokens, with a separate Python-specialized variant trained on 100B Python-only tokens

Summary

Code Llama is a Llama 2 model that was further trained primarily on source code and long documents, with a separate version available that was trained almost exclusively on Python. This establishes the model's training lineage and areas of specialization, which is key context when assessing what it can and can't do well compared to the general-purpose base model.