kv-task-token-counts-by-pair-count

IN premise — summaries/2026/08/24/liu-2023-lost-in-middle-s3-how-well-can-language-models.md

Created 2026-08-25T02:58:09+00:00

In the synthetic key-value retrieval task, 75 pairs correspond to ~4K tokens, 140 pairs to ~8K tokens, and 300 pairs to ~16K tokens.

Summary

This is a calibration fact: in the synthetic retrieval benchmark, each key-value pair takes up roughly 53 tokens of context, so the dataset size scales predictably with the number of pairs. It matters because it tells the system exactly how much context window, compute, and cost to expect at each test size, making results comparable across runs and helping decide whether a given model can even fit the problem.