bilingual-pair-filtering-requires-single-token-and-top1

IN premise — summaries/2026/08/24/park-2023-linear-representation-s2-the-lion-is-the.md

Created 2026-08-24T17:11:02+00:00

Park et al. (2023) filter bilingual translation pairs by keeping only top-1 mutual correspondences from a public word-translation dictionary and excluding pairs that are multi-token in the LLaMA-2 vocabulary (32K tokens), yielding 205–231 valid contexts per translation concept pair across four tested language pairs.

Summary

This is a factual note about how a specific study narrowed its translation dataset: by demanding that each word pair be a single token in LLaMA-2 and that the dictionary link be mutual and unique, the researchers were left with only about 200 to 230 usable examples per concept pair across four language combinations. That small, tightly filtered sample is the ground everything else in the paper builds on, so any claim about cross-lingual transfer or alignment in that work is bounded by these strict inclusion rules.