ginc-vocabulary-sizes-tested

IN premise — summaries/2026/08/24/xie-2021-icl-bayesian-sR-references-chunk-2.md

Created 2026-08-25T02:58:57+00:00

GINC experiments test vocabulary sizes of 50, 100, and 150, with larger vocabularies improving ICL accuracy because each hidden state is more likely to emit a distinct symbol.

Summary

The GINC experiments compared three vocabulary sizes (50, 100, and 150) and found that bigger vocabularies lead to better in-context learning accuracy. The practical implication is that symbol distinctness matters: with more symbols available, each internal state is less likely to accidentally collide with another, so the system can differentiate its outputs more reliably.