xie-2021-mixture-of-hmms-setting

IN premise — summaries/2026/08/24/xie-2021-icl-bayesian-s1-introduction.md

Created 2026-08-25T02:58:55+00:00

The pretraining distribution in Xie et al. (2021) is modeled as a mixture of HMMs: p(o₁,...,o_T) = ∫ p(o₁,...,o_T|θ) p(θ) dθ, where θ parameterizes the HMM transition matrix.

Summary

The model assumes the training data was generated by a blend of many different Markov transition patterns rather than one single fixed pattern, with a prior distribution deciding how much weight each pattern carries. This sets the statistical foundation for everything downstream: if the true data-generating process doesn't look like such a mixture, the model's predictions and inferences will be systematically biased.