prh-infonce-optimal-equals-kpmi-plus-input-dependent-offset
IN premise — summaries/2026/08/24/huh-2024-prh-sR-references-chunk-3.md
Created 2026-08-24T17:10:57+00:00
The Bayes optimal solution of InfoNCE with temperature τ recovers K_PMI plus an input-dependent offset c_X(x_a) (for τ=1) or an additional scale factor (for general τ), and negative pairs are sampled i.i.d. from the marginals P̃(x) = ∫ P_coor(x, x⁺) dx⁺ rather than from the joint distribution.
Summary
InfoNCE, the standard contrastive loss used in representation learning, is not just a heuristic: at its optimal setting it recovers a well-defined mutual-information quantity (K_PMI) plus a small correction that depends on the particular input, and the way negative samples are drawn (from the marginal distribution rather than the joint) is baked directly into that formula. Knowing the exact correction and the sampling rule tells us precisely how far a contrastive-trained representation sits from the true information-theoretic optimum, and whether that gap is uniform across inputs or varies with them.