evaluation-geometry-is-editing-coordinate-system
OUT derived (depth 7)
Created 2026-08-25T03:13:05+00:00 · Reviewed 2026-08-25T04:02:18+00:00
The optimal evaluation metric for an embedding model (cosine/Spearman pipeline) is simultaneously the optimal coordinate system for specifying knowledge edits, because both are readouts of the same universal covariance geometry
Justifications
SL — evaluation-geometry-predicts-editability establishes that eval and edit share the same second-moment geometry; geometry-as-universal-semantic-currency elevates that geometry to a model-independent semantic object. Together: the metric you optimize in MTEB/SBERT IS the coordinate frame in which ROME-style edits operate, unifying evaluation and intervention into a single geometric principle.
Antecedents (all must be IN):
- OUT evaluation-geometry-predicts-editability — The convergence of evaluation geometry (cosine/Spearman in SBERT/MTEB) and editing geometry (covariance whitening in ROME) on the same second-moment structure means that improving evaluation alignment and enabling reliable editing are two operational views of the same geometric optimization over the residual-stream covariance.
- OUT geometry-as-universal-semantic-currency — The covariance geometry is the single operational definition of "meaning" in LLMs, simultaneously determining what can be measured (evaluation via cosine/Spearman), what can be modified (rank-one editing via C⁻¹k*), what converges across architectures (SAE/Park universality), and what is hierarchically structured (feature neighborhoods as theorem instantiations)—making it a model-independent semantic currency rather than an architecture-specific artifact.
Dependents
These beliefs depend on this one:
- OUT geometric-closed-loop-eval-edit-navigate — The MTEB evaluation coordinate system, ROME's rank-one editing, and SAE feature navigation form a single closed geometric loop: evaluation identifies the whitened directions to read, ROME modifies one whitened direction to write, and SAE neighborhood traversal navigates between whitened directions—each operation is a different linear functional on the same covariance-geometric space, unified by the Riesz map.
- OUT mteb-simultaneously-measures-read-and-edit-addressability — MTEB's "no dominant model" observation is simultaneously a statement about readout diversity AND edit-addressability diversity: different embedding models optimize different directions in the same shared editable geometric space, making the benchmark a dual measure of read-quality and edit-coverage.