embedding-inversion-email-tweet-success-rates

IN premise — summaries/2026/08/24/jha-2025-vec2vec-s7-ablations.md

Created 2026-08-24T17:10:58+00:00

Translated embeddings enable extraction of ~80% of identifiable information from emails and ~67% from tweets for certain model pairs

Summary

When you convert text representations from one AI model's format into another's, you can still recover roughly four-fifths of the meaningful content from emails and a bit over two-thirds from tweets. This matters because it tells the system that cross-model text reconstruction is largely reliable for structured messages, with a modest penalty for very short, informal formats, so downstream reasoning can trust these translations without expecting catastrophic information loss.