vec2vec-inversion-leakage-80-percent-emails-67-percent-tweets
IN premise — summaries/2026/08/24/jha-2025-vec2vec-s4-experimental-setup.md
Created 2026-08-24T17:10:57+00:00
Zero-shot inversion of vec2vec-translated embeddings recovers document content (names, dates, financial data) in up to 80% of Enron emails and 67% of tweets, with LLM-judge (GPT-4o) information-leakage accuracy of 30–67% across model pairs on a 50-tweet subset.
Summary
Converting sensitive documents into a shared, universal vector format does not actually protect their contents. An attacker can reverse the translation and recover names, dates, and financial figures from a large majority of emails and tweets, meaning the embedding layer is not a meaningful privacy barrier.