vec2vec-training-compute-176-gpu-days
IN premise — summaries/2026/08/24/jha-2025-vec2vec-sR-references.md
Created 2026-08-24T17:10:58+00:00
Total training compute for the vec2vec experiments was ~176 GPU days across 25 fully-trained models (1–7 days each), 30 partially-trained models (stopped at 2 days), and one Qwen+GTE pair at 20 days on A100
Summary
The vec2vec experiments cost roughly 176 GPU days in total, spread across 25 models trained to completion, 30 that were cut short after 2 days, and one unusually expensive 20-day run on A100 hardware. This is the baseline compute budget for all those results, meaning any conclusions drawn from the experiments were produced within that spending envelope, and the one Qwen+GWE pair alone consumed more compute than the 30 short runs combined.