prh-dataset-pools-wit-1m-laion-15m

IN premise — summaries/2026/08/24/koepke-2026-back-into-cave-s3-experimental-setup.md

Created 2026-08-24T17:11:01+00:00

The PRH evaluation datasets are constructed as WIT-1M (2,389,146 final pool → 1M sampled) and LAION-15M (17,298,107 final pool → 15M sampled), with the query set fixed at WIT-1024 and gallery scaled to 1K, 10K, 50K, 100K, 500K, 1M (WIT), and 10M, 15M (LAION-400M).

Summary

The evaluation is run against two fixed pools of images — roughly 2.4 million and 17.3 million — from which 1 million and 15 million are sampled, always queried by the same 1,024-item test set, with the searchable gallery grown in steps from 1,000 up to 15 million. This matters because it locks down exactly what data, in what size, is being measured, so any performance numbers reported under these conditions are directly comparable across runs and experiments.