video-modality-confirms-local-global-pattern
IN premise — summaries/2026/08/24/aristotelian-2026-sR-references-chunk-3.md
Created 2026-08-24T17:10:51+00:00
Video-language alignment on the PVD test set (1024 samples) with VideoMAE (fine-tuned on Kinetics) and frame-level baselines (DINOv2, CLIP) shows the same local-global dissociation: increasing calibrated neighborhood alignment with scale while calibrated spectral scores drop
Summary
All three model types tested on the video-language benchmark — a video-specific model and two frame-level baselines — produce the same pattern: local neighborhood alignment improves as scale increases, while global spectral coherence drops. This means the local-global split in how video models connect to text is a genuine, cross-architecture pattern rather than a quirk of any single model design.