prh-early-cnn-layers-converge-first
IN premise — summaries/2026/08/24/huh-2024-prh-s2-representations-are-converging-chunk-1.md
Created 2026-08-24T17:10:56+00:00
Early CNN layers (Gabor-like filters) converge across architectures and datasets before later layers do; later layers remain more task-specific, as shown by model stitching results from Lenc & Vedaldi (2015).
Summary
The basic visual features that deep networks learn first (edges, textures, simple orientations) end up looking almost the same regardless of what architecture or dataset you train on, while the deeper, more abstract features are where each network truly specializes. This matters because it suggests that the "groundwork" of perception is somewhat fixed and transferable, so the real design choices and task adaptation happen in the upper layers.