fv-extraction-head-count-gptj

IN premise — summaries/2026/08/24/todd-2023-function-vectors-s2-function-vector-causal-effects-cannot-be-recovered-from-the-.md

Created 2026-08-25T02:58:41+00:00

For GPT-J, the FV is constructed from |A| = 10 attention heads (where performance plateaus), using 100 clean 10-shot prompts for computing mean activations and 25 corrupted 10-shot prompts per task for computing AIE.

Summary

For the GPT-J model, the system's fingerprint is built from exactly 10 attention heads, which is the point where adding more heads no longer improves identification accuracy. The fingerprint is calculated using 100 clean prompts to establish a baseline of normal behavior and 25 deliberately corrupted prompts per task to measure how the model responds to adversarial input, giving a concrete, reproducible recipe for detecting whether GPT-J is present.