capability-risk-dual-scaling-proved-systemic

OUT derived (depth 3)

Created 2026-06-21T10:20:46+00:00 · Reviewed 2026-06-21T10:54:59+00:00

GPT-2's early demonstration of capability-coupled risks (1-7% memorization, staged release over misuse concerns) proved prescient rather than incidental — the same dual scaling pattern compounded into today's multi-layered safety challenge spanning capability classification (Level 3), government suspension directives (Fable 5/Mythos 5), architectural vulnerabilities (prompt injection), and calibration failures (over-refusal), confirming capability-risk co-scaling as a systemic property, not an early-stage artifact.

Justifications

SL — GPT-2's early warning proved predictive of systemic dual scaling, not an isolated incident

Antecedents (all must be IN):

  • IN gpt2-foreshadowed-capability-risk-dual-scaling — GPT-2's staged release over misuse concerns, combined with measurements showing 1-7% exact duplicate training data in its outputs, illustrated early tensions between scaling language models and managing associated risks such as memorization and potential misuse.
  • IN llm-safety-is-multi-layered-unsettled-challenge — LLM safety operates across multiple interdependent layers — capability risk classification (Opus 4 at Level 3), architectural vulnerabilities (prompt injection as inherent design flaw), regulatory intervention (Fable 5/Mythos 5 suspension), and behavioral calibration trade-offs (Opus 4.7 over-refusal complaints) — with no single layer providing comprehensive coverage and each layer creating tensions with the others.