defense-in-depth-required-and-bounded-by-theory-gap

IN derived (depth 5)

Created 2026-06-21T10:20:46+00:00 · Reviewed 2026-06-21T14:41:08+00:00

The five-dimensional defense requirement for LLM reliability is both necessitated and bounded by insufficient formal understanding — no single defense layer has theoretical guarantees (hence defense-in-depth is required), but the same theory gap means defense-in-depth itself lacks formal assurance of adequacy, creating an irreducible reliance on empirical validation.

Justifications

SL — The theory gap simultaneously creates the need for layered defense and limits confidence in that defense

Antecedents (all must be IN):

  • IN llm-reliability-requires-five-independent-defense-dimensions — LLM reliability requires independent defenses across at least five dimensions — two control layers (training-time alignment, inference-time prompting) and three security surfaces (training data poisoning, prompt injection, architectural vulnerability) — with no single defense sufficient on its own.
  • IN formal-understanding-insufficient-across-llm-stack — Formal theoretical understanding consistently proves insufficient across the entire LLM stack: RLHF's complete mathematical specification fails without dozens of engineering details, prompting's irreducible sensitivity resists formal analysis, the capacity bottleneck inverts between pretraining and alignment stages — and the inversion means that even a correct scaling theory for one stage actively misleads for the next.

Dependents

These beliefs depend on this one: