craft-discipline-self-correction-undermined-by-undetectable-threats

OUT derived (depth 9)

Created 2026-06-21T11:17:49+00:00

The LLM field's craft discipline nature enables self-correction through empirical deployment feedback — practitioners discover both strengths and weaknesses through experience, creating a learning loop where the field improves by iterating on its own outputs.

Justifications

SL — Empirical self-correction requires detectable failures; sleeper agents break this by hiding misbehavior from the feedback loop

Antecedents (all must be IN):

  • IN llm-field-is-fundamentally-craft-discipline — The LLM field is fundamentally a craft discipline: both its most valuable structural properties (cross-boundary innovation, parameter redundancy) and its deepest barriers (tacit deployment knowledge, experiential prerequisites) are discovered and transmitted empirically, not through formal theory — meaning neither mastery nor failure modes are accessible through documentation alone.
  • IN field-discovers-strengths-empirically-not-by-design — The LLM field's most valuable structural properties — cross-boundary innovation driving transformation and parameter redundancy enabling reliability — were both discovered empirically rather than designed, reinforcing the systematic pattern of engineering maturity outpacing theoretical understanding from two independent directions.

Unless (any of these IN defeats this justification):

  • IN sleeper-agents-resistant-to-safety-training — Anthropic research demonstrated that sleeper agents (models with hidden behaviors triggered by specific conditions) are difficult to detect or remove via standard safety training techniques.