constitutional-ai-principles-not-per-example-feedback
IN premise — entries/2026/06/21/wiki-LLaMA-chunk-4.md
Created 2026-06-21T09:50:09+00:00
Constitutional AI (Anthropic) uses a set of written principles for alignment rather than requiring individual human feedback for each training example.
Summary
Instead of needing humans to judge millions of individual training examples, Anthropic's approach relies on a shared set of written principles that the model uses to critique and revise its own responses. This matters because alignment becomes more scalable and auditable — you can read and debate the rules directly rather than depending on opaque human preference labels for every single interaction.
Dependents
These beliefs depend on this one:
- IN constitutional-ai-is-complete-alternative-alignment-path — Constitutional AI, developed by Anthropic, uses written principles rather than per-example human feedback and employs AI-generated feedback (RLAIF) based on those principles in place of human preference labels, representing a principle-driven approach to alignment that differs from standard RLHF in its feedback mechanism.