constitutional-ai-principles-not-per-example-feedback

IN premiseentries/2026/06/21/wiki-LLaMA-chunk-4.md

Created 2026-06-21T09:50:09+00:00

Constitutional AI (Anthropic) uses a set of written principles for alignment rather than requiring individual human feedback for each training example.

Summary

Instead of needing humans to judge millions of individual training examples, Anthropic's approach relies on a shared set of written principles that the model uses to critique and revise its own responses. This matters because alignment becomes more scalable and auditable — you can read and debate the rules directly rather than depending on opaque human preference labels for every single interaction.

Dependents

These beliefs depend on this one: