c2023-eval-protocol-100-lg-queries

IN premise — summaries/2026/08/24/cohen-2023-ripple-effects-sA-acknowledgments.md

Created 2026-08-25T02:57:57+00:00

The RIPPLE EDITS evaluation protocol collects model responses to 100 random LG (long-geometry/logical generalization) queries and categorizes each into correct, abstain, or noise

Summary

This is the setup for how the system grades a model's ability to tackle complex logical or geometric puzzles: it poses 100 varied reasoning problems and sorts every answer into one of three buckets — got it right, said "I don't know," or produced gibberish. That three-way split matters because it lets the system tell apart genuine ignorance from careless hallucination, which drives very different follow-up actions.