chameleon-models-evaluated

IN premise — summaries/2026/08/24/xie-2024-chameleon-sloth-s0-abstract-chunk-1.md

Created 2026-08-25T02:58:58+00:00

The LLM knowledge conflict study evaluates 3 closed-source models (ChatGPT, GPT-4, PaLM2) and 5 open-source models (Qwen-7B, Llama2-7B/70B, Vicuna-7B/33B) in zero-shot settings.

Summary

This study tested eight different language models — a mix of large proprietary ones and smaller open-weight ones — to see how they handle contradictory information, without giving them any extra training or examples. It matters because the findings the system relies on are scoped to these specific models under those specific conditions, so results may not generalize to other architectures or to fine-tuned settings.