melllo-gpt35-vs-gptj-scores
IN premise — summaries/2026/08/24/zhong-2023-mquake-s6-related-work.md
Created 2026-08-25T02:59:06+00:00
MeLLo achieves 91.1/85.5 (CF/T at 1 edit) with GPT-3.5 versus 38.9/30.7 with GPT-J, indicating performance scales with base model instruction-following ability.
Summary
MeLLo's output quality is heavily dependent on which base model it sits on top of: swapping GPT-J for GPT-3.5 roughly doubles its performance on editing and faithfulness metrics. In practice, this means the system's results are largely inherited from the underlying model's ability to follow instructions, so improvements to MeLLo's own logic or architecture matter less than choosing a stronger base model.