helm-multi-metric-benchmark

IN premiseentries/2026/06/21/wiki-Generative_pre-trained_transformer-chunk-2.md

Created 2026-06-21T09:50:09+00:00

HELM (Holistic Evaluation of Language Models) from Stanford CRFM evaluates models across accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency

Summary

Stanford's HELM framework judges language models on a wide range of qualities beyond just correctness, including fairness, safety, consistency, and efficiency. This matters because it sets the standard that "good performance" is a multi-dimensional judgment, so any reasoning built on this premise treats model quality as more than a single accuracy score.

Dependents

These beliefs depend on this one: