LLM Judges Are Lying: 67% of Evaluations Are Inconsistent
New research reveals that LLM-as-judge frameworks suffer from per-instance inconsistency masked by aggregate metrics. The paper proposes conformal prediction sets as a diagnostic tool, but the findings suggest that current evaluation pipelines are unreliable.











