Measuring LLM quality with no single answer: metric limits, LLM-as-a-judge (modes, biases, mitigations), human agreement, Elo arenas, HELM.