How do we measure quality of generated content?
Traditional metrics (BLEU, ROUGE) correlate poorly with human judgment for LLM outputs. LLM-as-judge — using GPT-4 or Claude to score outputs against defined criteria — correlates well with human evaluation and scales to thousands of outputs per day. RAGAS provides specialised metrics for RAG systems. Always validate your evaluation methodology against human judgments on a calibration set before using it for production monitoring.