Q: What are AI evals, and what is LLM-as-judge?
Evals are how you test an AI system. They answer the question every stakeholder eventually asks: "how do we know it's working?"
Normal software has unit tests: same input, same expected output, pass or fail. LLM outputs don't work that way. Generative AI is non-deterministic: the same question can produce a different answer every run, quality is a spectrum, and a hundred different answers can all be valid. So instead of testing single cases, you build an eval: a set of test inputs, a definition of what good looks like, and a way to score the system's outputs against it. Run the eval after every change, and you know whether you made things better or worse. Without evals, you're changing prompts and hoping.
How do you score outputs at scale? Sometimes with simple checks (did it return valid JSON, did it include the required fields). But for judging quality, the practical answer is LLM-as-judge: you use another LLM (very Inception, I know) to grade the outputs against a rubric you write. It sounds circular, but it works, because judging an answer against clear criteria is an easier task than producing it. You spot-check the judge against your own ratings until you trust it.
If you have a QA, SDET, or testing background, take note: evals are the most natural entry point into AI engineering, and teams are desperate for people who take them seriously.