Google describes a quality flywheel for coding agents

Google AI agent evaluation helps teams improve multi-turn coding agents Use AutoRaters, custom metrics, and human review to validate every change

Google says it has created a developerfacing skill that helps coding agents evaluate and improve multiturn AI agents through a structured quality loop. The approach combines data preparation, inference, grading, failure analysis, and iteration. It uses Google’s adaptive AutoRaters alongside custom metrics so teams can measure whether a change actually improves behavior instead of just seeming better on a few examples. The company says the workflow is designed to keep evaluation and optimization separate, with human approval still required. It can use synthetic scenarios for early testing and production traces for ongoing monitoring, and it is available through two packages for different development stacks.