Agent and model evaluations in Gemini Enterprise Agent Platform reach general availability
Gemini Enterprise Agent Platform evaluations unify agent testing with live monitoring Compare models, track drift, and catch failures faster with 20+ metrics
Google announced that agent and model evaluations in Gemini Enterprise Agent Platform are now generally available. The update gives developers a single system to measure and compare agents and models during development and after launch, using consistent metrics across offline experiments and live production traffic.
The platform includes more than 20 prebuilt metrics covering quality, safety, grounding, tool use, trajectory, summarization, translation, and other tasks. Developers can also define custom codebased or LLMbased metrics, run experiments locally or serverside, and review failures through traces, logs, and issue clustering.
Google also highlighted online monitoring features that evaluate live traffic, show score trends, and send drift alerts. The evaluation tools are available through the Agent Platform SDK, agentscli, the console, and ADK. The service is generally available now, with standard pricing applying to modelbased evaluation calls and storage used for serverside runs.