{"id":17357,"date":"2026-08-03T00:37:19","date_gmt":"2026-08-03T00:37:19","guid":{"rendered":"https:\/\/techtrendfeed.com\/?p=17357"},"modified":"2026-08-03T00:37:19","modified_gmt":"2026-08-03T00:37:19","slug":"agent-and-mannequin-evaluations-in-gemini-enterprise-agent-platform-are-actually-ga","status":"publish","type":"post","link":"https:\/\/techtrendfeed.com\/?p=17357","title":{"rendered":"Agent and Mannequin Evaluations in Gemini Enterprise Agent Platform are actually GA"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p><img decoding=\"async\" class=\"banner-image\" src=\"https:\/\/storage.googleapis.com\/gweb-developer-goog-blog-assets\/images\/unnamed_8.original.png\" alt=\"unnamed (8)\"\/>  <\/p>\n<div class=\"inner-block-content rich-content\">\n<p data-block-key=\"pxh3z\">Agent high quality should be measured throughout growth towards the circumstances you wrote, and after launch towards the duties the agent truly carried out. <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/agent-evaluation\"><b>Agent and mannequin evaluations<\/b><\/a> in Agent Platform are actually typically obtainable, letting you measure and examine brokers and fashions at each levels, on one engine with constant metrics. If you use constant high quality scoring on native experiments and reside site visitors, a drift in manufacturing factors to an issue with the agent moderately than with the best way it was measured.<\/p>\n<h3 data-block-key=\"bhzve\" id=\"what's-generally-available\"><b>What&#8217;s typically obtainable<\/b><\/h3>\n<ul>\n<li data-block-key=\"sv45\"><b>Metrics.<\/b> Begin from greater than 20 <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/manage-metrics\">pre-built metrics<\/a> spanning high quality, security, grounding, agent software use and trajectory, and reference-based scoring for duties like summarization and translation. <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/manage-metrics#predefined-metrics\"><b>Adaptive rubrics<\/b><\/a> tailor the judging standards to every case as a substitute of making use of one brittle llm-as-judge immediate throughout inputs that do not deserve the identical questions. You too can outline your individual code-based or <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/manage-metrics#register_a_custom_llm_metric\">LLM-as-a-judge metrics<\/a> and retailer them in a single versioned, org-wide place so scoring stays constant and comparable over time.<\/li>\n<li data-block-key=\"3j2kt\"><b>Experiments.<\/b> Run them client- or server-side. Server-side retains each artifact in Cloud Storage, so runs are auditable and reproducible. Experiments combine with <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/evaluate-simulated#generate_scenarios_in_the_console\"><b>case technology<\/b><\/a> to bootstrap an analysis dataset, a <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/evaluate-simulated\"><b>consumer simulator<\/b><\/a> to play out multi-turn circumstances with out scripting every reply, and an <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/adk.dev\/evaluate\/environment_simulation\/\"><b>surroundings simulator<\/b><\/a> to face in for the programs the agent calls, so you&#8217;ll be able to emulate a failing or gradual backend with out affecting manufacturing.<\/li>\n<li data-block-key=\"dddo3\"><b>On-line displays and telemetry integrations.<\/b> <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/evaluate-online\">Steady analysis on reside manufacturing site visitors<\/a> grades the traces you already accumulate and produces <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/quality-alerts\">score-over-time charts and drift alerts<\/a>, while not having to arrange customized knowledge processing pipelines.<\/li>\n<\/ul>\n<p data-block-key=\"4l62k\">You&#8217;ll be able to attain evaluations from the <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/models\/evaluation-genai-sdk\">Agent Platform SDK<\/a>, <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/github.com\/google\/agents-cli\">agents-cli,<\/a> the<a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/pantheon.corp.google.com\/agent-platform\/agent-evaluation\"> Analysis<\/a> part in Agent Platform Google Cloud console, and immediately from <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/google.github.io\/adk-docs\/evaluate\/\">ADK<\/a>.<\/p>\n<h2 data-block-key=\"pd4nn\" id=\"experiments\"><b>Experiments<\/b><\/h2>\n<p data-block-key=\"bc5ed\">The first unit of labor in analysis is the <b>experiment<\/b>, a dataset of eval circumstances and a set of metrics and scores to run and assessment as you iterate. The analysis service comes with a versatile UI to outline metrics, kick off new runs, and assessment the outcomes, all the best way right down to a single failure, the place you&#8217;ll be able to open the agent&#8217;s full hint and session log to see precisely what occurred.<\/p>\n<\/div>\n<div class=\"inner-block-content video-block\">\n<p>        <video autoplay=\"\" loop=\"\" muted=\"\" playsinline=\"\" poster=\"https:\/\/storage.googleapis.com\/gweb-developer-goog-blog-assets\/original_videos\/wagtailvideo-64iumtpe_thumb.jpg\"><source src=\"https:\/\/storage.googleapis.com\/gweb-developer-goog-blog-assets\/original_videos\/Demo_1.mp4\" type=\"video\/mp4\"><p>Sorry, your browser does not help playback for this video<\/p>\n<p><\/source><\/video><\/p>\n<p>Evals Worksheet \u2014 assessment and run agent evaluations.<\/p>\n<\/div>\n<div class=\"inner-block-content rich-content\">\n<p data-block-key=\"b0jw6\">You&#8217;ll be able to run experiments <b>domestically<\/b> for quick iterations, or, in case your agent is already deployed on Agent Platform with telemetry enabled, you&#8217;ll be able to grade present <b>classes and traces<\/b>, or run the agent with the consumer simulator enabled to create traces earlier than your customers produce them. Each experiment artifact is saved transparently in <b>Cloud Storage<\/b>, so you&#8217;ll be able to model and audit your runs for compliance and posterity.<\/p>\n<p data-block-key=\"fev87\">For giant analysis jobs, the system helps <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/view-results#failure-analysis\"><b>subject clustering<\/b><\/a>: it teams eval failures into interpretable, actionable clusters towards your individual <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/view-results#loss_pattern_taxonomies\"><b>taxonomy<\/b><\/a> of failure causes. If you have not developed a taxonomy but, you need to use a pre-built one now we have for adaptive rubrics, which covers the widespread methods brokers go unsuitable.<\/p>\n<h2 data-block-key=\"jzxv8\" id=\"evaluation-metrics\"><b>Analysis metrics<\/b><\/h2>\n<p data-block-key=\"9vt2o\">Greater than 20 <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/manage-metrics\"><b>pre-built metrics<\/b><\/a> ship with the service. <b>Computation-based metrics<\/b> rating deterministically towards a ground-truth reference: <b>ROUGE<\/b> for summarization, <b>BLEU<\/b>, <b>MetricX<\/b>, and <b>COMET<\/b> for translation, precise match for extractive QA.<\/p>\n<p data-block-key=\"a16ub\">Past these, an <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/models\/evaluation-overview#adaptive-rubrics\"><b>adaptive rubric<\/b><\/a> is a sophisticated LLM-judge metric workflow co-developed with our analysis companions at Google DeepMind. It creates case-specific <b>move\/fail checks<\/b> (rubrics) from the eval case definition, the developer instruction, and the software declarations. It then grades the traces towards these rubrics, offering the decision and rationale per rubric.<\/p>\n<\/div>\n<div class=\"inner-block-content\">\n<div class=\"image-wrapper\">\n<p>                <img decoding=\"async\" class=\"regular-image\" src=\"https:\/\/storage.googleapis.com\/gweb-developer-goog-blog-assets\/images\/image1_3.original.png\" alt=\"image1 (3)\"\/><\/p>\n<p>\n                        Adaptive rubric with standards generated per analysis case.\n                    <\/p>\n<\/p><\/div><\/div>\n<div class=\"inner-block-content rich-content\">\n<p data-block-key=\"pxh3z\">We designed and calibrated variants for what you often wish to learn about an agent. <b>Job Success<\/b> grades purpose achievement throughout a dialog from observable outcomes and confirmations within the agent&#8217;s responses. <b>Instrument Use High quality<\/b> evaluates software choice, argument correctness, and schema compliance. <b>Security<\/b> scores the response towards content material insurance policies spanning hate speech, harassment, harmful content material, sexually express materials, and PII, returning the insurance policies violated. <b>Trajectory High quality, Remaining Response High quality, Hallucination, Grounding<\/b>, and picture and video high quality metrics are additionally obtainable. Pure-language tips <b>steer rubrics technology<\/b> towards standards you specify, with the ensuing rubric group reviewable and reusable throughout brokers.<\/p>\n<p data-block-key=\"38sv6\">When the pre-built metrics do not match, you&#8217;ll be able to convey your individual. A <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/manage-metrics#register_a_custom_code_metric\"><b>code-based metric<\/b><\/a> is a Python operate, overlaying precise textual content matches, JSON-shape checks, and anything expressible in code. An <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/manage-metrics#register_a_custom_llm_metric\"><b>LLM-as-a-judge metric<\/b><\/a> carries your individual standards, score scale, and decide mannequin. When working domestically you need to use any mannequin from any supplier, and server-side analysis runs help any <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cloud.google.com\/model-garden\">Mannequin Backyard<\/a> mannequin, together with all Gemini and Anthropic fashions.<\/p>\n<p data-block-key=\"2rl2o\">Nonetheless a metric is outlined, it lands in the identical versioned, org-wide registry, working unchanged in offline experiments and on-line displays. We&#8217;re actively increasing protection to extra specialised duties and enter modalities, and we welcome your suggestions.<\/p>\n<h2 data-block-key=\"pg80n\" id=\"online-monitors-and-telemetry-integrations\"><b>On-line displays and telemetry integrations<\/b><\/h2>\n<p data-block-key=\"dr7uf\">An agent that is inexperienced throughout your entire take a look at suite can nonetheless drift every week after launch, when confronted with inputs no person wrote a case for. That is not a testing failure a lot as a reality of manufacturing the place actual duties are larger and stranger than any suite you&#8217;ll be able to write up entrance.<\/p>\n<p data-block-key=\"5kteq\">The metrics you constructed when creating your agent ought to hold working after your agent launches. In the event you already <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/optimize\/evaluation\/evaluate-offline#evaluate_a_single_trace_or_session\">accumulate traces and classes by means of Cloud Hint<\/a>, you&#8217;ll be able to consider them immediately with only one click on within the traces UI. Higher nonetheless is to arrange an <b>on-line monitor<\/b> to grade the reside site visitors because it arrives. To keep away from grading each request and hold prices manageable, you need to use the built-in sampling and focused filters. When displays are lively you&#8217;ll be able to see scores over time in built-in dashboards and even arrange drift alerts for e mail, Slack, or different channels.<\/p>\n<\/div>\n<div class=\"inner-block-content\">\n<div class=\"image-wrapper\">\n<p>                <img decoding=\"async\" class=\"regular-image\" src=\"https:\/\/storage.googleapis.com\/gweb-developer-goog-blog-assets\/images\/2._2.original.png\" alt=\"2. (2)\"\/><\/p>\n<p>\n                        On-line displays present eval scores over on manufacturing traces.\n                    <\/p>\n<\/p><\/div><\/div>\n<div class=\"inner-block-content rich-content\">\n<h2 data-block-key=\"7s8ku\" id=\"case-generation-and-simulation\"><b>Case technology and simulation<\/b><\/h2>\n<p data-block-key=\"age22\">Each analysis wants take a look at circumstances to run towards, and writing all of them by hand is gradual and tends to overlook the non-obvious ones. The analysis service can generate them for you, with a <b>case generator<\/b>, a <b>consumer simulator<\/b>, and an <b>surroundings simulator<\/b>.<\/p>\n<p data-block-key=\"1dlfg\">The <b>case generator<\/b> seeds artificial eval circumstances from the agent&#8217;s directions and instruments, so you are not ranging from a clean web page or capping your protection at what you thought to sort out. For ADK brokers, it could pair with a <b>consumer simulator<\/b> to succeed in the eventualities which are laborious to provide by hand: you outline a persona and a brief dialog plan, and the simulator performs that consumer throughout a full multi-turn alternate, so you&#8217;ll be able to consider actual back-and-forth with out scripting every reply.<\/p>\n<p data-block-key=\"5ad1n\">The <b>surroundings simulator<\/b> stands in for the programs your agent calls. Level it at a software, give it the response you need \u2014 mocked knowledge, a pressured error, added latency \u2014 and it intercepts that decision through the run, so you&#8217;ll be able to take a look at how the agent handles a failing or gradual backend with out touching manufacturing.<\/p>\n<h2 data-block-key=\"9fv4f\" id=\"pricing-and-availability\"><b>Pricing and availability<\/b><\/h2>\n<p data-block-key=\"7iaps\">Agent and mannequin evaluations are typically obtainable. For supported areas and enterprise safety features see <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/scale#enterprise-security\">areas and safety features desk<\/a>. You pay normal charges for the mannequin calls behind LLM-as-a-judge and the opposite model-based metrics \u2013 plus Cloud Storage for the artifacts a server-side run retains. Code-based and computation metrics add no further value. Your datasets and traces all the time keep in your venture.<\/p>\n<h2 data-block-key=\"8wo05\" id=\"get-started\"><b>Get began<\/b><\/h2>\n<p data-block-key=\"3do5o\">Analysis meets you in no matter software you already construct brokers with. It is a part of the <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/docs.cloud.google.com\/gemini-enterprise-agent-platform\/models\/evaluation-genai-sdk\"><b>Agent Platform SDK<\/b><\/a>, with the identical operations obtainable over the <b>REST API<\/b> for the languages and pipelines the SDK does not attain.<\/p>\n<p data-block-key=\"1m42k\">In the event you work from the command line, <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/github.com\/google\/agents-cli\"><b>agents-cli<\/b><\/a> makes eval a first-class command subsequent to those you already use to deploy brokers and examine telemetry (ADK-Python brokers right now). You&#8217;ll be able to even run the loop straight out of your coding agent: a reusable talent walks Claude Code and related instruments by means of <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/developers.googleblog.com\/driving-the-agent-quality-flywheel-from-your-coding-agent\/\">the total agent-quality flywheel<\/a>. For groups constructing on <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/google.github.io\/adk-docs\/evaluate\/\"><b>ADK<\/b><\/a>, analysis is constructed into the framework \u2014 outline eval units and run them towards your agent domestically as you develop, together with inside pytest for CI. And once you&#8217;d moderately assessment a run with out writing code, the <b>Evals Worksheet<\/b> internet UI is the grid coated above.<\/p>\n<p data-block-key=\"ecelt\">We won&#8217;t wait to see what you ship.<\/p>\n<\/div><\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>Agent high quality should be measured throughout growth towards the circumstances you wrote, and after launch towards the duties the agent truly carried out. Agent and mannequin evaluations in Agent Platform are actually typically obtainable, letting you measure and examine brokers and fashions at each levels, on one engine with constant metrics. If you use [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":17359,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[56],"tags":[75,3128,6950,295,358,630],"class_list":["post-17357","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-software","tag-agent","tag-enterprise","tag-evaluations","tag-gemini","tag-model","tag-platform"],"_links":{"self":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/17357","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=17357"}],"version-history":[{"count":1,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/17357\/revisions"}],"predecessor-version":[{"id":17358,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/17357\/revisions\/17358"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/media\/17359"}],"wp:attachment":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=17357"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=17357"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=17357"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}<!-- This website is optimized by Airlift. Learn more: https://airlift.net. Template:. Learn more: https://airlift.net. Template: 69d9690a190636c2e0989534. Config Timestamp: 2026-04-10 21:18:02 UTC, Cached Timestamp: 2026-08-03 02:27:51 UTC -->