iAdvizeDocs
Agents

Evaluation

Run scripted shopper scenarios against an agent version and review pass/fail results with transcripts, instead of testing by hand in the playground.

Evaluation runs a library of scripted shopper conversations against one specific agent version and grades each one (pass or fail, with the full transcript) so you can check a version behaves correctly without replaying every case by hand in the playground.

A run always executes against the chosen version's own live instructions, connectors, and chat tools, never a stand-in configuration. Once launched, a run keeps its result forever, pinned to that exact version; creating a new version later never changes a past run's outcome.

Where to find it

  • Evaluations in the dashboard sidebar (after Playground, before Connectors) opens the hub: every past run across every agent on the site, newest first.
  • The Evaluations tab on an agent's own detail page shows that agent's run history and a New run button that pre-fills the agent on the launch screen.

Launch a run

Open Evaluations (from the sidebar, or the New run button on an agent's own Evaluations tab) and choose New run.

Pick the agent and the version to test. The version defaults to the champion, or the newest version if the agent has none.

Tick the scenarios to run, grouped by category. You can mix native and custom scenarios in the same run, up to 50 scenarios.

Choose Run evaluation. You're taken to the results screen, which updates on its own while the run is in progress.

Native and custom scenarios

Every scenario belongs to one of two sources:

  • Native scenarios ship with the product, curated by iAdvize, and carry a Native lock badge. You can select or leave out a native scenario for a given run, but you can't edit or delete one. They're grouped into an Assistant category (prompt injection resistance, brand-voice consistency, language matching) that applies to every agent, and a Sales category (product discovery, product comparison, promotions, edge cases) that only applies to Shopping assistant agents.
  • Custom scenarios are your own, grouped under a Custom category you organize into subcategories. Add one with + Add scenario on the launch screen, and edit an existing one with the pencil icon next to it.

Author a custom scenario

On the launch screen, under Custom, choose + Add scenario.

Pick an existing subcategory or create a new one to file the scenario under.

Write the intent (a short label for what the scenario tests) and one or more shopper turns: the message(s) sent to the assistant, in order.

Optionally add an expected behavior description: free text describing what a good reply looks like. When set, an LLM judge grades the assistant's final reply against it.

Optionally add one or more deterministic checks against the assistant's final reply: a specific tool must have been called, specific text must or must not appear, shown products must stay under a price ceiling, or the reply must be in a given language.

Save. The scenario is added to the checklist and auto-selected for this run.

A scenario needs at least one of the two

A custom scenario can rely on deterministic checks alone, an expected-behavior description alone, or both together. It fails if any deterministic check fails, or (when it has an expected-behavior description) if the judge's score doesn't clear the bar.

Read the results

The run history (hub and agent tab) shows each run's status (Pending, Running, Completed, or Failed) and, once every scenario has a verdict, how many passed out of the total. While a run is still in progress, it shows how many scenarios have finished instead.

Open a run to see:

  • A pass ratio per category (Assistant, Sales, Custom).
  • An All / Failed filter over the individual scenario results.
  • For each scenario: its pass/fail verdict, the outcome of every deterministic check it declared, the judge's written reasoning (only for scenarios with an expected-behavior description), and the full transcript of the replayed conversation.

Runs execute in the background

A run doesn't block your browser tab: it runs as a background job, and the results screen refreshes itself automatically until every scenario has a verdict. You'll also get an email once it's done: a pass/fail summary for a completed run, or a failure notice, either way with a link back to the run.

On this page