iAdvizeDocs
Conversation labeling

Conversation labeling

How an AI judge reads finished conversations in the background and records structured labels about them, how to read them in a conversation's Analysis section, when it runs, and how the data is handled.

Conversation labeling is a background job that reads a conversation once it has gone quiet and records a structured set of labels about it: whether the shopper's need was met, what the conversation was about, whether the assistant's replies were accurate, on brand and within your instructions, what the shopper objected to, and what was missing. An AI model, the judge, answers a fixed list of questions about the conversation. Every answer is a yes/no, a pass/fail, or one value from a closed list, and every negative answer must point at the messages or products it's based on.

Each record also carries scores and a review priority computed from those answers, so the labels can single out the conversations worth a person's review. You read them in the Analysis section of a conversation's detail page (see The Analysis section), and you can find conversations by them in the Conversations list (see Find conversations by label).

Where this stands today

  • Conversations on the Web channel (your storefront), on Hosted Page (your public link) and in the Playground are labeled, by the hourly pass in production and on demand with Label now in any environment. There's no setting to turn it off or to pick channels. The exception is a conversation whose shopper objected to measurement: it isn't labeled from the next pass on.
  • The labels are in Beta. They're shown in the Analysis section of a conversation, marked "AI-generated labels, not yet validated", and read by the Conversations list's Resolution column, filters and Needs review tab. You can also read them over the REST API and through the dashboard MCP server, which filters conversations by label and returns each label record.

For the full list of labels and what each value means, see the Label reference. For how they're named in the dashboard, see How labels appear in the dashboard.

What a label record contains

Each time a conversation is labeled, one label record is stored with it:

  • The judge's answers, grouped into resolution, conversation types, service and trust criteria, shopper signals, purchase objections, gaps, and the order issue. See the Label reference.
  • The evidence behind each negative answer: the identifiers of the messages or products it relies on. Never quoted text.
  • Facts read directly off the transcript by code, which the judge is told to take as established: the language the shopper wrote in, the ratings given on replies, which catalog searches returned results, tool calls that errored, empty or error replies, replies that sent the shopper to email, phone or a form, repeated messages, the products shown, the instruction blocks the assistant loaded, the page the conversation started from, and the starter question tapped, if any.
  • Values computed from the labels, again by code: a quality, a service and a trust score, a review priority, and a few summary values such as why a need went unmet. See Values computed from the labels.
  • Which judge produced it: the taxonomy version (currently v1.1), the judging configuration, and the last message it read.

Which conversations are labeled

A conversation is labeled when all of these hold:

  • It's on the Web, Hosted Page or Playground channel. The judge reads shopper conversations as well as the ones you and your team have while testing an agent.
  • The shopper hasn't objected to measurement. A conversation where the shopper objected isn't labeled from the next pass on (see Privacy and data processing).
  • It has gone quiet. No new message for 30 minutes at an hourly pass, and for 1 minute when someone clicks Label now.
  • Its last activity is recent. Within the last 30 days. A conversation idle for longer is never picked up, including the ones that already existed when labeling started.
  • The assistant replied at least once. A conversation with only a shopper message has nothing to judge.
  • Its current state isn't labeled yet. A conversation already labeled, with nothing new since, is skipped.
  • It's under its labeling limits (see below).

Every eligible conversation is labeled. There's no sampling.

When it runs

In production, labeling runs once an hour. Each pass labels the conversations that have become eligible since the last one, conversations never labeled before going first, oldest activity first. Each pass handles a bounded number of conversations per organization; if more are waiting, the rest are picked up by the following passes.

So a conversation is typically labeled at the first hourly pass after it has been idle for 30 minutes. Each conversation is judged on its own, in the background.

You don't have to wait for the next pass: Label now on a conversation starts labeling it straight away, in every environment. On a test (non-production) environment there's no hourly pass at all, so that button is the only way labeling runs there.

New activity means a new label

A label describes the conversation as it was when the judge read it. If the conversation changes afterwards, it's labeled again:

  • A new message (the conversation was resumed).
  • A new rating, or a revised one, on any of its assistant replies (see Rate a message). The judge reads the ratings, so a thumbs-down given after the first label is reflected in the next one.

The new label becomes the conversation's current label. The previous one is kept, marked as superseded, never edited or overwritten. A conversation always has at most one current label per judging configuration.

A conversation is labeled at most 5 times. After its fifth label, further activity no longer triggers a new one, and the fifth stays current.

When labeling fails

The judge's answer is checked before anything is stored: every label present, every value from its allowed list, the consistency rules respected, and every piece of evidence pointing at a message or product that is really in the conversation. If the first answer breaks a rule, the judge is told what was wrong and gets one more try.

An answer that still doesn't pass is stored as a failed run with a reason code (for example, the answer didn't match the label list, or cited evidence that isn't in the conversation). A failed run is never patched with default values, never becomes the current label, and holds no conversation text.

  • A conversation that fails is tried again at the next pass. After 3 failures on the same state of the conversation, it's left unlabeled until it has new activity.
  • When the model provider is briefly unavailable, the attempt is retried automatically. If it still can't get through, the conversation waits a few hours before it's tried again. These attempts don't count toward the 3 failures.
  • A very long conversation can exceed what the model accepts in one request. It then fails the same way every time and stays unlabeled.

It never touches the conversation

Labeling happens after the fact, separately from the chat. It doesn't change what the assistant says, in that conversation or any other, and it doesn't feed back into your agent's configuration.

It also doesn't change the conversation itself: the transcript, its title and its last-activity date stay as they were. So labeling doesn't move a conversation in the Conversations list, and doesn't extend how long it's kept.

What the judge reads

To judge a conversation against your own setup, the judge receives:

  • The transcript: every message, including the assistant's tool calls and their results.
  • The ratings on assistant replies: thumbs-up or thumbs-down, and the issue category picked with a thumbs-down. Not the free-text comment.
  • The configuration of the agent version that answered: its instructions and instruction blocks, and the tools it had enabled, including the tools of its connectors as last listed. Never a connector's credentials.
  • Your site's details: its vertical, its site information and its tone of voice.
  • The facts computed by code listed above.

The transcript is treated as data, never as instructions: a message asking the judge to change a label or its rules is ignored.

Privacy and data processing

The model. The judge is an OpenAI model, reached through the Vercel AI Gateway. Every call:

  • is restricted to OpenAI as the provider, so no other model provider sees the conversation,
  • is pinned to the Gateway's European inference zone, and the region that actually served it is recorded with the label,
  • carries zero data retention and a no-training instruction.

The sub-processor page lists OpenAI for this purpose, with what it processes and the limits of these commitments (for example, OpenAI's own handling of requests its abuse-detection systems flag).

Where labels live. In our database, next to the conversation they describe. A label record contains the answers, the evidence identifiers and the computed facts; the conversation text itself stays in the conversation.

Labels leave with their conversation. A label is never kept longer than the conversation it describes:

  • deleting a conversation, from the dashboard or by the shopper from the chat panel, deletes its labels with it (see Delete a conversation),
  • when a conversation reaches the end of its 12-month retention period, its labels are purged with it.

The shopper's measurement objection. When a shopper objects to measurement, that objection is recorded on their conversation, and once recorded it stays. From the next pass on the conversation isn't labeled, including when it's resumed or rated later. A labeling already under way when the objection arrives may still finish. Labels written before the objection are kept, not deleted: the objection stops new labeling, it doesn't remove what already exists. A conversation with no recorded objection is labeled, which includes conversations stored before the objection was recorded.

The Analysis section

Open a conversation from the Conversations list. The rail beside the transcript is a stack of expandable sections. Analysis, with a Beta badge, is one of them, open by default alongside Details (see the rail). Any member who can open the conversation can read it: no permission is needed, only Label now is gated.

From top to bottom, a labeled conversation shows:

  • Resolution. Green for Resolved, amber for Partly resolved, red for Unresolved. No real need and Visitor left read in plain text.
  • Review priority, only when the record has a review reason. High (shown in red) means a trust criterion failed, Medium that the need was not met, the shopper was frustrated, or a person was needed, Low that the assistant used an unneeded fallback. For High and Low, the callout names the failed criteria behind it; for Medium it names none, because that reason comes from the resolution and the shopper, not from a criterion. This is the review priority stored with the record. A conversation with no review reason shows no callout.
  • Quality, Service and Trust scores, each as passed/judged (for example "3/4 passed"). Only criteria judged pass or fail are counted; a criterion that didn't apply is left out, and a score with nothing judged isn't shown. The panel shows counts, not the decimal shares stored with the record.
  • A "Show all labels" switch, off each time you open the page and never remembered.
  • The labels, grouped by block: Conversation type, Criteria, Signals, Objections, Content gaps, Order issue. With the switch off, a block lists only what needs attention: failed criteria and every label answered yes or with a value. Passed criteria are summarized as a count ("3 passed"). A block with nothing to list stays collapsed, title only. With the switch on, every block opens and lists every answered label, including passed and not applicable criteria and the "no" answers.
  • Colors. Only three things are colored: the resolution (above), a criterion (green when passed, red when failed) and a content gap answered yes (amber). Everything else reads in plain text, including a purchase objection the shopper raised. An objection is a normal part of a buying conversation, not a fault, so it isn't flagged.
  • A provenance line: when it was labeled, in your time zone, and the taxonomy version it was written with, followed by "AI-generated labels, not yet validated."

An answer that cites messages shows each one under its row as a short excerpt, the first 60 characters of the message. Click it and the transcript scrolls to that message and outlines it, with a "Cited by" tag naming the label and its answer, for example "Cited by: Stuck to known facts · Failed". The outline stays until you click another evidence link or click in the transcript. On a narrow screen, the Details panel closes so the transcript is in view.

  • A message that isn't in the transcript any more reads "Message no longer available" and can't be clicked.
  • A cited product shows its name when the conversation's tool results show it, and "Product" otherwise. It isn't a link.

What the section says when labels are missing or out of date

StateWhat you see
Waiting"Not labeled yet. Labels appear 30 to 90 minutes after the conversation goes quiet." That's the production schedule: a 30-minute quiet period, then the next hourly pass. On a test environment, only Label now runs it.
Out of dateThe labels of an earlier version, dimmed, under "This conversation continued after it was labeled. These labels describe an earlier part of it." It's labeled again at a later pass, or with Label now.
Last attempt failedA Failed badge, what went wrong in plain words, the reason code (for example provider_error), when it was attempted and "It will be retried automatically." Only the first reason code is shown.
Labeled, then the shopper objectedThe labels stay, under "The shopper objected to measurement after this conversation was labeled. It will not be labeled again."
Unavailable"Labels are unavailable right now." The labels couldn't be read, or the record can't be read against a known taxonomy version. The rest of the page is unaffected.
Not eligibleOne sentence for the first condition that isn't met, listed below.

A conversation that isn't labeled shows the first of these that applies, in this order:

  1. "The shopper objected to measurement, so this conversation is never analyzed."
  2. "The assistant never replied, so there is nothing to label."
  3. "This conversation is older than 30 days and is not labeled."
  4. "This conversation has been labeled the maximum number of times." (5 labels, see New activity means a new label.)
  5. "Labeling failed the maximum number of times for this conversation." (3 failures, see When labeling fails.)
  6. "Only a sample of conversations is labeled, and this one is not in it."
  7. "Conversations on this channel are not labeled."

The last two sentences exist in the dashboard, but with today's configuration every eligible conversation and all three channels are labeled (no sampling), so neither appears.

Find conversations by label

The Conversations list reads each conversation's current label in three places. All three are in Beta and read-only.

  • A Resolution column shows the resolution, or "Not labeled" when the conversation has no current label. See The Resolution column.
  • Two filters, Resolution and Issues, narrow the list to the conversations with a given resolution or problem. A conversation with no current label never matches one. See Filter by label and how the filters map to labels.
  • A Needs review tab lists the conversations that have a review priority, trust first, with a line saying whether labeling is keeping up. See Needs review.

None of this corrects a label or assigns a conversation. Human correction isn't available yet.

Label now

Label now is a button beside the Analysis title. It labels this one conversation straight away instead of waiting for the next hourly pass.

  • Who sees it. Members with the Conversations: Edit permission (the built-in Admin role has it; see Roles). Everyone else doesn't see the button at all.
  • When it's there. When a run could start now: the conversation is waiting, out of date, or its last attempt failed and it can be tried again. When the label is already up to date, the button is disabled, and a tooltip says "Already labeled on its latest version." In every other state (not eligible, unavailable, labeled then objected), there's no button.
  • What it labels. The conversation you're on, on any of the three labeled channels, in every environment, production included. It uses the same rules as above, with a 1-minute quiet period instead of 30.
  • A conversation still in progress. One quiet minute is enough, so a click can label a conversation someone paused but hasn't finished. Its next message makes it eligible again. Each new label counts toward its limit of 5.
  • What you see. A message, "Labeling started. The labels appear in a few minutes.", the button turns to "Labeling…" and stays disabled, and the section says "Labeling in progress. The labels appear here when it finishes." The labels are produced in the background: the page checks for them every 8 seconds and shows them as soon as they land, or shows the failure if the attempt fails, and the notice clears. A reload keeps the notice. If nothing has arrived after about five minutes, the notice goes away and the section returns to its usual state. Your browser remembers the notice for that tab only: a colleague opening the same conversation, or you in another tab or browser, doesn't see it.

If the request doesn't go through, a message says why:

MessageWhat it means
"You do not have permission to label conversations."Your role doesn't include Conversations: Edit. The attempt is recorded in the audit log.
"This conversation was not found."The conversation doesn't exist in this site of your organization.
"Already labeled on its latest version."Its label was already up to date when the request arrived.
"This conversation cannot be labeled right now."A condition isn't met, or the conversation was active within the last minute, so nothing was queued.
"Labeling could not be started. Try again."The request couldn't be handled.

Each request that passes these checks is recorded in the audit log as Labeling requested for a conversation. The entry targets the conversation and carries how many conversations were selected and queued, never any conversation text.

The Conversations list has no bulk button for this. Labeling on demand is per conversation. To see whether the hourly pass is keeping up, use the freshness line on the list's Needs review tab.

Label status

A conversation is always in one of four states. The state is worked out when you read it, from the conversation's runs, and the API returns it as labelStatus on every conversation.

The Analysis section of a conversation in the dashboard words more situations than these four statuses, so the two don't map one to one. Rely on the status when you integrate.

StatusWhat it meansWhat to do as an integrator
labeledThe conversation has a current label record. labels is filled on the labels endpoint, and labelSummary on the conversation.Read the labels. Compare labeledAt with the conversation's updatedAt to see whether a message arrived after the label.
pendingNothing is stored yet and the conversation isn't ruled out. It may be waiting for its quiet period or for the next hourly pass. A conversation the pass can't label, such as one the assistant never replied to, stays pending until it is 30 days old.Check again later. Don't treat it as an error.
failedNo current label record, and at least one labeling attempt was stored as failed. The API doesn't say why. It may be retried at a later pass (see When labeling fails).Check again later. If it stays failed, it is unlabeled until it has new activity.
not_eligibleNothing is stored, and this conversation will not be labeled as things stand.Don't wait for labels. The API doesn't say why.

What to know:

  • labeled can be out of date. A new message or rating doesn't remove the current label record. The conversation keeps reading labeled with its previous labels until the next pass replaces them (see New activity means a new label). Comparing labeledAt with updatedAt catches a new message only: a rating given after the label makes it out of date for the next pass but doesn't change updatedAt. After a conversation's fifth label, the fifth stays current.
  • A label isn't erased when a conversation becomes ineligible. A conversation that already has a current label record reads labeled, even if it would no longer be selected today.
  • not_eligible doesn't give its reason. The status means this conversation will not be labeled. The API doesn't say why.
  • A failed attempt after a label changes nothing. If a conversation already has a current label record, a later failed attempt leaves it labeled.

Read labels over the API

Three endpoints return labels. All take an API key. Labels are read under the current taxonomy, v1.1: the API returns nothing written under another version.

Labels on a conversation

The conversation list and Get a conversation carry labelStatus and, for a labeled conversation, a short labelSummary: the resolution, the review reason, the three scores and the list of criteria the judge failed. For the other three statuses labelSummary is null.

The full record

GET /api/v1/sites/{siteId}/conversations/{id}/labels answers with the status and the record:

{
  "data": {
    "conversationId": "conv_01kvnwmagvf27rc9cze5rm9e1r",
    "labelStatus": "labeled",
    "taxonomyVersion": "v1.1",
    "labels": {
      "labeledAt": "2026-09-12T09:30:02.000Z",
      "resolution": "unresolved",
      "reviewReason": "service",
      "reviewPriority": 3,
      "scores": { "quality": 0.9, "service": 0.8, "trust": 1 },
      "answers": {
        "visitor_frustrated": {
          "value": false,
          "source": "judge",
          "evidence": []
        },
        "no_unjustified_fallback": {
          "value": "fail",
          "source": "judge",
          "evidence": ["msg_01kvnwmagvf27rc9cze5rm9e1r"]
        }
      },
      "tokens": ["no_unjustified_fallback:fail", "resolution:unresolved"]
    }
  }
}

The values are illustrative. answers and tokens are cut here: the real answers has an entry for every label of the Label reference.

  • answers has one entry per label, keyed by the label's name. value is true or false, pass, fail or not_applicable, or one value from a closed list. A conditional label that wasn't answered is null. evidence lists the identifiers the answer relies on, never quoted text, and is empty when there is none.
  • source is judge on every answer today. The field exists so a person's correction could appear later without a change of shape.
  • scores are the quality, service and trust scores, each from 0 to 1, or null when none of its criteria was judged. reviewReason is trust, service, fallback, or null. reviewPriority is a number where lower means reviewed first, or null when the conversation isn't in the review queue. The numbers 0 and 1 are the trust reason, 2 and 3 service, 4 and 5 fallback; in each pair the lower number is a conversation with purchase intent. See Review priority. The dashboard's High, Medium and Low come from the review reason alone (trust, service, fallback). reviewPriority is the finer queue rank, an integer where lower comes first, so don't derive one from the other.
  • tokens are the filter tokens of this conversation, derived values included. They are what the labels.* filters of the list match against.
  • labeledAt is when the record was written.

When the status isn't labeled, labels is null and the response carries no reason. A conversation of another site or organization, and an unknown id, answer 404 with the code not_found.

The record leaves out internal run data: the model, its cost, the failure reasons and the facts read off the transcript.

The vocabulary

GET /api/v1/labels/vocabulary lists what the filters accept. It needs an API key but no site id:

{
  "data": {
    "taxonomyVersion": "v1.1",
    "blocks": [
      { "key": "resolution", "displayName": "Resolution" },
      { "key": "criteria", "displayName": "Criteria" }
    ],
    "labels": [
      {
        "key": "resolution",
        "displayName": "Resolution",
        "block": "resolution",
        "kind": "categorical",
        "values": [
          "resolved",
          "partly_resolved",
          "unresolved",
          "no_real_need",
          "visitor_left"
        ],
        "valueNames": {
          "resolved": "Resolved",
          "partly_resolved": "Partly resolved",
          "unresolved": "Unresolved",
          "no_real_need": "No real need",
          "visitor_left": "Visitor left"
        }
      },
      {
        "key": "grounded",
        "displayName": "Stuck to known facts",
        "block": "criteria",
        "kind": "criterion",
        "values": ["pass", "fail"],
        "valueNames": { "pass": "Passed", "fail": "Failed" }
      }
    ],
    "derivedTokens": [
      "sales_phase:pre_sales",
      "content_gap:true",
      "containment:contained"
    ],
    "documentationUrl": "https://www.iadvize.ninja/docs/product/conversation-labeling"
  }
}

The lists are cut here. Each label has a kind, and a token is <key>:<value>:

kindTokens it can produce
boolean<key>:true only. A false answer produces no token: filter it with labels.none.
criterion<key>:pass and <key>:fail. not_applicable produces none.
categorical<key>:<value> for each value of the list. A conditional label that wasn't answered produces none.

derivedTokens are the values computed from the labels, such as containment:contained or unmet_reason:catalog_gap. They are filtered like any other token. A token that isn't in the vocabulary makes a filter fail with invalid_filter.

Labels are the ones the Label reference explains. The taxonomy's instructions and rules stay private: the vocabulary gives keys, kinds, blocks, values and display names only.

displayName, valueNames and the blocks names are English text, there to explain a label to a person or to an assistant answering one. They are not identifiers and can be reworded: match on key and on the tokens, never on a displayName.

Questions

On this page