A/B testing a widget
Run your draft against your live version with real traffic split between them, and get a verdict you can trust: the traffic floor, the detection limit, and the one-session-at-a-time limitation to know before you start.
A widget's Experiment tab (the seventh and last tab in its editor, right after Performance) lets you test a change before publishing it to every shopper. Your current live version (the champion) and your draft (the candidate) run side by side, each seeing a share of real traffic, and the tab tells you which one actually gets more shoppers to start a conversation.
This is different from the plain version comparison on the Performance tab, which is explicitly not a controlled test: it just compares two numbers from two different time windows. An experiment runs both versions at the same time, which is what makes the comparison trustworthy.
Before you can start
The tab shows a readiness checklist:
- A published version is live: you need something to compare against.
- A draft is ready to test: the change you want to try.
If either is missing, Start experiment stays disabled. There's one more condition, checked only when you actually try to start: your draft has to target the same pages as your live version. If your draft shows on different page types than your live widget, the checklist tells you so and blocks the start.
Why page targeting can't be part of the test
The test's headline number only counts shoppers who could have seen either version: the ones on a matching page. That only means the same thing for both versions if both target the same pages. If the draft targeted different pages, the two versions would be compared on different audiences and the result would be meaningless, even though the numbers would still show up as if it worked. So appearance, copy, size, placement and format are all testable. Which pages the widget appears on is not. That's a limitation of this first version, not an oversight.
The tab also shows how many views this widget got over the last 7 days, useful context before you commit, since that traffic is what will determine how long the test takes.
Starting a test
Start experiment opens a dialog that shows:
- The two versions being compared (champion and candidate, by version number).
- Traffic split: how much of your traffic goes to the candidate: 10%, 20% (the default), 30%, 40%, or 50%. The rest goes to the champion.
- The traffic floor: the minimum number of sessions each version needs before any result renders (see below).
- An estimated duration, calculated from this widget's own recent traffic at the split you chose. If the widget has no observed traffic yet, the dialog says a duration can't be estimated rather than guessing.
Once started, the split can't be changed. If you want a different split, end the current test and start a new one (see Ending a test for how). This isn't a UI limitation: each shopper's version is decided by a calculation run fresh on every visit, not stored anywhere, so changing the split partway through would silently move shoppers already in one version into the other, corrupting the count.
Publishing is paused while a test runs. You can still edit and save a new draft (that becomes a separate, third version that isn't part of the test), but you can't publish anything until the test ends, because publishing would replace the champion out from under a live comparison.
The method, briefly
The test uses mSPRT, an "always-valid" statistical method. In plain terms: you can check the result at any time (the moment you start the test, once an hour, once a day) without that habit making the result less trustworthy. That's not true of a simpler test, where checking early and often inflates how often it wrongly calls a winner. The trade-off is that always-valid testing needs more traffic to reach the same confidence a one-time check would need: roughly 60% more, in our case.
The traffic floor: 5,000 sessions per side
No result of any kind renders until both the champion and the candidate have each reached 5,000 assigned sessions: a session assigned to that version on a page the widget targets, whether or not the shopper ever actually noticed the widget. Below that, the tab shows both versions' raw numbers and how far each is from the floor, but no verdict, no percentage, no "leaning" indicator.
This isn't an arbitrary gate. At fewer sessions, the always-valid method can't reliably tell a real difference apart from ordinary day-to-day noise. 5,000 per side is the point where it can detect a solid improvement (see the numbers below) without the extra traffic cost of a stricter floor.
What "not yet conclusive" means, and doesn't mean
Once both versions clear the floor, the tab keeps evaluating and eventually reaches one of two states: conclusive (a winner is named) or, if you stop checking before that happens, still not yet conclusive.
"Not yet conclusive" is not "no difference." It means the test hasn't collected enough evidence yet to tell a real difference apart from noise, not that the two versions perform the same. A small real improvement can sit in "not yet conclusive" for a long time, simply because the test isn't sized to catch something that small.
Here's how reliably a real difference is caught, at the 5,000-session floor:
| If the candidate's true improvement is | The test catches it |
|---|---|
| No real difference (a false alarm) | 0.8% |
| +1 point | 7.8% |
| +2 points | 39% |
| +3 points | 78% |
| +5 points | 99.7% |
| +8 points or more | ~100% |
Read it this way: a large winner is almost never missed. A solid, +3-point improvement is caught about four times out of five. A one-point improvement is missed roughly nine times out of ten. A "not yet conclusive" result on a widget you changed in a small way is genuinely uninformative about whether that small change helped: it isn't evidence that it didn't.
Reading a running test
Below the floor, the tab shows both versions' assigned sessions and start rate side by side, how close each is to the 5,000-session floor, and a Stop experiment button. No verdict and no guardrail status yet, but you don't have to wait for the floor to end a test you've already decided against.
Once both versions clear the floor, the tab adds a Guardrails row confirming nothing unusual has been detected on the candidate so far (see below), and Promote candidate becomes visible, disabled until the result is conclusive, so you can always see where the test is headed, even before it gets there.
One traffic limitation to know about
Traffic is split by browser session, not by shopper. Someone who reopens your site later, or who has your storefront open in two tabs at once, can land in either version each time: the test has no way to recognize a shopper across sessions and keep them on the same version. This is a property of how the split works, not a bug, and it's shown directly in the tab so it's never a surprise.
Ending a test
A merchant ends a test one of two ways:
- Stop experiment (available any time, including below the floor) or Keep champion (offered once the result is conclusive). Same effect either way: the champion stays live, and your tested candidate goes back to being an editable draft (unless you already started a newer draft in the meantime, in which case the tested version is archived instead, so your newer work is never overwritten). Choosing to keep the champion over a statistically sound winner is a legitimate call (brand, tone, a reason that doesn't reduce to a number), and it's offered as an equal action next to promoting, never as a discard link.
- Promote candidate: available once the result is conclusive and the candidate is the winner. This publishes the candidate as the new live version, the same as a normal publish; the previous champion is archived and stays restorable.
If the candidate is the conclusive winner, the tab shows the measured effect (how many points better it did) alongside the promote/keep-champion choice.
When a test stops itself
A test can end automatically, without anyone clicking anything:
- A guardrail catches a problem with the candidate. The system watches the candidate for signs it's actively hurting the experience (failing to load noticeably more than the champion, or shoppers abandoning after opening the chat noticeably more than on the champion) and reverts all traffic to the champion the moment it sees one, rather than waiting for a merchant to notice. The exact thresholds aren't shown; what matters is that this is a safety net, not part of the result.
- The traffic split itself looks broken. If the actual share of sessions landing on each version drifts away from what you configured, that's a sign the split mechanism itself may be malfunctioning, in which case every number the test has produced is unreliable, not just disappointing. This ends the test as invalidated rather than merely paused, and no rate or figure from it is shown as a result.
Either way, the tab names what happened and states plainly that the collected numbers can't be trusted as a result, never a rate that looks like a real outcome.
Switching the widget off, or deleting it, also ends its running test, for a reason you already know, since you're the one who did it.
Where else this shows up
- The Versions tab labels the version under test Under test rather than Draft, and shows a banner that publishing is paused.
- The site's widgets list marks a widget currently under test, without a second column or a per-widget verdict; that stays on the widget's own Experiment tab.
- The site's Engagement report lists every widget with a test currently running, read-only, with a link back to that widget's Experiment tab, useful since a test can run for weeks and is easy to forget about between visits.
Widget performance
One widget's own funnel: views, opens, starts, its start rate, messages per conversation, and a version-over-version comparison that is explicitly not a controlled test.
Drafts, publishing and versions
Editing a widget writes a draft that shoppers never see. Publish sends it, the previous version is archived rather than discarded, and Restore brings one back for review.