Skip to main content

A/B testing your videos

Written by Irek Khasianov

In short: we hide your video widgets from a small, random slice of your visitors and compare what those visitors spend against everyone else. After a few weeks you get a straight answer to "are these videos actually making me money?"

How it works

When you launch a test, every visitor to your store is sorted into one of two groups:

Group

What they see

Holdout

No video widgets at all — your store as if you'd never installed Storista

Videos on

Your store exactly as it is today

Which group a visitor lands in is decided from the anonymous ID already stored in their browser, so the same person always sees the same version — no flicker, no "the videos were there yesterday" support emails. The split happens at our edge network before the page finishes loading, so there is no delay and nothing flashes on screen.

We then measure both groups on:

  • Conversion rate — what share of visitors bought

  • Revenue per visitor (RPV) — the headline number, because it captures both "did they buy" and "how much did they spend"

Choosing your split

You decide how big the holdout is. This is the most important choice you'll make, and it's the one most people get backwards.

Split

What it means

Trade-off

50 / 50

Half your visitors see no videos

Fastest answer, most traffic affected

90 / 10

Only 1 in 10 sees no videos

Safest-feeling, but much slower

A smaller holdout is not the safer choice — it's the slowest one. The smaller group is what decides how long the test takes, because both groups need enough visitors before the difference between them means anything. A test that finishes in 68 days at a 50/50 split can take over 300 days at a 10% holdout. That's not a test; that's a year of waiting.

When you move the slider, we show you the estimated duration right there. If it warns that the test is impractical, move the slider closer to 50%.

How long it takes at your traffic level

Duration is driven by one thing: how quickly the smaller group fills up with visitors. So it depends on your traffic, not on how long the test has been switched on.

Rough guide, assuming a 2% store conversion rate and looking for a 15% improvement:

Monthly store visitors

50 / 50 split

90 / 10 holdout

1,000,000

~4 days

~2.5 weeks

500,000

~1 week

~5 weeks

100,000

~5 weeks

~5 months

50,000

~9 weeks

~11 months

10,000

~11 months

not realistic

Two things to read off that table:

  • High-traffic stores get answers fast. At a million visitors a month you can run an even split, have a trustworthy result inside a week, and be back to normal before the month is out. There's little reason to use a small holdout at that volume.

  • Low-traffic stores have to commit. Under ~50,000 visitors a month, a 10% holdout will not produce a conclusion in any useful timeframe. Run 50/50 or don't run the test.

These are estimates for detecting a modest 15% difference. If videos are doing much more than that for your store — or much less — you'll usually see it sooner, because a big difference stands out from the noise faster than a small one. Your store's own numbers also matter: a higher conversion rate shortens the test, a very spiky order value (a few huge orders among many small ones) lengthens it.

To calculate that estimate we need your store's visitor count and conversion rate, so the first time you open the form we'll ask permission to read your Shopify Analytics. It's read-only, we only use it for this estimate, and you can decline — the test still runs perfectly, you just won't see a duration prediction up front.

Sensitivity: how small a difference to look for

Next to the split you'll find a second dial — how big a difference the test should be built to detect. It is the other half of the duration equation, and it moves the numbers hard, because the required traffic grows with the square of the precision you ask for:

Setting

Visitors needed per group

Relative duration

Detect small changes (10%)

~80,700

~6x the fastest option

Balanced (15%) — recommended

~36,700

~2.7x

Detect only big changes (25%)

~13,800

fastest

(At a 2% store conversion rate.)

Pick "big changes only" if your traffic is modest and the question you actually have is "are these videos pulling their weight at all?" — you'll get an answer in weeks instead of months, at the cost of not being able to detect a small-but-real improvement. Pick "small changes" only if you have the traffic to spend on precision.

Picking a page (optional)

You can name a specific page — a product, a collection, your home page — as the page you care about.

This does not limit where the test runs. The holdout group has videos hidden everywhere, on every page, for the whole test. What the page setting does is give you a second, cleaner view of the results: "all visitors" versus "only the visitors who actually reached this page."

That second view is usually the one to trust. Plenty of visitors never land anywhere a widget appears, and including them dilutes the real effect.

What to expect after launching

Every test moves through the same stages, shown as a badge on the results page.

Stage

What it means

What to do

Collecting data

First day or two. Nothing reliable to show yet.

Let it run.

Gathering results

Numbers are visible but still moving. They can briefly look dramatic.

Don't act on them, however good they look.

Leaning

A real direction has appeared, but we're still holding something back — a full week of data, 300 orders per group, or the reading holding steady two days in a row.

Useful signal. Fine to act on if the stakes are low; otherwise wait.

Significant

The result held up. This is the call.

Act on it.

No difference

Two weeks in, nothing meaningful showed up.

A real answer. Move on.

Two things are worth knowing about why we make you wait.

A single good day is not a result. If we called a winner the first time the numbers crossed the line, we'd be wrong about one test in three — we simulated it. So a reading has to hold before we call it, which is why you may see "Leaning" for a while on a test that already looks decided.

We use the same minimums the rest of the industry does — at least a full week (to cover a whole weekly cycle) and at least 300 orders in each group before any verdict. If your test is stuck at "Leaning", the results page tells you which of those is still missing.

How long the early phase lasts is purely a matter of traffic. A store doing a million visitors a month is through it in a couple of days; at 50,000 visitors a month it's the first month or two. Use the table above to know which one you are — acting on an early swing is the single most common way to draw the wrong conclusion.

A "no difference" result is a real, useful answer, not a failed test. It means the effect is smaller than the test was built to see.

Reading the results page

  • Visitors / Conversions / Revenue per group — the raw counts. Revenue is your actual revenue, untouched.

  • Lift — how much better (or worse) the "Videos on" group did. Only meaningful once the badge says Significant.

  • Stage badge — see the table above. We use the 95% confidence standard, adjusted for the fact that you can check the results every day.

  • p-value under each metric — informational. A low one on its own is not a verdict; the stage badge is.

  • Chart — conversion rate and revenue per visitor per day, per group.

  • Tabs — "All visitors" and, if you named a page, "Visited <that page>".

One detail about revenue

A single enormous order can swing a whole group's average and make a test look decided when it isn't. For the statistics only, we cap unusually large orders (above the 97.5th percentile of that day's orders, applied identically to both groups). The revenue figures you see are never capped — that's your real money. This is standard practice and it makes tests conclude about 14% sooner.

Ending a test

When you end a test, you choose whether to roll out a winner. Rolling out simply means the holdout stops: everyone sees the videos again.

Ended tests keep their results, so you can come back to the numbers whenever you want. We also process one final day of data after you stop, so the last day isn't cut off halfway.

Rules while a test is running

Two things are locked once a test is live, and it's worth knowing why:

  • You can't change the split, the groups, or the target page. Your visitors have already been sorted using the current settings, and every day of results so far was measured with them. Changing them mid-test wouldn't adjust the test going forward — it would silently rewrite the days already recorded. To change the setup, end the test and start a new one.

  • You can't pause a live test. Pausing would put the videos back in front of the holdout group while the test kept counting them as if they'd never seen one, which quietly poisons the result. End the test instead — you keep all the data collected so far.

You can still rename a live test at any time.

One test at a time. Two overlapping holdouts split each other's groups, and neither produces a usable number, so we only allow one running at a time.

Troubleshooting

"My widgets disappeared!" — If you're in the holdout group, that's the test working correctly. Open your store in a private/incognito window to get sorted again, or check the test's results page, where both group sizes are shown.

"The numbers look identical." — With a small holdout in the first weeks, they often do. Check the estimated duration on the test; if it's very long, end the test and relaunch with a larger holdout.

"Revenue here doesn't match my Shopify reports." — These numbers are built from visitor sessions attributed to the test window only, and converted to a single currency. They will not tie out to Shopify's totals, and aren't meant to — what matters is the comparison between the two groups, which is measured identically on both sides.

"Can I see which group I'm in?" — Not in the dashboard. Assignment is anonymous by design; we never store a per-person record of it.

Did this answer your question?