Skip to content

Retail and consumer goodsBI, AI and ML for the business

A/B test with its p-value

One page converts better than the other, but the difference may be pure chance. In Muvia a statistical test written in Python is recomputed every morning on the cumulative data, Analysis shows when the difference becomes significant and Athena writes the conclusion in a notebook.

Illustrative scenario: it describes a typical case, not a customer project.

B converts better. Is it real, or chance?

The team put a second version of the page online and split traffic in half. After a few days B is ahead and someone already wants to adopt it; a few days later it is behind. Checking the conversion rate every morning and deciding on a hunch is the fastest way to pick the wrong page.

What you need is a statistical test redone on the cumulative data, with the confidence interval next to the difference, and a clear rule for when the difference is solid enough to decide.

Who it's for
E-commerce and product managers, marketing, analysts who design experiments.
Visits and conversions of the two variants from a REST API, orders and revenue from the online store's database on PostgreSQL and the experiment log from Google Sheets flow into Muvia, where a Python function runs the statistical test every morning; out come the result table, an analysis of the p-value over time, the notebook with the conclusion and Athena's answers.

In Muvia, step by step

Real product screens, recorded on a project with sample data.

1 of 5

The video · One page converts better than the other, but the difference may be pure chance. In Muvia a statistical test written in Python is recomputed every morning on the cumulative data, Analysis shows when the difference becomes significant and Athena writes the conclusion in a notebook.

What you get

Decisions made on numbers
A page is adopted when the difference is significant, not when it seems to have been ahead for a few days.
Uncertainty written alongside
Difference, confidence interval and p-value sit on the same row, updated every morning.
The test reruns itself
The flow recomputes the result on the cumulative data every day, without anyone reopening a spreadsheet.
A conclusion ready to share
The notebook with results and decision is written by Athena and always reads the latest result.

For the technical team

How it is built in Muvia

  1. 1

    Collect visits and conversions

    The sources bring visits, conversions and revenue per variant hour by hour; a query puts them into one series for the experiment.

  2. 2

    Write the test in Python

    A Python function runs a two-proportion z-test on the cumulative data: the two variants' rates, the difference in points with its 95% interval, and the p-value. The significance level is a parameter.

  3. 3

    Schedule the test

    A flow runs it every morning at 7 and writes the result to a managed table, one row per day.

  4. 4

    Follow the experiment in Analysis

    The two pages' rates and the p-value are stacked over time: you see the day the curve drops below the threshold and stays there.

  5. 5

    Let Athena write the conclusion

    With one question, Athena saves the comparison query in the Lab and creates a notebook with the results table and the decision.

Python function

import numpy as np
import pandas as pd
from scipy.stats import norm

def transform(inputs, params, ctx):
    d = inputs["rows"].copy()
    d["day"] = pd.to_datetime(d["hour"]).dt.floor("D")
    g = d.pivot_table(index="day", columns="variant",
                      values=["visits", "conversions"], aggfunc="sum").cumsum()
    va, vb = g["visits"]["A"], g["visits"]["B"]
    ca, cb = g["conversions"]["A"], g["conversions"]["B"]
    pa, pb = ca / va, cb / vb
    p = (ca + cb) / (va + vb)
    z = (pb - pa) / np.sqrt(p * (1 - p) * (1 / va + 1 / vb))
    se = np.sqrt(pa * (1 - pa) / va + pb * (1 - pb) / vb)
    out = pd.DataFrame({
        "rate_a": 100 * pa, "rate_b": 100 * pb,
        "difference_pp": 100 * (pb - pa),
        "ci_low_pp": 100 * (pb - pa - params["z_critical"] * se),
        "ci_high_pp": 100 * (pb - pa + params["z_critical"] * se),
        "p_value": 2 * norm.sf(np.abs(z)),
    }).round(4).reset_index()
    out["significant"] = (out["p_value"] < params["alpha"]).astype(int)
    return {"rows": out}
The two-proportion z-test on cumulative data, day by day: difference, 95% interval and p-value.
The data it needs
  • Hourly visits and conversions per variant, read through a REST API
  • Orders and revenue per variant, from the online store's database on PostgreSQL
  • Experiment log with variants, dates and hypotheses, from Google Sheets
Parts of Muvia used

Bring us a question you can't answer today.

We start from a real question your business has and walk the path from source to dashboard with your systems, not a demo dataset.