Skip to content

Healthcare and clinical researchBI, AI and ML for the business

Clinical trial: the interim analysis

A trial's visits, measurements and adverse events come from the data capture system, the randomisation list from a file. Muvia joins them patient by patient, a Python function reruns the statistical comparison at every visit, and the interim report for the committee is always up to date.

Illustrative scenario: it describes a typical case, not a customer project.

Every interim analysis starts again from a different spreadsheet.

In a clinical trial, measurements, visits, adverse events and randomisation sit in different systems. Every time the committee asks for an update someone exports, joins, reruns the tests and lays out the report, and two exports a month apart are never done the same way.

Writing the join and the statistical comparison once, and rerunning them at every visit on the same data, makes each report comparable with the last and leaves biostatisticians time for the new questions.

Who it's for
Trial biostatisticians and data managers, medical monitors, sponsors and monitoring committees.
Exports from the trial's data capture system over SFTP, the randomisation list in Excel and adverse events from a PostgreSQL database flow into Muvia, where the team's function compares the arms; out come the table of results by visit, a trial dashboard and the interim report as a PDF for the committee.

In Muvia, step by step

Real product screens, recorded on a project with sample data.

1 of 5

The video · A trial's visits, measurements and adverse events come from the data capture system, the randomisation list from a file. Muvia joins them patient by patient, a Python function reruns the statistical comparison at every visit, and the interim report for the committee is always up to date.

What you get

Comparable reports
The join and the tests are written once: every report is built the same way, on the latest data.
Statistics you can reread
The comparison is a function biostatisticians can reread version by version, not a chain of manual steps.
Efficacy and safety together
Change from baseline and adverse events are read in one place, by visit and by arm.
More time for new questions
The recurring analysis runs by itself; the team works on what the committee asks next.

For the technical team

How it is built in Muvia

  1. 1

    Bring the trial data into Muvia

    Exports from the data capture system arrive as files over SFTP, the randomisation as an Excel file and adverse events with an incremental copy. Every source becomes a table.

  2. 2

    Work out the change from baseline

    A query joins measurements and randomisation and works out, for every patient and visit, the change from baseline. Save it as a dataset with stable fields.

  3. 3

    Compare the arms in Python

    A function using scipy works out, for every visit, the difference between arms, Welch's t, the 95% confidence interval and Cohen's d. A flow runs it every Monday and writes the results to a managed table.

  4. 4

    Look at the two populations

    A dashboard shows the distribution of changes by visit and arm, the mean trend and adverse events by type.

  5. 5

    Prepare the committee report

    A notebook with text, indicators, the per-visit table and the adverse events becomes a PDF for the committee, rebuilt on the latest data.

Python function

import numpy as np
import pandas as pd
from scipy import stats

def transform(inputs, params, ctx):
    df, rows = inputs["rows"], []
    for visit, g in df.groupby("visit"):
        a = g.loc[g["arm"] == params["active_arm"], "change"].dropna()
        p = g.loc[g["arm"] == params["control_arm"], "change"].dropna()
        t, pval = stats.ttest_ind(a, p, equal_var=False)
        va, vp = a.var() / len(a), p.var() / len(p)
        dof = (va + vp) ** 2 / (va ** 2 / (len(a) - 1) + vp ** 2 / (len(p) - 1))
        diff, half_ci = a.mean() - p.mean(), stats.t.ppf(0.975, dof) * np.sqrt(va + vp)
        rows.append({"visit": visit, "difference": diff,
                     "ci_low": diff - half_ci, "ci_high": diff + half_ci,
                     "t": t, "p": pval,
                     "cohen_d": diff / np.sqrt((a.var() + p.var()) / 2)})
    ctx.logger.info("visits analysed: %d", len(rows))
    return {"rows": pd.DataFrame(rows)}
The comparison between arms rerun at every visit, always the same way: difference, Welch's t with its 95% interval, and Cohen's d.
The data it needs
  • Visits and measurements per patient, such as systolic blood pressure, exported from the trial's data capture system over SFTP
  • Randomisation list with each patient's arm, from an Excel file
  • Adverse events with type, severity and date, from a PostgreSQL database
Parts of Muvia used

Bring us a question you can't answer today.

We start from a real question your business has and walk the path from source to dashboard with your systems, not a demo dataset.