Data, AI and platformBI, AI and ML for the business
The calculation that costs: CPU by project and flow
Every night a flow recalculates margins and nobody knows what it costs. Muvia measures the CPU of every project, flow and step, shows which calculation weighs most, and lets you rewrite it and compare before and after on the same runs.
Illustrative scenario: it describes a typical case, not a customer project.
A slow calculation makes no noise. It just consumes.
A nightly flow that recalculates margins works: the numbers are there in the morning. Nobody notices that it walks through the rows one by one and uses a hundred times the CPU it needs, until there are ten such flows and the night is no longer long enough.
To decide what to optimise you need to know how much each project, flow and step consumes, and be able to show that the rewritten version gives the same results.
- Who it's for
- Data platform owners, data engineers and whoever writes the functions; IT and management deciding where to invest.
In Muvia, step by step
Real product screens, recorded on a project with sample data.
The video · Every night a flow recalculates margins and nobody knows what it costs. Muvia measures the CPU of every project, flow and step, shows which calculation weighs most, and lets you rewrite it and compare before and after on the same runs.
What you get
- You know where the CPU goes
- Project, flow and step: consumption reads from the top down, with no separate tool.
- You optimise what weighs
- The heaviest calculation stands out, and the effort goes where it matters instead of where it is convenient.
- Before and after, proven
- The runs of both versions sit on the same dashboard, with the same rows out.
- The decision stays written
- The notebook keeps why the calculation was costly and what was decided, next to the numbers.
For the technical team
How it is built in Muvia
- 1
Read the usage
In the company settings, the Usage page shows the month's CPU by project; from a project you drill down to the flow and from there to its individual steps.
- 2
Find the heaviest calculation
In the nightly margins flow, at 2:30 am, almost all the CPU goes into the Python function. The runs list the duration and version of each step.
- 3
Rewrite the function
The new version joins and groups instead of walking row by row; Athena can draft it. A trial run writes nothing, and the flow uses the new version from the following night.
- 4
Compare before and after
A dashboard reads the run log: CPU seconds per run with the first and the second version, and the same rows out.
- 5
Write down what you decide
A notebook explains why it cost so much, what changed and which other flows to look at, and stays in the project for whoever comes next.
Python function
import pandas as pd
def transform(inputs, params, ctx):
lines = inputs["lines"]
prices = inputs["price_list"][["item", "unit_cost"]].drop_duplicates("item")
df = lines.merge(prices, on="item", how="left")
df["revenue"] = df["price"] * (1 - df["discount_pct"] / 100) * df["quantity"]
df["cost"] = df["unit_cost"].fillna(0) * df["quantity"]
df["day"] = pd.to_datetime(df["day"])
out = (df.groupby(["day", "region", "category"], as_index=False)
.agg(lines=("revenue", "size"), revenue=("revenue", "sum"),
cost=("cost", "sum")))
out["margin"] = out["revenue"] - out["cost"]
revenue = out["revenue"].where(out["revenue"] != 0)
out["margin_pct"] = (out["margin"] / revenue * 100).round(2)
ctx.logger.info("margins: %d order lines, %d groups", len(df), len(out))
return {"rows": out}- The data it needs
- Order lines with price, discount and quantity from the ERP on SQL Server
- The price list with each item's cost, from an Excel file
- Customers and regions from Salesforce
- The CPU each run consumed, measured by Muvia
More cases in this line
All use cases- Data science on company dataYour data scientist trains wherever they like, in their own Jupyter with their own tools. Muvia gives them governed data to start from, runs the model on a schedule next to the data, and puts the scores into dashboards, alarms and reports.
- Athena: from a question to a query and a notebookOne question in plain language and the work is done: Athena saves the query in the Lab, writes the notebook with the ranking and the conclusion, and explains from the data what happened. Everything stays in the project, can be checked, and is one keystroke away.
- Data from business systems, checked before useFiles, databases and online services connect from one catalogue, and each becomes a table with no schema to design. Before the data reaches a report, written rules keep the bad rows out and set them aside in quarantine with the reason.
Bring us a question you can't answer today.
We start from a real question your business has and walk the path from source to dashboard with your systems, not a demo dataset.