Data science on company data
Your data scientist trains wherever they like, in their own Jupyter with their own tools. Muvia gives them governed data to start from, runs the model on a schedule next to the data, and puts the scores into dashboards, alarms and reports.
Illustrative scenario: it describes a typical case, not a customer project.
A good model that stays in a notebook changes no decision.
The data scientist builds a model that estimates which invoices will be paid late. In their Jupyter it works. Then it needs fresh data every day, someone to run it, somewhere to write the results and a way to get them to the credit team: that is where the model stops.
Muvia draws a clear line. Training stays outside, in whatever environment the team prefers. Inside are the governed data to start from, the scheduled run of the model and the results in front of the people who decide.
- Who it's for
- Data scientists and analysts who build models; credit control, treasury and management accounting who use the results.
What you get
- Train where you like
- The team keeps its own environment and tools for training; Muvia does not impose a way of working.
- Features with one definition
- Training and inference read the same dataset with stable fields, fed by the sources and protected by roles and permissions.
- Run in Muvia, next to the data
- No server to look after for the model: inference runs on a schedule, in a sandboxed runtime, beside the tables it reads.
- Results in front of the business
- Scores land in the dashboards, alarms and reports the credit team already uses, not in a file to forward.
For the technical team
How it is built in Muvia
- 1
Explore governed data
In the SQL editor, query the tables your sources already feed, through read-only access. The column profile shows nulls, distinct values, min, max, quantiles and a histogram before you think about a model.
- 2
Prepare the features
One query computes, for each invoice, the customer's average days late, overdue amount against credit limit, invoices in the last year and open tickets. Save it as a dataset with stable fields: the same definition feeds training and inference.
- 3
Train in your own environment
From your own Jupyter, pull the dataset's rows through the read-only public API with an API key, as paged JSON. Train with scikit-learn or the tool you prefer and export the model to ONNX, or to joblib if it is scikit-learn. Muvia has no training interface and no model registry: training stays yours.
- 4
Attach the model to a function
Create a Python function and attach the model file to its version. The runtime is locked down and sandboxed, with no network: numpy, pandas, pyarrow, scipy, scikit-learn, onnxruntime and joblib. Athena can draft the code, and a trial run writes nothing.
- 5
Schedule inference
A flow runs the function every night on open invoices and writes the scores into a managed table. Dashboards, alarms and reports read it like any other data; a retrained model is a new version of the function.
Python function
import numpy as np
import onnxruntime as ort
FEATURES = ["avg_days_late", "overdue_to_credit", "invoices_12m",
"open_tickets", "amount"]
def transform(inputs, params, ctx):
session = ort.InferenceSession(ctx.file("late_payment.onnx"))
df = inputs["rows"]
x = df[FEATURES].to_numpy(dtype=np.float32)
_, proba = session.run(None, {session.get_inputs()[0].name: x})
out = df[["invoice_number", "customer_code", "due_date"]].copy()
out["late_risk"] = [p[1] for p in proba]
high = out["late_risk"] >= params["threshold"]
out["band"] = np.where(high, "high", "normal")
ctx.logger.info("scores computed: %d invoices", len(out))
return {"rows": out}- The data it needs
- Invoices, due dates and payments from the ERP on SQL Server
- Customers, industry and contacts from Salesforce
- Open support tickets, read from a REST API
- Credit limits and payment terms, from an Excel file
More cases in this line
All use cases- Margins and management accountingInvoices and costs sit in the ERP, customers and sales reps in the CRM, the budget in a spreadsheet. In Muvia they become one margin by customer and by product, with the Monday report and an alarm when the numbers drift from budget.
- Demand and stock forecastingSales history and stock levels are already in the ERP. Your data scientist builds a forecasting model with scikit-learn; Muvia runs it every night next to the data and puts forecast, actual sales and understock in front of the people who reorder.
- Sales and pipelineThe CRM knows what is being negotiated; the ERP knows what was ordered and invoiced. Muvia joins them: pipeline and conversion with one definition, the customers who are ordering less, and Athena answering the sales team's questions.
Bring us a question you can't answer today.
We start from a real question your business has and walk the path from source to dashboard with your systems, not a demo dataset.