Train wherever you like. Run in Muvia, next to the data.
For a data scientist, Muvia is where a model stops living in a notebook on a laptop: data already connected and permissioned, Python with pandas, scipy and scikit-learn, models that run on a schedule in the cloud or on the edge node, and results that land in the dashboards of the people who decide.
The model was trained outside Muvia and uploaded with the function. The flow runs it every Monday and writes the forecasts to a table the dashboard reads.
From raw data to a model in production.
A data scientist's working cycle, and where each step happens.
In Muvia
In your environment
1In Muvia
Explore
Query the sources in SQL. Every column's profile shows nulls, distinct values, minimum, maximum, quantiles and a histogram.
2In Muvia
Build features
Moving averages, standard deviation, EWMA, lag and lead: windows are computed in queries and saved to tables that keep themselves current.
3In your environment
Train
In your own Jupyter, with the tools you already use, pulling data from the read-only API with a key. For simple models you can also fit a scikit-learn estimator inside the function on every run.
4In Muvia
Put it in production
Upload the model, as ONNX or joblib, with a Python function. A flow runs it on a schedule, runs a Replay of history when needed and writes results to a managed table.
5In Muvia
Get results to the people who decide
Forecasts and scores become dashboards, alerts and PDF reports, with the same permissions as the rest of the data. Athena answers questions about them too.
One fixed Python environment, the same in the cloud and on the floor.
Every function receives its data in inputs["rows"] as a pandas DataFrame and returns {"rows": DataFrame}. It runs in an isolated environment with no network and no installs: what works today works the same next month, on the server and on the edge node, amd64 or arm64.
If your model uses other libraries, you can often export it to ONNX and run it with onnxruntime.
Available libraries
pandas
tables and transformations
numpy
numerical computing
pyarrow
columnar data
scipy
statistics and signals
scikit-learn
models and preprocessing
onnxruntime
models exported to ONNX
joblib
saved scikit-learn models
Everyday statistics come ready-made.
Many questions need no code: the most common analyses are query templates and catalogue functions you fill in.
Deviation from baseline (z-score)
Moving average and rate of change
Comparison with the previous period
Threshold crossings, time out of limits, downtime
Consumption over a period, from a counter
Shift summary and rankings
Comparison of two variables
Weighted health index from 0 to 100
The same model, on the edge node.
With Muvia Industrial, the function you wrote for the cloud also runs on the node on the shop floor, next to the machine: the same Python environment, the same libraries, the result computed before the data leaves.
Athena drafts query steps and function code from a request, and a function can ask a language model for an answer within a budget. You review it and decide what goes to production.