Data Science Skills by Probabl
Tell us your problem. Let's experiment together.
Bring your agent and your model. Our skills and libraries organize, build, and evaluate the project, with the expertise from more than a decade of building scikit-learn.

How to get started
Your agent and your model.
Our skills.
Install our CLI
Use the package manager you already have.
❯
pip install skore-cli❯
conda install -c conda-forge skore-cli❯
uv tool install skore-cli❯
pixi global install skore-cliInstall our skills
Add our skills to your coding agent.
❯
skore skills installStart your agent
Open the coding agent you already use.
❯
pi❯
opencode❯
claude❯
codex
How it works
From your data science problem.
To deeper experiments.
01 Do it right from the start
Setup up your project
We do the heavy lifting for you: scaffolding a proven data science workspace, managing a reproducible Python environment with your favorite package manager, recording every dependency, and keeping the project under version control.
A structured workspace ready for new experiments and production
Dependencies and experiment history preserved with Git

02 Get the right insights
Explore your data
We dig into every data table: inspecting, analyzing, and exploring it to surface insights and catch issues early, researching ideas from the literature, and turning every finding into code for a detailed report you can return to.
Insights and data issues caught before any modeling
A detailed report with code you can come back to

03 Get evidence you can trust
Build and evaluate models
We take it from data source to prediction: connecting the tables with skrub, modeling with a scikit-learn-compatible library, and evaluating in skore with a structured report of the right metrics, issue checks, and insights on the data, pipeline, and results.
Any data source, one scikit-learn-compatible pipeline
A skore report of metrics, checks, and insights

04 Keep learning
Iterate on new ideas
We turn the reports into the next experiment: collecting what the evaluation already showed, deciding what to do next, and searching the literature for ideas that build on what you know and what is still worth exploring.
Next steps drawn from the reports and the evaluation
Literature research on what you know and could explore

Benchmark
A working model, sooner.
Same model, same harness, same data. The skills change how long it takes to reach a working predictive model.
Time to a result
DeepSeek V4 Flash in OpenCode. Six open-data tasks. One point is one run.
A working model is a valid result that beats a trivial baseline. The hand-written baseline is the stronger bar. One drinking-water run with the skills never reached it, so that row has 17 runs.
Mean rounds
1.28 / 1.67
Rounds to a working model, with the skills first.
Round timeouts
6 / 28
Rounds that hit the time limit, with the skills first.
Final quality
Level
17 of 18 runs with the skills match a hand-written baseline, and 18 of 18 do without. The run that falls short is drinking-water.
Scale with Skore
Start locally.
Scale when you need it.
Your team needs more velocity
In the enterprise, velocity matters. Our skills bring methodological backing and statistical thinking. Skore is the platform that scales data science experiments: a remote agent on your LLM provider, remote compute on your infrastructure, remote storage so no result is lost, and a Skore UI where data scientists investigate the findings.
Discover Skore- 01
Remote agent
Connect your agent to the LLM provider you already use.
02Remote compute
Scale experiments onto your infrastructure.