generated: '2026-09-16' method: searched source: https://github.com/probabl-ai/skills note: 'Provider-published Agent Skills (probabl-ai/skills, catalog .catalog.json), saved verbatim. These are data-science workflow skills for coding agents (skrub/scikit-learn/skore), not REST call recipes, so they do not list Skore Hub operationIds. Installed with `skore skills install` (skore-cli). Re-verified 2026-09-16: all 14 skills//SKILL.md paths present upstream.' catalog: https://github.com/probabl-ai/skills/blob/main/.catalog.json install: pip install skore-cli && skore skills install skills: - file: probabl-audit-ml-pipeline.md name: audit-ml-pipeline description: 'Owns the `audit/` folder: one `# %%` (jupytext percent) Python file per experiment, aligned 1:1 with `experiments/NN_.py` and `journal/NN_.md`, that loads the experiment''s skore report **read-only** and uses bare-last-expression cells whose `__repr__` carries the audit''s signal. The agent executes the audit file via the bundled in-process runner (`audit-ml-pipeline/scripts/run_cells.py` — IPython `InteractiveShell.run_cell`), which streams a markdown digest of each cell''s stdout + last-expression repr to stdout (optionally also to a file). The digest fuels narrative work (the `JOURNAL.md` Status + History update, follow-up questions about a past experiment, cross-experiment comparison). Stops at "audit/NN_*.py is placed, executed, and the digest is available." Never calls `skore.evaluate(...)` or `project.put(...)`.' upstream: https://github.com/probabl-ai/skills/blob/main/skills/audit-ml-pipeline/SKILL.md - file: probabl-build-ml-pipeline.md name: build-ml-pipeline description: Declare the pipeline from data source to predictor as a **skrub DataOps graph** (not as a bare `sklearn.Pipeline`). Every step is either a pure-Python function (stateless) attached via `.skb.apply_func`, or a sklearn-compatible estimator (stateful) attached via `.skb.apply`. Stops at the declared object — no fit, split, tuning, persistence, or evaluation. upstream: https://github.com/probabl-ai/skills/blob/main/skills/build-ml-pipeline/SKILL.md - file: probabl-data-science-python-stack.md name: data-science-python-stack description: Opinionated Python stack for data-science / ML work — one library per job, organized into tiers (mandatory / user choice / optional / transitive). SKILL.md is the index; per-library `references/.md` files carry scope, "pick this when" / "pick something else when", and pairings. upstream: https://github.com/probabl-ai/skills/blob/main/skills/data-science-python-stack/SKILL.md - file: probabl-evaluate-ml-pipeline.md name: evaluate-ml-pipeline description: 'Methodology for evaluating a single sklearn-compatible learner (in particular, the `SkrubLearner` produced by `build-ml-pipeline`). Owns: which entry point to call (`skore.evaluate` first, the explicit report classes when needed), which cross-validator to pick from scikit-learn''s catalogue, how to consume the structural metadata (`groups`, `times`, …) attached at build time via `.skb.mark_as_X(split_kwargs=...)`. Stops at "what does the report say". Defaults (metrics, plots) come from skore; only override on explicit user request.' upstream: https://github.com/probabl-ai/skills/blob/main/skills/evaluate-ml-pipeline/SKILL.md - file: probabl-explore-ml-data.md name: explore-ml-data description: Owns data understanding BEFORE any model is designed. Places and executes `data/eda.py` (a jupytext `# %%` script) via the shared in-process runner, reads the streamed digest, then writes a persisted `data/eda.md` report (plus linked `data/eda_.html` skrub `TableReport` pages) and the `## Data understanding (EDA)` section of `journal/JOURNAL.md`. The point is to surface the dataset facts — shape, dtypes, missingness, cardinality, target balance / skew, datetime / group structure, feature associations — that JUSTIFY the later learner / splitter / metric decisions, so the user understands *why* the modelling choices are made. Uses `skrub.TableReport` for dataframe overviews and the shared runner `audit-ml-pipeline/scripts/run_cells.py`. Stops at "EDA executed, `data/eda.md` + HTML written, JOURNAL EDA section updated." Never designs the model, never edits `src//`, never modifies the user's raw data files. upstream: https://github.com/probabl-ai/skills/blob/main/skills/explore-ml-data/SKILL.md - file: probabl-iterate-from-skore.md name: iterate-from-skore description: Source the next ML experiment proposal by **reading the audit digest** at `scratch/audit//audit.md` (produced by `audit-ml-pipeline` at § 4 record-outcome). For every row in the digest's `## Checks summary` whose `severity` is `issue` or `tip`, follow the row's `documentation_url` to draft a Backlog row whose `Item` is the mitigation the docs recommend. The `## Metrics summary` provides context for the human summary paragraph but does not drive Backlog rows on its own. Returns the enriched Backlog rows + a one-paragraph summary back to `iterate-ml-experiment`, which writes the rows into `JOURNAL.md` and re-presents the sourcing menu so the user can promote a `B` row. Stops at "Backlog enriched, summary returned"; never writes a per-experiment design note, never picks the "winning" finding — the user picks via `B`. upstream: https://github.com/probabl-ai/skills/blob/main/skills/iterate-from-skore/SKILL.md - file: probabl-iterate-from-user.md name: iterate-from-user description: 'Source the next ML experiment proposal from the user via one of three entry points selected by `AskUserQuestion`: (a) a scientific article URL the agent must read and synthesize, (b) a resource link or path (GitHub issue / spec file / reference repo), or (c) free-text the user types directly. In every branch, the agent reads the source, synthesizes its understanding of what to implement, and confirms with the user *before* returning the Proposal block. Hand the confirmed Proposal back to `iterate-ml-experiment`, which writes it into `journal/NN_short_name.md` and seeks the user''s design-note approval. Stops at "Proposal returned, user-confirmed"; never writes a design note, never authors acceptance criteria.' upstream: https://github.com/probabl-ai/skills/blob/main/skills/iterate-from-user/SKILL.md - file: probabl-iterate-ml-experiment.md name: iterate-ml-experiment description: 'Owns the iteration loop on top of an ML workspace: the `journal/JOURNAL.md` index and the per-experiment `journal/NN_short_name.md` design notes that must be drafted and approved by the user **before** `experiments/NN_short_name.py` is created. Drives the propose → iterate → approve → implement → record loop; dispatches to `iterate-from-skore` / `iterate-from-user` for sourcing.' upstream: https://github.com/probabl-ai/skills/blob/main/skills/iterate-ml-experiment/SKILL.md - file: probabl-organize-ml-workspace.md name: organize-ml-workspace description: 'Decide where files live in an ML experimentation project: reusable code in `src//`, one `# %%` script per experiment in `experiments/`, design notes + index in `journal/`, reports in `reports/`, agent-only probes in `scratch/`. Owns the layout, the file-creation rules (one file per experiment, ask before editing), and the jupytext `# %%` script convention. Never imposes `data/` — the user owns that.' upstream: https://github.com/probabl-ai/skills/blob/main/skills/organize-ml-workspace/SKILL.md - file: probabl-python-api.md name: python-api description: 'Look up the public API of a Python package against the *installed version* and cache what''s worth keeping. Four shapes by question type: (0) cache hit under `scratch/api///`; (1) `inspect.signature` + `pydoc.render_doc` for a symbol; (2) `dir` / `pkgutil.iter_modules` for a module surface; (3) WebSearch + WebFetch of versioned docs for narrative ("how", "which", "what does X return when Y"). Never write a symbol from training-data memory — recognition is not a lookup.' upstream: https://github.com/probabl-ai/skills/blob/main/skills/python-api/SKILL.md - file: probabl-python-code-style.md name: python-code-style description: 'Owns Python code style for this stack: ruff for lint + format, numpydoc for docstrings. Three responsibilities — (1) place the project''s `ruff.toml` from the bundled template once the stack and workspace are in place, (2) run ruff against any Python files Claude has just generated or edited, and (3) contextualize each touched file''s comments to the data-science problem — rewriting any leftover template / workflow prose (skill names, gates, runner, digest, guard-rails) into concise, problem-specific docs so the user''s committed files read like a colleague wrote them, not like a generated scaffold. Stops at "the touched files pass `ruff check` and document the problem, not the process."' upstream: https://github.com/probabl-ai/skills/blob/main/skills/python-code-style/SKILL.md - file: probabl-python-env-manager.md name: python-env-manager description: 'Single source of truth for "which Python environment manager does this project use, and how do I install a package with it?". Owns the detection table (pixi / uv / poetry / hatch / conda+mamba / pip+venv), the install / remove / upgrade commands per manager, and the bootstrap path when no manager is in place (default recommendation: pixi). Stops at "the install command was issued with the right manager and the package is importable".' upstream: https://github.com/probabl-ai/skills/blob/main/skills/python-env-manager/SKILL.md - file: probabl-smoke-test-ml-pipeline.md name: smoke-test-ml-pipeline description: 'Owns the smoke test contract for an ML experiment: a small, diagnostic-by-construction pytest that fits the experiment''s learner on a portion of the real `data/` source and predicts on a *disjoint* portion that deliberately carries **no pre-history buffer**. The assertion is structural — the number of predictions must equal the number of rows in the predict grid. A pipeline that loads-then-features-then-splits will silently drop the cold-start rows of the predict slice and the test will fail with a row-count mismatch; a pipeline that marks X early and references upstream history nodes from feature steps will pass trivially. The smoke test is the executable proof of the X-marker placement rule from `build-ml-pipeline`.' upstream: https://github.com/probabl-ai/skills/blob/main/skills/smoke-test-ml-pipeline/SKILL.md - file: probabl-test-ml-pipeline.md name: test-ml-pipeline description: 'Owns the `tests/` folder of an ML workspace and the pairing rule between an experiment and its tests. Lightweight router: every test category has its own subskill (`smoke-test-ml-pipeline` is the only one for v1; `regression-test-ml-pipeline` / `distribution-test-ml-pipeline` etc. plug in as siblings as the workspace grows). This skill places an empty `tests//` folder, enforces the stem-pairing rule between `tests//test_NN_.py` and `experiments/NN_.py`, and dispatches to the matching subskill when the user asks for a test.' upstream: https://github.com/probabl-ai/skills/blob/main/skills/test-ml-pipeline/SKILL.md