---
name: compute-embeddings
description: >-
Compute cell embeddings for a single-cell dataset with a Helical foundation model, then
locate the outputs. Use when the user wants embeddings, representations, or features
from a dataset already in the platform catalogue. For training a model on labelled data,
use fine-tune-model instead.
---
# Compute embeddings
Turns a catalogue dataset plus a foundation model into an embedding matrix and a UMAP.
**The one thing to understand before starting:** `start_embedding_run` consumes real
compute and starts immediately. Nothing downstream will ask the user to confirm — there is
no approval screen for the user to click through. **You are the
confirmation step.** State the dataset, the model, and what the run will do, and get an
explicit yes before calling it.
You also cannot choose a project: scope comes from the signed-in user, and no tool accepts
a project or conversation identifier.
## 1. Select the dataset
`list_datasets({ name?, organism?, tissue?, disease?, limit? })` — `name` and `author`
match as case-insensitive substrings; `organism`, `tissue` and `disease` match an exact
element of the dataset's array field, so `organism: "human"` only matches if that exact
string is present. Note `total` is the count across all pages, not the rows returned.
`get_dataset({ id })` for the full record, including `cellCount` and `geneCount` — worth
reporting, because run time scales with them.
A user's own `.h5ad` can be added with `initiateDatasetUpload` → `completeDatasetUpload`
→ `registerDataset`; follow those tools' descriptions for splitting the file. Once
registered, it is embedded like any catalogue dataset.
## 2. Choose a model
`list_models({ modelType?, status? })` — promoted models by default; pass
`status: "all"` to see unpromoted candidates too. Pass the returned **`name`** as `model`.
A bare base-model name is accepted only if it is a known identifier (`scgpt`,
`gf-12L-40M-i2048`, `tf_sapiens`, `helix-mRNA`, and similar). For a fine-tuned model use
its registered `_v` name, or give `model` plus `model_version`. An unrecognised
bare name is rejected with a 400 rather than guessed at, and a promoted model belonging to
another workspace returns 403.
## 3. Price it, then confirm, then run
```
estimate_embedding_run({ datasetId, model, batch_size, modalities?, model_version? })
```
Free, starts nothing, safe to call repeatedly. Returns the token count, the price, the
assumptions behind it, and a `quote_id`. **Show the user the tokens and the price**, and
get an explicit yes. Then:
```
start_embedding_run({ ...the same arguments, quote_id })
```
The quote is what fixes the price, and the run tool will not accept a call without one.
Change any argument and the quote no longer applies — estimate again.
- **`batch_size` is required.** The API's own description says it can be omitted; it
cannot — omitting it is a 400. 8–32 is a reasonable starting range.
- **`modalities` must include `"sc"` for TranscriptFormer models** (`tf_sapiens`,
`tf_metazoa`, `tf_exemplar`), which fail at run time without it. For other single-cell
models it can be omitted.
Do not batch several runs behind one confirmation, and do not treat an earlier "sounds
good" about the plan as approval for the run itself.
Each call starts a separate run: a retry after a timeout may duplicate work, so check
`list_runs` before re-issuing.
## 4. Follow it through
`list_runs({ dagIds: ["embedding"], state: ["running"] })` to find the run, then
`get_run_details({ runId })` until it reaches a terminal state. Runs take minutes to hours;
poll at a sensible interval and keep the user informed rather than going silent.
## 5. Report the outputs
`get_run_details({ runId })` returns `artifacts[]` with `artifact_type`, `display_name`
and `s3_key`.
- `list_files({ path })` on the run's output directory to see what was produced.
- `list_s3_files({ path? })` lists an object-storage prefix **non-recursively**, across the
caller's projects. Use it when an `s3_key` from `artifacts[]` needs to be located or
confirmed to exist; use `list_files` for walking a run's output directory.
- `read_file({ path, maxBytes? })` reads **UTF-8 text only, up to 1 MiB**. An embedding
matrix is a binary `.npy` — it cannot be read through this tool. Report its path and let
the user fetch it from the Helical console; do not pretend to have inspected it.
- For the UMAP, `list_umaps({ datasetId })` then `get_umap({ runId })` returns parsed
coordinates and labels. The payload can be very large: summarise it, do not echo it.
## Hand-off
- Want the model trained on labels first? → `fine-tune-model`, then return here.
- Want to knock out, knock down or overexpress genes? → `in-silico-perturbation`: edit the
counts locally, upload, and embed the perturbed file here.
- Judging whether an embedding is good — separation of the conditions of interest in the
UMAP versus a zero-shot baseline — is interpretation, not something these tools report.
## Conventions
`datasetId`, never a path · scope is derived from the signed-in user, you cannot pass a
project · estimate → show the price → explicit yes → run with the quote · poll the run · read only under
`/projects//data` · never print credentials or raw upstream errors.