SelfBench
Find the best models for your repo.
Private coding-agent benchmarks built from your repository's own merged pull requests.
Public benchmarks tell you how a model does on someone else's code. SelfBench tells you how it does on yours: it turns your merged PRs into tasks with hidden tests, runs agents and models on them, and plots accuracy against cost.
## Results
Every open-source repository released on **[selfbench.dev](https://selfbench.dev)** gets a live leaderboard: each model, harness, and reasoning setting placed by accuracy and cost per task, with the Pareto frontier drawn through the settings nothing else beats on both.
**Browse the leaderboards:** [vercel/next.js](https://selfbench.dev/vercel/next.js) · [supabase/supabase](https://selfbench.dev/supabase/supabase) · [earendil-works/pi](https://selfbench.dev/earendil-works/pi) · [getsentry/sentry](https://selfbench.dev/getsentry/sentry) · [PostHog/posthog](https://selfbench.dev/PostHog/posthog) · [pingdotgg/t3code](https://selfbench.dev/pingdotgg/t3code) · [vercel/vercel](https://selfbench.dev/vercel/vercel) · **[all repositories →](https://selfbench.dev)**
## How it works
For each merged PR, SelfBench rebuilds the task from the commit before the change: the PR's own request becomes the instruction, and an authoring agent writes hidden tests and a reference solution. A task is accepted only if the tests fail without a solution, pass with the original implementation, pass again on a rerun, and survive an independent review. Every accepted task is a native [Harbor](https://harborframework.com/) task.
## Using SelfBench
Everything happens in the web app at [app.selfbench.dev](https://app.selfbench.dev):
1. **Sign in** with GitHub and **connect a repository**.
2. **Batch Generation**: choose how many easy, medium, and hard tasks to build, or use **Add PRs** on the Dataset page to build one task from each pull request you pick. A batch takes hours; it keeps running after you close the page.
3. **Dataset**: inspect each task (instruction, environment, hidden tests, reference patch, pipeline artifacts) and approve or reject it.
4. **Run**: pick models, harnesses, and a sandbox, and run them on the approved tasks.
5. **Results**: compare accuracy against cost, and open any trial's transcript and scores.
6. **Releases**: publish a public repository's results to [selfbench.dev](https://selfbench.dev).
Models and sandboxes run on your organization's own keys under **Credentials**. Read released results through the [Public Results API](docs/api-reference/overview.mdx), or automate your workspace with an [API key](docs/api-reference/workspace.mdx).
## Self-hosting
SelfBench's reference deployment runs on GCP: Cloud Run serves the API, while GKE Autopilot runs the Temporal workflow worker and KEDA-scaled Harbor jobs, with Cloud SQL and GCS. A Cloud Run worker pool remains available as an alternative.
1. Provision a GCP project, billing, Terraform state bucket, and GitHub Actions Workload Identity Federation.
2. Configure Terraform inputs and store each runtime secret value in its own Secret Manager secret.
3. Apply the environment with Terraform, or configure the protected GitHub `dev` and `prod` environments to deploy through Actions.
4. Point your domain at the provisioned load balancer and configure GitHub OAuth for the app URL.
For prerequisites, exact Terraform commands, runtime configuration, GitHub Actions setup, and GKE worker setup, see the [self-hosting and infrastructure guide](infra/README.md).
## Development
Requires [Bun](https://bun.sh/) 1.3.14+ and Docker with Compose.
```bash
bun install --frozen-lockfile
bun run validate
```
Run the whole stack (API serving the app, worker, Temporal, Postgres, local Docker sandboxes) from a checkout:
```bash
cp .env.example .env # GitHub OAuth app, session secret, credential key, managed keys
SELFBENCH_PUBLIC_URL=https://your-tunnel.example docker compose --profile sandbox up -d --build
docker compose port api 8080
```
Compose names the project after the checkout directory and publishes ephemeral host ports, so worktrees run side by side. Set `SELFBENCH_PUBLIC_URL` to the origin the browser opens (a tunnel or reverse proxy) before `up`, and register `/auth/github/callback` on the OAuth app.
For frontend work, `bun run dev:site` runs the API and Vite with hot reload (secrets in `.env.site`; see `scripts/dev-site.ts`).
## Documentation
- [Mintlify docs](docs/): guides, public results API, and workspace API (`cd docs && npm ci && npm run dev`)
- [How it works](docs/concepts/task-generation.mdx): the pipeline and what makes a task valid
- [Public Results API](docs/api-reference/overview.mdx) · [Workspace API](docs/api-reference/workspace.mdx)
- [Infrastructure](infra/README.md)
## License
[MIT](LICENSE) © 2026 Mupt AI.