@cyanheads/census-mcp-server
Query U.S. Census Bureau data, variables, and geography via MCP. STDIO or Streamable HTTP.
8 Tools
[](./CHANGELOG.md) [](./LICENSE) [](https://github.com/users/cyanheads/packages/container/package/census-mcp-server) [](https://modelcontextprotocol.io/) [](https://www.npmjs.com/package/@cyanheads/census-mcp-server) [](https://www.typescriptlang.org/) [](https://bun.sh/)
[](https://github.com/cyanheads/census-mcp-server/releases/latest/download/census-mcp-server.mcpb) [](https://cursor.com/en/install-mcp?name=census-mcp-server&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsIkBjeWFuaGVhZHMvY2Vuc3VzLW1jcC1zZXJ2ZXIiXX0=) [](https://vscode.dev/redirect?url=vscode:mcp/install?%7B%22name%22%3A%22census-mcp-server%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22%40cyanheads%2Fcensus-mcp-server%22%5D%7D)
[](https://www.npmjs.com/package/@cyanheads/mcp-ts-core)
**Public Hosted Server:** [https://census.caseyjhand.com/mcp](https://census.caseyjhand.com/mcp)
---
## Tools
8 tools covering the full Census data workflow — from dataset discovery and variable search through geography resolution and ranked comparisons:
| Tool | Description |
|:-----|:------------|
| `census_list_datasets` | Browse available Census Bureau datasets (ACS5, ACS1, Population Estimates, Decennial, County Business Patterns, Economic Census, Nonemployer Statistics) with vintage years and dataset codes. |
| `census_list_geographies` | List the geography levels supported by a dataset and year, with parent requirements and example FIPS values. |
| `census_search_variables` | Keyword search across variable labels and concept groups. On ACS, returns estimate and margin-of-error codes together. |
| `census_get_variable` | Fetch full metadata for one or more variable codes — label, concept, predicate type, universe, MOE sibling. |
| `census_list_predicate_values` | List the codes a filter dimension accepts (`EMPSZES`, `LFO`, `POPGROUP`, `NAICS2017`…), from the dataset dictionary or a live wildcard enumeration. |
| `census_resolve_geography` | Convert place names (e.g., "King County, WA") or street addresses to Census FIPS identifiers via TIGERweb and Census Geocoder. |
| `census_query_data` | Query a Census dataset for variables at a specific geography. Returns estimates with MOE, suppression codes resolved to readable reasons, and predicate filtering for the business datasets. |
| `census_compare_geographies` | Rank and compare variables across multiple geographies — all counties in a state, all states nationally, or a named set. Sorted table output, with the same predicate filtering. |
### `census_list_datasets`
Browse available Census Bureau datasets.
- Returns dataset codes, names, descriptions, and available vintage years
- Covers ACS5, ACS5 Data Profiles, ACS5 Subject Tables, ACS1, ACS1 Data Profiles, Population Estimates, Decennial Redistricting (P.L. 94-171), Decennial DHC, County Business Patterns (`cbp`), Economic Census (`ecnbasic`), and Nonemployer Statistics (`nonemp`)
- Each description names the filter predicates the dataset requires and the geography levels it publishes — both vary by dataset
- Accepts an optional keyword filter
- Dataset codes (e.g., `acs/acs5`) are the values to pass to other tools
- `available_years` is exhaustive, not a sample: any other year fails with `year_not_available` before a request goes out, naming the years that do work. It is narrower than what the Census API hosts — `pep/charv` reaches its 2020-2022 estimates through the `YEAR` filter inside the 2023 vintage, and the `cbp`/`nonemp` vintages left out reject the `NAME` column every query here sends
---
### `census_search_variables`
Search Census variables by keyword.
- Full-text search across label and concept fields with relevance scoring (exact concept match > label match > partial)
- On ACS datasets, returns estimate (E suffix) and margin-of-error (M suffix) codes together so both can be requested in one query — no other family publishes margins of error, and an E-final code there is an ordinary code
- Also surfaces the predicate codes a dataset filters on, such as `NAICS2017` in `cbp`
- Configurable limit (default 20, max 100); `total_matches` indicates how many matched before the limit
- Cache-backed: variables.json is fetched once per dataset+year with a configurable TTL (default 24h)
---
### `census_list_predicate_values`
List the codes a filter dimension accepts, so a `predicates` map can be written without guessing.
- Two routes, picked by where the answer lives: a dimension with a published value list is read from the dataset dictionary, one without is enumerated live by wildcarding it on the data endpoint. `NAICS*` and `POPGROUP` always publish one (thousands of codes — narrow them with `query`); on the current vintages `EMPSZES`, `LFO`, `RCPSZES`, `TAXSTAT`, and `TYPOP` publish none, so the live route is the only place their codes appear
- A dictionary value list is a classification shared across Census products, not a record of what one dataset serves — `dec/ddhca` declares 5,543 `POPGROUP` codes and publishes 2,996, `cbp` declares 6,694 `NAICS2017` codes and publishes 2,003. The declared list is checked against the dataset's own published rows and the dead codes are dropped; `source` says whether that check ran and the notice says how many were withheld. A keyword that matched only withheld codes names them, so "total population" on `dec/ddhca` reports that `001` is declared and serves nothing rather than reading like a typo
- Keyword `query` matches code and label; results are sorted by code and a truncated list is disclosed rather than passed off as complete
- `ecnbasic` publishes `TAXSTAT` and `TYPOP` per industry, so `within_naics` scopes the enumeration — and the notice says the result is complete for that industry alone. A per-industry dimension is left unchecked for the same reason, since an unscoped check would withhold codes a scoped query does return
- Live enumerations are cached per dataset, year, dimension, industry scope, and probe measure
---
### `census_resolve_geography`
Convert place names and addresses to Census FIPS identifiers.
- Named places (e.g., "King County, WA", "Seattle, WA", "California") resolved via TIGERweb MapServer
- Street addresses resolved to tract level via Census Geocoder
- Auto-detects the geography level — state for an abbreviation or spelled-out state name, county for "County"/"Borough"/"Parish", tract for "Tract", otherwise place falling back to county; `geography_type` overrides it
- Also resolves metropolitan/micropolitan statistical areas, combined statistical areas, and consolidated cities — never auto-detected, since their names overlap city names, so each needs an explicit `geography_type`. The value is the level's own Census API name, so it feeds `geography_level` unchanged
- Optional `county_fips` pins a tract name to one county, since a tract name is unique only inside its county. Only county and tract sit within a county, so it restricts resolution to those two levels rather than being dropped on a layer that cannot apply it
- Prefers an exactly-named match, so "Kansas City, MO" does not resolve to North Kansas City
- Never picks between matches: anything still matching more than one geography comes back as `ambiguous_name`, with every candidate carrying the code resolving it would have returned, plus the state that separates same-named places
- Returns `state_fips` (→ `parent_fips`) and `fips_summary` (→ `geography_fips`) ready to pass to other tools; a statistical area omits `state_fips`, since it can span several states and takes no parent
---
### `census_query_data`
Query a Census dataset for one or more variables at a specific geography.
- Requires FIPS codes — use `census_resolve_geography` first for place names
- Use `geography_fips: "*"` to return all geographies at the level within the parent
- The level and its parents are checked against the dataset's own geography metadata before the query runs: a missing `parent_fips` returns `parent_required` naming what to add, and a parent the level does not sit within returns `parent_not_accepted` naming the input to drop — neither reaches the API as an opaque 400
- `parent_fips` and `county_fips` are zero-padded to the widths the Census matches on, so `"5"` and `"05"` both find Arkansas; either also takes `"*"`, which is what reaches every block group in a state. `geography_fips` takes its width from `geography_level` and is passed through as given
- Each row carries both `geography_fips` (bare level code, round-trips back into this tool) and `geography_geoid` (level plus parents, nationally unique)
- A query that matches nothing returns `no_data` with dataset-aware recovery, not a retried upstream error
- Optional `predicates` map for the datasets that filter on one — `{"NAICS2017": "5112"}` narrows a `cbp` count to software publishers, and `census_list_predicate_values` supplies the codes. Keys are validated against the dataset's own variables before the query
- Dimensions left unset are named in a notice and their applied default is echoed per row in `applied_filters`. That label is load-bearing: `cbp` defaults `NAICS2017` to the all-industries total, but `dec/ddhca` defaults `POPGROUP` to one population group and `ecnbasic` defaults its NAICS dimension to a single sector, so an unfiltered value can read like a total without being one. A dimension that publishes no label attribute (`pep/charv` `YEAR`, the `nonemp` NAICS codes before 2012) has no default to echo, and the notice says so rather than leaving it looking undefaulted
- One geography can come back on more than one row: `pep/charv` publishes an April 1 estimates base alongside its July 1 estimate, and `MONTH` is what separates them — not `YEAR`, which both rows carry. Each row names its record in a `record` field and on its rendered heading, and the notice gives the predicate that pins one (`{"MONTH": "7"}`)
- Suppression codes (geography too small, data not collected, etc.) resolved to human-readable reasons
- A cell that holds text rather than a number keeps it, under `value`, so a null `estimate` says which of three things it is: `suppressed` is a number the Census withheld, a `value` alongside it is text (`GEO_ID` returns `"0500000US53033"`), and neither is an empty cell
- Variable labels enriched from cache and surfaced alongside estimates
- Requires `CENSUS_API_KEY`
---
### `census_compare_geographies`
Rank and compare variables across multiple geographies.
- Fetches all geographies at a level (e.g., all WA counties) in one API call, then sorts and slices
- Optional `within` parameter to constrain to a parent FIPS; omit for national comparison
- Optional `geographies` list to filter to specific geographies — full GEOIDs (`"53033"`, `"06037"`) work across states; bare level codes (`"033"`) need `within` to disambiguate. Entries matching no row, and bare codes that matched more than one state, are named in a notice
- Same pre-query level and parent validation as `census_query_data`, reported against `within` / `within_county`
- Configurable sort variable, direction, and limit (default 50, max 500)
- Same `predicates` map as `census_query_data`, applied to every geography — without it the ranking runs on whatever default the API picks, named in the notice and echoed per row in `applied_filters`
- A dataset that publishes several records per geography is refused rather than ranked twice: a rank is a statement about one geography, so `pep/charv` without a pinned record fails with `ambiguous_rows` naming `MONTH` and the code to pass. With one pinned, each geography ranks once and the row says which record it is
- Suppressed values sorted to end of results and labeled rather than passed through as negative sentinels
- Same `value` field as `census_query_data` for a text cell; text has no ordering, so sorting on a column of it leaves every row tied
- Requires `CENSUS_API_KEY`
---
## Features
Built on [`@cyanheads/mcp-ts-core`](https://www.npmjs.com/package/@cyanheads/mcp-ts-core):
- Declarative tool definitions — single file per tool, framework handles registration and validation
- Unified error handling — handlers throw, framework catches, classifies, and formats with recovery hints
- Structured logging with optional OpenTelemetry tracing
- STDIO and Streamable HTTP transports
Census-specific:
- In-process variable cache with configurable TTL — variables.json fetched once per dataset+year, searched client-side
- Three-API backend: Census Data API for data queries, TIGERweb for named-place resolution, Census Geocoder for address-to-tract
- Automatic retry with backoff on all external API calls
- FIPS formatting helpers — zero-padded state, county, and tract codes ready to pass between tools
Agent-friendly output:
- Workflow-oriented tool surface — `fips_summary` and `state_fips` return values are ready to pass as `geography_fips` and `parent_fips` to the next tool
- Suppression codes decoded — Census negative sentinel values (e.g., `-666666666`) surfaced as human-readable reasons instead of raw numbers
- Recovery hints on errors — ambiguous geography names include candidate lists; missing API key errors include registration URL
---
## Getting started
> **API key:** Register a free key at [api.census.gov/data/key_signup.html](https://api.census.gov/data/key_signup.html). Variable search and geography resolution work without a key; data queries (`census_query_data`, `census_compare_geographies`) require one.
Add the following to your MCP client configuration file:
```json
{
"mcpServers": {
"census-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/census-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info",
"CENSUS_API_KEY": "your-census-api-key"
}
}
}
}
```
Or with npx (no Bun required):
```json
{
"mcpServers": {
"census-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/census-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info",
"CENSUS_API_KEY": "your-census-api-key"
}
}
}
}
```
Or with Docker:
```json
{
"mcpServers": {
"census-mcp-server": {
"type": "stdio",
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "MCP_TRANSPORT_TYPE=stdio",
"-e", "CENSUS_API_KEY=your-census-api-key",
"ghcr.io/cyanheads/census-mcp-server:latest"
]
}
}
}
```
For Streamable HTTP, set the transport and start the server:
```sh
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 CENSUS_API_KEY=... bun run start:http
# Server listens at http://localhost:3010/mcp
```
### Prerequisites
- [Bun v1.3.0](https://bun.sh/) or higher (or Node.js v24+).
- A Census API key — register free at [api.census.gov/data/key_signup.html](https://api.census.gov/data/key_signup.html). Required for `census_query_data` and `census_compare_geographies`; other tools work without it.
### Installation
1. **Clone the repository:**
```sh
git clone https://github.com/cyanheads/census-mcp-server.git
```
2. **Navigate into the directory:**
```sh
cd census-mcp-server
```
3. **Install dependencies:**
```sh
bun install
```
4. **Configure environment:**
```sh
cp .env.example .env
# edit .env and set CENSUS_API_KEY
```
---
## Configuration
| Variable | Description | Default |
|:---------|:------------|:--------|
| `CENSUS_API_KEY` | **Required for data queries.** Register free at api.census.gov/data/key_signup.html. | — |
| `CENSUS_DEFAULT_YEAR` | Default vintage year when no year is specified. | `2024` |
| `CENSUS_VARIABLE_CACHE_TTL_HOURS` | Hours to cache variables.json per dataset+year in memory. | `24` |
| `MCP_TRANSPORT_TYPE` | Transport: `stdio` or `http`. | `stdio` |
| `MCP_SESSION_MODE` | HTTP session mode: `stateful`, `stateless`, or `auto`. `auto` resolves to `stateful`; the Docker image sets `stateless`. | `auto` |
| `MCP_HTTP_PORT` | Port for HTTP server. | `3010` |
| `MCP_AUTH_MODE` | Auth mode: `none`, `jwt`, or `oauth`. | `none` |
| `MCP_LOG_LEVEL` | Log level (`debug`, `info`, `notice`, `warning`, `error`). | `info` |
| `OTEL_ENABLED` | Enable OpenTelemetry instrumentation. | `false` |
See [`.env.example`](./.env.example) for the full list of optional overrides.
---
## Running the server
### Local development
```sh
# One-time build
bun run rebuild
# Run the built server
bun run start:stdio
# or
bun run start:http
```
Run checks and tests:
```sh
bun run devcheck # Lint, format, typecheck, security audit
bun run test # Vitest test suite
bun run lint:mcp # Validate MCP definitions against spec
```
### Docker
```sh
docker build -t census-mcp-server .
docker run --rm -e CENSUS_API_KEY=your-key -p 3010:3010 census-mcp-server
```
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to `/var/log/census-mcp-server`. OpenTelemetry peer dependencies are installed by default — build with `--build-arg OTEL_ENABLED=false` to omit them.
---
## Project structure
| Path | Purpose |
|:-----|:--------|
| `src/index.ts` | `createApp()` entry point — registers tools and initializes services. |
| `src/config/server-config.ts` | Census-specific env var parsing and validation with Zod. |
| `src/mcp-server/tools/definitions/` | Tool definitions (`*.tool.ts`). |
| `src/services/census-api/` | Census Data API client — data queries, suppression code mapping, retry logic. |
| `src/services/geography/` | Geography resolution — TIGERweb named-place lookup and Census Geocoder address-to-tract. |
| `src/services/variable-cache/` | In-process variables.json cache with TTL and keyword search. |
| `tests/` | Vitest tests mirroring `src/` structure. |
---
## Development guide
See [`CLAUDE.md`](./CLAUDE.md) for development guidelines and architectural rules. The short version:
- Handlers throw, framework catches — no `try/catch` in tool logic
- Use `ctx.log` for request-scoped logging, `ctx.state` for tenant-scoped storage
- Register new tools via the barrel in `src/mcp-server/tools/definitions/index.ts`
- Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields
---
## Contributing
Issues and pull requests are welcome. Run checks and tests before submitting:
```sh
bun run devcheck
bun run test
```
---
## License
Apache-2.0 — see [LICENSE](LICENSE) for details.