# PDF SPEC MCP Server [![CI](https://github.com/shuji-bonji/pdf-spec-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/shuji-bonji/pdf-spec-mcp/actions/workflows/ci.yml) [![npm version](https://img.shields.io/npm/v/@shuji-bonji/pdf-spec-mcp)](https://www.npmjs.com/package/@shuji-bonji/pdf-spec-mcp) [日本語版 README はこちら](README.ja.md) An MCP (Model Context Protocol) server that provides structured access to ISO 32000 (PDF) specification documents. Enables LLMs to navigate, search, and analyze PDF specifications through well-defined tools. > [!IMPORTANT] > **This is a specification *reference*, not a rule engine.** > It retrieves and structures the text of ISO 32000 — clauses, tables, definitions, and > `shall`/`should`/`may` requirements. It does **not** examine a PDF file, and it cannot tell > you whether a document conforms to anything. Conformance verdicts come from > [pdf-verify-mcp](https://github.com/shuji-bonji/pdf-verify-mcp) > (`validate_conformance` / `evaluate_policy`). > > The distinction matters because three different things get conflated: > **declaration** — a label the file wrote about itself ("I am PDF/A" in the metadata). Writing it is not evidence / > **conformance** — whether the file actually meets the standard. There is no way to prove it in full; you can only find where it breaks the rules / > **validation** — what a validator (veraPDF and the like) reports against the checks it implements. A pass means "this inspection did not fail", not "the file conforms to the standard". > Reading a `shall` here tells you what the standard requires — not whether your file meets it. > > **A search that returns nothing means "cannot answer", not "no such requirement."** > ISO 19005 (PDF/A) and ETSI PAdES are outside this corpus; see `list_specs` → `coverage.gaps`. ### What each PDF family server does — and does not do | Server | Does | **Does not** | |---|---|---| | **pdf-spec-mcp** (this) | Search, retrieve and extract requirements from 17 PDF-related documents | **Is not a rule engine.** Does not define business rules, inspect PDF files, or validate schemas. ISO 19005 (PDF/A) is not part of the corpus | | [pdf-reader-mcp](https://github.com/shuji-bonji/pdf-reader-mcp) | Extract text / tables / structure tree / fonts / annotations / images / signature *fields* | **Does not verify cryptography.** Does not read the incremental-update history, does not map object IDs to coordinates, does not OCR | | [pdf-writer-mcp](https://github.com/shuji-bonji/pdf-writer-mcp) | Create, page operations, tagging, forms, annotations, metadata, attachments, PDF/A-3b scaffolding | **Does not sign.** Does not make the file meet the standard — it can write a *label*, not conformance | | [pdf-verify-mcp](https://github.com/shuji-bonji/pdf-verify-mcp) | Conformance validation (delegated to veraPDF), cryptographic signature verification, tamper detection, policy verdicts | **Does not prove the file meets the standard** (it can only find where it breaks the rules). Does not vouch for the signer's identity. Does not judge whether the content is true | > [!IMPORTANT] > **PDF specification files are NOT included in this package.** > You must obtain the PDF specification documents separately and place them in a local directory. > > **Download from:** [PDF Association — Sponsored Standards](https://pdfa.org/sponsored-standards/) > > See "[Setup](#setup)" for details. ## Features - **Multi-spec support** — Auto-discovers and manages up to 17 PDF-related documents (ISO 32000-2, PDF/UA, Tagged PDF guides, etc.) - **Structured content extraction** — Headings, paragraphs, lists, tables, and notes from any section - **Full-text search** — Keyword search with section-aware context snippets - **Requirements extraction** — Extracts normative language (shall / must / may) per ISO conventions - **Definitions lookup** — Term definitions from Section 3 (Definitions) - **Table extraction** — Multi-page table detection with header merging - **Version comparison** — Diff PDF 1.7 vs PDF 2.0 section structures - **Bounded-concurrency processing** — Parallel page processing for large documents - **On-disk index cache** — The search index and the full requirements scan are built once per PDF and reused by every later process (ISO 32000-2: ~6 s → ~0.2 s) ## Architecture ```mermaid graph LR subgraph Client["MCP Client"] LLM["LLM
(Claude, etc.)"] end subgraph Server["PDF Spec MCP Server"] direction TB MCP["MCP Server
index.ts"] subgraph Tools["Tools Layer"] direction LR T1["list_specs"] T2["get_structure"] T3["get_section"] T4["search_spec"] T5["get_requirements"] T6["get_definitions"] T7["get_tables"] T8["compare_versions"] end subgraph Services["Services Layer"] direction LR REG["Registry
Auto-discovery"] LOADER["Loader
LRU Cache"] SVC["PDFService
Orchestration"] CMP["CompareService
Version Diff"] end subgraph Extractors["Extractors"] direction LR OUTLINE["OutlineResolver
TOC & Section Index"] CONTENT["ContentExtractor
Structured Extraction"] SEARCH["SearchIndex
Full-text Search"] REQ["RequirementExtractor"] DEF["DefinitionExtractor"] end subgraph Utils["Utils"] direction LR CACHE["LRU Cache"] CONC["Concurrency"] VALID["Validation"] end end subgraph PDFs["PDF Spec Files (obtained separately)"] direction LR PDF1["ISO 32000-2
(PDF 2.0)"] PDF2["ISO 32000-1
(PDF 1.7)"] PDF3["TS 32001–32005
PDF/UA, etc."] end LLM <-->|"stdio / JSON-RPC"| MCP MCP --> Tools Tools --> Services Services --> Extractors Services --> Utils LOADER --> PDFs REG -->|"Filename pattern
auto-discovery"| PDFs style Client fill:#e8f4f8,stroke:#2196F3 style PDFs fill:#fff3e0,stroke:#FF9800 style Tools fill:#e8f5e9,stroke:#4CAF50 style Services fill:#f3e5f5,stroke:#9C27B0 style Extractors fill:#fce4ec,stroke:#E91E63 style Utils fill:#f5f5f5,stroke:#9E9E9E ``` ### Layer Overview | Layer | Responsibility | | -------------- | ---------------------------------------------------------------------------------- | | **Tools** | MCP tool schema definitions & handlers (input validation) | | **Services** | Business logic (PDF registry, loader, orchestration) | | **Extractors** | Information extraction from PDFs (TOC, content, search, requirements, definitions) | | **Utils** | Shared utilities (cache, concurrency, validation) | ## Setup ### 1. Obtain PDF Specification Files > [!WARNING] > PDF specifications are **copyrighted documents** and are not included in this package. > Download them from the sources below and place them in a local directory. | Document | Source | | ---------------------------- | ----------------------------------------------------------------------------------------------- | | ISO 32000-2 (PDF 2.0) | [PDF Association](https://pdfa.org/resource/iso-32000-pdf/) | | ISO 32000-1 (PDF 1.7) | [Adobe (free)](https://opensource.adobe.com/dc-acrobat-sdk-docs/pdfstandards/PDF32000_2008.pdf) | | TS 32001–32005, PDF/UA, etc. | [PDF Association — Sponsored Standards](https://pdfa.org/sponsored-standards/) | All 17 files below are supported. You do not need all of them — place only the specs you need (at minimum, ISO 32000-2 is recommended). ``` pdf-specs/ │ │ ── Standards ───────────────────────────── ├── ISO_32000-2_sponsored_EC3.pdf # iso32000-2 : PDF 2.0 EC3 (recommended; falls back to -ec2.pdf) ├── ISO_32000-2-2020_sponsored.pdf # iso32000-2-2020 : PDF 2.0 original ├── PDF32000_2008.pdf # pdf17 : PDF 1.7 (for version comparison) ├── pdfreference1.7old.pdf # pdf17old : Adobe PDF Reference 1.7 │ │ ── Technical Specifications (TS) ───────── ├── ISO_TS_32001-2022_sponsored_EC3.pdf # ts32001 : Hash extensions (SHA-3) ├── ISO_TS_32002-2022_sponsored_EC3.pdf # ts32002 : Digital signature extensions (ECC/PAdES) ├── ISO_TS_32003-2023_sponsored.pdf # ts32003 : AES-GCM encryption ├── ISO-TS-32004-2024_sponsored.pdf # ts32004 : Integrity protection ├── ISO-TS-32005-2023-sponsored.pdf # ts32005 : Namespace mapping │ │ ── PDF/UA (Accessibility) ──────────────── ├── ISO-14289-1-2014-sponsored.pdf # pdfua1 : PDF/UA-1 ├── ISO-14289-2-2024-sponsored.pdf # pdfua2 : PDF/UA-2 │ │ ── Guides ──────────────────────────────── ├── Tagged-PDF-Best-Practice-Guide.pdf # tagged-bpg : Tagged PDF Best Practice ├── Well-Tagged-PDF-WTPDF-1.0.pdf # wtpdf : Well-Tagged PDF ├── PDF-Declarations.pdf # declarations: PDF Declarations │ │ ── Application Notes ───────────────────── ├── PDF20_AN001-BPC.pdf # an001 : Black Point Compensation ├── PDF20_AN002-AF.pdf # an002 : Associated Files └── PDF20_AN003-ObjectMetadataLocations.pdf # an003 : Object Metadata ``` ### 2. Install This package ships a CLI binary (`pdf-spec-mcp`) intended to be launched by an MCP client. **You do not need to install it manually** — just point your MCP client to `npx @shuji-bonji/pdf-spec-mcp@latest` as shown in the next step. If you want to run it directly from the shell (e.g. for debugging): ```bash PDF_SPEC_DIR=/path/to/pdf-specs npx -y @shuji-bonji/pdf-spec-mcp@latest ``` Or install it globally (optional): ```bash npm install -g @shuji-bonji/pdf-spec-mcp PDF_SPEC_DIR=/path/to/pdf-specs pdf-spec-mcp ``` ### 3. Configure MCP Client #### Environment Variable | Variable | Description | Default | | -------------------- | --------------------------------------------------------------------- | ------------------------------------------ | | `PDF_SPEC_DIR` | Directory containing PDF specification files | (required) | | `PDF_SPEC_CACHE_DIR` | Where the on-disk index cache lives (see [Index cache](#index-cache)) | `${XDG_CACHE_HOME:-~/.cache}/pdf-spec-mcp` | | `PDF_SPEC_CACHE` | Set to `off` to neither read nor write the index cache | on | #### Claude Desktop Add to `claude_desktop_config.json`: ```json { "mcpServers": { "pdf-spec": { "command": "npx", "args": ["-y", "@shuji-bonji/pdf-spec-mcp@latest"], "env": { "PDF_SPEC_DIR": "/path/to/pdf-specs" } } } } ``` > [!IMPORTANT] > **Use `@latest` (or pin a version).** `npx -y ` without a version keeps running whatever > it cached the first time — `-y` only skips the install prompt, it does not check for updates. > A bare specifier will happily run a months-old release. `@latest` makes npx check the registry > on each start; pin `@0.4.0` instead if you want reproducibility. > To clear a stale cache: `rm -rf ~/.npm/_npx`. #### Cursor / VS Code Add to `.cursor/mcp.json` or VS Code MCP settings: ```json { "mcpServers": { "pdf-spec": { "command": "npx", "args": ["-y", "@shuji-bonji/pdf-spec-mcp@latest"], "env": { "PDF_SPEC_DIR": "/path/to/pdf-specs" } } } } ``` ## Index cache Two operations walk every page of a specification: the first `search_spec` on a spec builds its full-text index (ISO 32000-2, 1023 pages: about 6 s on a laptop), and `get_requirements` without a `section` scans every section (about 11 s). Everything else opens only the pages it needs and answers in well under a second. Since 0.5.0 those two results are written to disk after the first build and read back by every later process — an MCP client that starts one server per session no longer pays the build each time. The second process answers the same `search_spec` in about 0.2 s and the full requirements scan in about 0.02 s, from byte-for-byte the same index. - **Location:** `${PDF_SPEC_CACHE_DIR:-${XDG_CACHE_HOME:-~/.cache}/pdf-spec-mcp}/v1//...json`. The whole 17-spec corpus is about 18 MB per package version. - **Key:** package version, `pdfjs-dist` version, spec id, and the SHA-256 of the PDF. A replaced PDF, an upgraded server, or an upgraded pdfjs all miss and rebuild. Entries of older versions are left in place (another install may still use them); `--clear-cache` removes everything. - **Failure is a miss, never an error:** an unreadable, truncated, or foreign file is rebuilt; an unwritable directory is reported once on stderr and the server carries on without a cache. - **It is derived from *your* copy of the PDFs and stays on your machine.** It is not part of the package and must not be redistributed — the specifications are copyrighted. Nothing about searching changes: the same in-memory structure is searched by the same code. Only where it comes from (built vs. read) does. ### Pre-building the cache The cache fills lazily, one spec at a time as tools touch it. To warm every spec up front — after installing, after upgrading, or from cron — run the CLI (it uses the same code path as the tools, processes specs sequentially, and exits): ```bash PDF_SPEC_DIR=/path/to/pdf-specs npx -y @shuji-bonji/pdf-spec-mcp@latest --build-cache # --spec=iso32000-2,pdf17 only these specs # --force rebuild even when a valid entry exists npx -y @shuji-bonji/pdf-spec-mcp@latest --cache-info # directory, key, entries npx -y @shuji-bonji/pdf-spec-mcp@latest --clear-cache # remove the directory ``` A full build of the 17-spec corpus takes about a minute on a laptop. ## Available Tools All tools accept an optional `spec` parameter to target a specific specification (default: `iso32000-2`). | Tool | Description | | ------------------ | ----------------------------------------------------------------- | | `list_specs` | List all discovered PDF specifications with metadata | | `get_structure` | Get section hierarchy (table of contents) with configurable depth | | `get_section` | Get structured content of a specific section | | `search_spec` | Full-text keyword search across a specification | | `get_requirements` | Extract normative requirements (shall/must/may) | | `get_definitions` | Lookup term definitions | | `get_tables` | Extract table structures from a section | | `compare_versions` | Compare PDF 1.7 and PDF 2.0 section structures | ### `list_specs` — Discover Specifications List all available specification documents. Use the returned IDs as the `spec` parameter in other tools. ```jsonc // List all specs { } // Filter by category { "category": "ts" } // Technical specs only { "category": "pdfua" } // PDF/UA only { "category": "guide" } // Guide documents only ``` ### `get_structure` — Table of Contents Get the section hierarchy (TOC tree) of a specification. ```jsonc // PDF 2.0 top-level sections only { "max_depth": 1 } // Expand to 2 levels { "max_depth": 2 } // TS 32002 (Digital Signatures) full structure { "spec": "ts32002" } // PDF/UA-2 structure { "spec": "pdfua2", "max_depth": 2 } ``` ### `get_section` — Section Content Get structured content (headings, paragraphs, lists, tables, notes) of a specific section. A parent section returns its entire subtree (its preamble followed by all subsections, in document order). Top-level clauses can be very large — prefer the most specific section number. ```jsonc // PDF 2.0 Section 7.3.4.2 (Literal Strings) { "section": "7.3.4.2" } // PDF 2.0 Annex A { "section": "Annex A" } // TS 32002 Section 5 { "spec": "ts32002", "section": "5" } // PDF/UA-2 Section 8 (Tagged PDF) { "spec": "pdfua2", "section": "8" } ``` ### `search_spec` — Full-text Search Search across a specification with section-aware context snippets. The first call on a spec builds its index (a few seconds); the index is then cached on disk (see [Index cache](#index-cache)). ```jsonc // Search PDF 2.0 for "digital signature" { "query": "digital signature" } // Limit results { "query": "font", "max_results": 5 } // Search within TS 32002 { "spec": "ts32002", "query": "CMS" } ``` ### `get_requirements` — Normative Requirements Extract normative requirements (shall / must / may) per ISO conventions. ```jsonc // All requirements in section 12.8 { "section": "12.8" } // Only "shall" requirements { "section": "12.8", "level": "shall" } // Only "shall not" requirements { "section": "7.3", "level": "shall not" } // PDF/UA-2 requirements { "spec": "pdfua2", "section": "8", "level": "shall" } ``` ### `get_definitions` — Term Definitions Look up term definitions from Section 3 (Definitions). ```jsonc // Search for "font" definitions { "term": "font" } // List all definitions { } // PDF/UA definitions { "spec": "pdfua2", "term": "artifact" } ``` ### `get_tables` — Table Extraction Extract table structures (headers, rows, captions) from a section. Multi-page tables are automatically merged. ```jsonc // All tables in section 7.3.4.2 (Table 3 — Escape sequences) { "section": "7.3.4.2" } // Specific table only (0-based index) { "section": "7.3.4.2", "table_index": 0 } // TS spec tables { "spec": "ts32002", "section": "5" } ``` ### `compare_versions` — Version Comparison Compare section structures between PDF 1.7 (ISO 32000-1) and PDF 2.0 (ISO 32000-2). Uses title-based automatic matching to detect matched, added, and removed sections. > [!NOTE] > This tool requires both PDF 1.7 (`PDF32000_2008.pdf`) and PDF 2.0 files in `PDF_SPEC_DIR`. ```jsonc // Diff section 12.8 (Digital Signatures) { "section": "12.8" } // Compare all top-level sections { } ``` ## Supported Specifications The server auto-discovers PDF files in `PDF_SPEC_DIR` by filename pattern matching: | Category | Spec IDs | Documents | | ------------------ | ---------------------------------------------------- | ------------------------------------------------------- | | **Standard** | `iso32000-2`, `iso32000-2-2020`, `pdf17`, `pdf17old` | ISO 32000-2 (PDF 2.0), ISO 32000-1 (PDF 1.7) | | **Technical Spec** | `ts32001` – `ts32005` | Hash, Digital Signatures, AES-GCM, Integrity, Namespace | | **PDF/UA** | `pdfua1`, `pdfua2` | Accessibility (ISO 14289-1, 14289-2) | | **Guide** | `tagged-bpg`, `wtpdf`, `declarations` | Tagged PDF, Well-Tagged PDF, Declarations | | **App Note** | `an001` – `an003` | BPC, Associated Files, Object Metadata | ## Directory Structure ``` src/ ├── index.ts # Entry point: MCP server on stdio, or the cache CLI ├── cli.ts # --build-cache / --clear-cache / --cache-info ├── config.ts # Configuration & spec patterns ├── errors.ts # Error hierarchy (PDFSpecError → sub-classes) ├── services/ │ ├── pdf-registry.ts # Auto-discovery of PDF files │ ├── pdf-loader.ts # PDF loading with LRU cache │ ├── pdf-service.ts # Orchestration layer │ ├── index-store.ts # On-disk cache for the search / requirements indexes │ ├── compare-service.ts # Version comparison │ ├── outline-resolver.ts # Section index builder │ ├── content-extractor.ts # Structured content extraction │ ├── search-index.ts # Full-text search index │ ├── requirement-extractor.ts │ └── definition-extractor.ts ├── tools/ │ ├── definitions.ts # MCP tool schemas │ └── handlers.ts # Tool implementations ├── types/ │ └── index.ts # Shared type definitions └── utils/ ├── concurrency.ts # mapConcurrent (bounded Promise.all) ├── text.ts # Text normalization ├── cache.ts # LRU cache ├── file-hash.ts # SHA-256 of a PDF (index cache key) ├── validation.ts # Input validation └── logger.ts # Structured logger ``` ## Development ```bash git clone https://github.com/shuji-bonji/pdf-spec-mcp.git cd pdf-spec-mcp npm install npm run build # Unit tests npm run test # E2E tests (requires PDF files in ./pdf-spec/) npm run test:e2e # Lint & format npm run lint npm run format:check ``` ## License [MIT](LICENSE)