--- name: paper description: Academic paper workflow — find, read/annotate (HTML), daily digest, cite. Use when user invokes /paper, needs paper annotation, literature search, citation generation, or daily paper discovery. --- # /paper — Academic Paper Workflow ## Default Config Edit this section to customize for your project. ```yaml output_dir: "./papers" # Obsidian vault users: set paper_output_dir in your project's CLAUDE.md # Drive/OneDrive: use the appropriate MCP, then set path (e.g. "gdrive://My Drive/Papers") venues: [CHI, ACL, EMNLP, NAACL, NeurIPS, ICML, ICLR, EDM, LAK, AIED, ITS, CSCW, SIGIR, CIKM] keywords: - cognitive load - adaptive learning - intelligent tutoring - conversational learning - LLM tutoring - learning analytics - behavioral signals - educational dialogue min_citations: 5 # /paper find: filter out papers with fewer citations daily_max: 10 # /paper digest: max papers per run annotation_lang: zh # zh = Chinese annotations | en = English; override with --lang openalex_email: "" # Optional. Add email to join OpenAlex Polite Pool (higher rate limits). semantic_scholar_api_key: "" # Optional. Free key: semanticscholar.org/product/api ``` Override any config value in your project's `CLAUDE.md` using keys: `paper_output_dir`, `paper_venues`, `paper_keywords`, `paper_annotation_lang`. --- ## Subcommands | Command | Usage | Description | |---------|-------|-------------| | `find` | `/paper find "cognitive load LLM"` | Search papers → Digest HTML with top N cards | | `read` | `/paper read /path/to/file.pdf` | Annotate a single PDF → full dual-column HTML | | `digest` | `/paper digest` | Daily new papers from arxiv (used by cron) | | `cite` | `/paper cite [[note-name]]` | Generate APA + BibTeX citation from existing annotation | **Flags:** - `--context "ML课第3周"` — override project context - `--questions "Q1:... Q2:..."` — switch to Question mode (Mode A) - `--questions` — interactive question mode (Claude asks you) - `--output /path/` — override output directory for this run - `--top N` — return N results (default: 5 for find, daily_max for digest); also accepts bare number after query - `--lang en` — override annotation language for this run - `--source arxiv` — search only arxiv (latest preprints, no citation filter) - `--source venues` — search only papers from config `venues` list (citation-weighted ranking) - `--local /path/` — read local PDFs from folder (no API calls); or `--local a.pdf b.pdf` for specific files > *Flags can be written with or without `--`. Claude accepts natural language equivalents (e.g. `top 5`, `source venues`, `local /path/`, `questions "..."`).* --- ## Context Resolution (in priority order) 1. `--context` flag in command 2. Current directory's `CLAUDE.md` (auto-loaded by Claude Code) 3. No context → use Mode B directly (zero-config, logic analysis mode) --- ## Data Sources ### Primary: OpenAlex Free, no API key required, 10 req/s rate limit. ``` https://api.openalex.org/works?search={query} &select=title,authorships,publication_year,cited_by_count,primary_location,doi &per-page=20 [&mailto={openalex_email} ← add if configured] ``` **After fetching:** immediately extract and keep only — title, first author, year, cited_by_count, venue name (`primary_location.source.display_name`), DOI. Discard all other fields from the response before further processing. ### Secondary: Semantic Scholar (only if `semantic_scholar_api_key` is set) ``` https://api.semanticscholar.org/graph/v1/paper/search?query={topic} &fields=title,authors,year,citationCount,venue,externalIds&limit=20 ``` Header: `x-api-key: {semantic_scholar_api_key}` ### Tertiary: arxiv (fallback + digest source) ``` https://export.arxiv.org/api/query?search_query=(cat:cs.CL+OR+cat:cs.HC+OR+cat:cs.AI+OR+cat:cs.LG) +AND+({keywords})&sortBy=submittedDate&sortOrder=descending&max_results=30 ``` Note: arxiv papers have no citation count. Label as `preprint`. ### Rate Limit Handling If a source returns 429: check `Retry-After` header → wait that many seconds. If no header: wait 2s → 5s → 10s (3 retries). After 3 failures → skip source, move to next in priority. Note which source was skipped in output. --- ## Subcommand: find 1. **Parse flags:** topic query, `--top N` (default: 5), `--lang`, `--source` (default: mixed), `--local` (path or file list) - `--top N` also accepts a bare number immediately after the query string (e.g. `/paper find "query" 10` → top 10) 2. **Resolve source strategy:** | `--source` / flag | API calls | Ranking formula | |-------------------|-----------|----------------| | *(default, mixed)* | Single OpenAlex call `sort=cited_by_count:desc`; arxiv fallback if <3 results | `0.5 × relevance_rank + 0.3 × log(cited_by_count+1) + 0.2 × recency_score` | | `arxiv` | arxiv API only (`sortBy=submittedDate`); no `min_citations` filter | `0.5 × relevance_rank + 0.5 × recency_score` | | `venues` | Single OpenAlex call; keep only papers where venue name matches any entry in config `venues` list (case-insensitive partial match on `primary_location.source.display_name`) | `0.5 × relevance_rank + 0.5 × log(cited_by_count+1)` | | `--local /path/` or `--local a.pdf b.pdf` | No API calls — read local PDFs only | Take first `--top N` files (alphabetical order) | **`--local` and `--source` are mutually exclusive.** If both are given, return an error. 3. **If `--local` flag present — execute local PDF steps instead of steps 3–6:** 1. Scan the given folder for all `.pdf` files, or use the explicitly listed files directly 2. If file count > `--top N`, take first N (alphabetical order) 3. For each PDF, extract: title, authors, year, abstract (first 300 words), conclusion paragraph (last 300 words) 4. Generate a Digest Card per PDF (same 4-sentence format: 研究问题 / 核心方法 / 关键发现 / 相关性) 5. Citation badge: fixed as `📄 Local PDF` (no OpenAlex lookup) 6. Card bottom link: `/paper read /absolute/path/to/file.pdf 精读` (use absolute path) 7. Save to: `{output_dir}/FindResults/find-local-YYYY-MM-DD-HHmm-{folder-slug}.html` 8. Print summary and exit — skip steps 4–8 below 3b. **Fetch papers (online mode):** - Add `&filter=cited_by_count:>{min_citations}` for default/venues mode (skip for arxiv) - After fetch: immediately discard all fields except title, first author, year, cited_by_count, venue name, DOI - On any source failure → follow Rate Limit Handling 4. **Deduplicate against existing files:** - Check `{output_dir}/FindResults/` and `{output_dir}/` root for `[FirstAuthor-YYYY-slug].html` - Mark existing papers as `[已有]`, include in summary with their path, skip card generation 5. **Rank top N** using the formula for the active source strategy 6. **Generate Digest HTML** — a single self-contained HTML file containing N paper cards. **Digest Card format** (one card per paper): ``` [Title] — [First Author] et al., [Year] — [Venue or arXiv] [Citation badge] [Source tag] 研究问题: [1 sentence — what problem does this paper solve?] 核心方法: [1 sentence — what approach do they use?] 关键发现: [1 sentence — what is the main result?] 相关性: [1 sentence — why this matters for your research, based on CLAUDE.md context] [DOI link if available] · → /paper read 精读 ``` - Citation badge uses tier logic from Paper Quality Bar section - Annotation language for cards follows `annotation_lang` config (or `--lang` flag) - HTML structure: simple cards layout, embed `references/template.css` in `