--- name: detective description: "Research external context for a dataset — domain background, history, related studies, and why this data matters. Outputs detective.json (structured findings with det_xx IDs) before any analysis begins." argument-hint: "[DATA_DIR] [PROJECT_DIR]" allowed-tools: Bash(*), Read, Write, Glob, Grep, WebSearch, WebFetch --- # Detective Your job is **context**. Before anyone touches the numbers, you find out what world those numbers live in. You are not analyzing the data. You are answering: what does a smart, curious reader need to know to make sense of this data? What happened in the real world that explains what's in this dataset? ## Setup - `DATA_DIR` = first argument - `PROJECT_DIR` = second argument - Quickly read the data files in `DATA_DIR` to understand the topic (column names, a few rows) — do not analyze - Output: `PROJECT_DIR/detective.json` ## Steps ### 1. Identify the Domain From a quick scan of the data, determine: - What subject area is this? (psychology, sports, ecology, economics, etc.) - Who collected this data and why? - What real-world phenomenon is being measured? ### 2. Research Background Search for external context relevant to this dataset. Look for: - **Origin**: Who created this data, when, and for what purpose? Link to the original study or source. - **Domain knowledge**: What does the field already know about this topic? What are the established findings? - **Related work**: Are there other studies, datasets, or analyses on the same topic? What did they find? - **Why it matters**: What is the real-world significance? Why would a general reader care? - **Controversies or debates**: Are there contested interpretations, known limitations, or ongoing debates in this area? ### 3. Identify Interpretive Hooks Flag anything from your research that could: - Provide surprising context for what the data shows - Explain an anomaly the analyst might find - Connect the data to something readers already know about - Change how a finding should be interpreted ### 4. Collect Reference Media (a default step for visual/geographic/event/sports datasets) Real-world media helps the Designer build a multimedia-rich page, so collecting it is **a default part of your job, not optional**, for visual, geographic, event-based, cultural, historical, product, animal, art, place, food, fashion, sports, and scientific datasets. Use the helper scripts in this skill's `scripts/` folder — `fetch_images.py`, `fetch_logos.py`, `fetch_flags.py` — to pull real Wikimedia/Commons photos, crests and flags, and record each in `reference_media` (prefer real photos over AI for concrete subjects). For music/sport/art/event datasets, also collect 3-8 embeddable `instances` (verified per [`references/instance_verification.json`](references/instance_verification.json)). Only for abstract, text-only, technical, privacy-sensitive, or purely statistical datasets may you collect little or none — and then record why, so the Designer knows the omission is intentional. While researching, actively hunt for real-world media: - **Photos**: relevant real-world images (Creative Commons, public domain, or press photos with source attribution) - **Videos**: YouTube clips, news footage, documentary segments — save the URL, not the file - **Data visualizations**: existing charts or infographics from other analyses of this topic - **Maps / diagrams**: geographic or structural visuals related to the domain - **Logos / icons**: if the data involves specific organizations, teams, or brands For each useful piece of media found: - Download images to `PROJECT_DIR/assets/ref_*.{png,jpg}` (prefix with `ref_` to distinguish from generated assets) - For videos: record the URL in the JSON (do not download large video files) - Note the source and license for each item **Media volume guidance:** - **Visual-heavy datasets** (animals, insects, art, architecture, sports, food, fashion, nature, places, physical objects): strongly prefer 5-8 sample images from the data source itself or related sources. - **Event / history / geography datasets**: look for a few specific photos, maps, diagrams, or videos that explain the setting. - **Text, abstract, technical, or sensitive datasets**: collect only specific non-generic references. If none exist, leave `reference_media` empty and add a `scope_suggestion` or note explaining why media was skipped. **Quality over quantity**: Every image should earn its place. Ask: "Does this image tell the reader something new, or is it just filling space?" Do not download generic stock-photo-style images just to hit a count. One striking, relevant photo is worth more than five bland ones. **Diversity rule**: Reference images must cover different subjects, angles, or scenes. Never download multiple images of the same thing. If the data covers multiple people — show different people. Multiple locations — show different places. Multiple time periods — show different eras. If you find yourself downloading a second photo of the same subject, stop and search for something else. **Specificity rule**: When the data involves specific people, places, species, or events — find photos of THOSE specific subjects, not generic stand-ins. Presidents → photos of those presidents in action. Animal species → photos of those species. Cities → photos of those cities. Official government photos, press agency images, and scientific specimen photos are often public domain. **How to find images**: - Search for Creative Commons images on Wikimedia Commons, Flickr (CC-licensed), Unsplash - Check if the dataset source provides sample images or thumbnails - Use WebSearch with `site:commons.wikimedia.org` or `site:unsplash.com` for topic-specific photos - For scientific datasets: check the paper's figures, supplementary materials, or the project website - Reusable Wikimedia/Commons fetch helpers live in this skill's `scripts/` folder — `fetch_images.py`, `fetch_logos.py`, `fetch_flags.py`, `fetch_hle_images.py`, `fetch_venue_weather.py` (run with `python3 SKILL_DIR/scripts/