--- name: question2report description: > Turn a natural-language financial question into a polished, self-contained HTML report. Covers the full pipeline: requirement analysis → scope negotiation → data fetching → data cleaning → quantitative analysis → visualization → beautiful HTML output. Use when the user asks for any financial data comparison, fund screening, index analysis, or strategy back-test that should end with a deliverable report. argument-hint: --- # Question → Report Skill Transform a user's free-form financial question into a production-quality, self-contained HTML report with embedded charts and tables. ## Pipeline ``` User Question │ ▼ ┌──────────────────┐ │ 1. ANALYZE │ Parse intent, identify assets, metrics, time range └────────┬─────────┘ │ ▼ ┌──────────────────┐ │ 2. CONFIRM │ Present analysis plan to user; agree on scope └────────┬─────────┘ │ ▼ ┌──────────────────┐ │ 3. DISCOVER API │ Explore the xalpha codebase to find suitable APIs └────────┬─────────┘ │ ▼ ┌──────────────────┐ │ 4. FETCH & CLEAN │ Write & run a Python script; handle errors & NaN └────────┬─────────┘ │ ▼ ┌──────────────────┐ │ 5. ANALYZE DATA │ Compute metrics appropriate to the question └────────┬─────────┘ │ ▼ ┌──────────────────┐ │ 6. GENERATE HTML │ Build a beautiful, self-contained HTML report └──────────────────┘ ``` --- ## Step 1 — Analyze the Question Parse the user's natural-language question and extract: - **Subject**: What assets, funds, indices, or strategies are being discussed? - **Comparison / benchmark**: Is there a reference to compare against? - **Time range**: Explicit dates, or implied ("last 3 years", "since inception"). Default to the **most recent 3 full calendar years** if unspecified. - **Desired output**: What kind of insights does the user want? (rankings, trend comparison, risk analysis, prediction accuracy, etc.) If fund codes or asset identifiers are not given, **research them** via web search or by exploring the xalpha codebase for relevant list/search APIs. ## Step 2 — Confirm Scope Before any data work, present a concise plan to the user: ``` 📋 Analysis Plan ───────────────────────────────── Subject : Assets : Period : → Analysis : Charts : ───────────────────────────────── Shall I proceed, or would you like to adjust? ``` **Wait for user confirmation** before proceeding. Adjust scope if requested. ## Step 3 — Discover Suitable APIs **Do NOT assume which xalpha APIs to use.** Instead: 1. Explore the codebase — read `xalpha/__init__.py`, outlines of key modules like `universal.py`, `info.py`, `toolbox.py`, `evaluate.py`, `indicator.py`, etc. 2. Identify the right functions for the user's question. The xalpha library is rich: it supports funds, indices, stocks, bonds, QDII, commodities, forex, PE/PB valuation, portfolio back-testing, holdings analysis, and more. 3. Check function signatures and docstrings to understand parameters, return types, and any known quirks (e.g. some classes don't accept `start` in `__init__`). 4. If you encounter an API error at runtime, read the traceback, explore the source for alternatives, and fix the script. Be resilient. ## Step 3.1 — Common xalpha APIs & Interfaces To accelerate discovery, prioritize these common interfaces in the `xalpha` package: ### Data Fetching (`xalpha.universal`) - `xa.get_daily(code, start=None, end=None)`: The "universal" historical data fetcher. - **A-Share**: `SH600000` (prefix SH/SZ + 6-digit code). - **HK-Share**: `HK00700`. - **US-Share**: `AAPL`, `MSFT`. - **Funds**: `F000001` (unit net value), `T000001` (accumulated net value). - **Valuation**: `peb-SH000300` (Index PE/PB, **requires JQData**), `peb-600000` (Stock PE/PB, handles 6-digit codes automatically, no JQ required). - `xa.get_rt(code)`: Fetches real-time price and basic metadata (name, market, etc.). ### Fund Analysis (`xalpha.info`) - `xa.fundinfo(code, path=None, priceonly=True)`: Core class for fund data. - `.price`: DataFrame with `date` and `netvalue`. - `.get_holdings(year, season)`: Quarterly holdings data. ### Backtesting & Portfolios (`xalpha.trade`, `xa.multiple`) - `xa.trade(fund_obj, status_df)`: Backtests a single fund/asset based on a transaction table (`status_df`). - `xa.itrade(fund_obj, status_df)`: For exchange-traded assets (stocks/ETFs). - `xa.multiple(trade_list)`: Aggregates multiple `trade`/`itrade` objects into a portfolio. - `.v_totvalue()`: Visualizes the portfolio total value curve. - `.combsummary()`: Generates a summary table of the portfolio performance. ### Evaluation & Comparison (`xalpha.evaluate`, `xa.toolbox`) - `xa.evaluate(asset_obj)`: Provides comprehensive performance metrics (Sharpe, Max Drawdown). - `xa.compare(list_of_objs, start=None)`: Compares multiple assets/strategies on a normalized (1.0) scale. ### Technical Indicators (`xalpha.indicator`) - `xa.indicator(daily_df)`: Wraps a daily price DataFrame to compute indicators like MA, RSI, MACD. ## Step 4 — Fetch & Clean Data Write a **single self-contained Python script** that: 1. Imports `xalpha` and any needed stdlib/pandas/numpy modules. 2. Fetches all required data using the APIs discovered in Step 3. 3. **Workspace Organization**: If you create scratch scripts (e.g., `test_api.py`, `diag.py`) or temporary files to debug, create them directly in the target report folder or move them there immediately. 4. Handles errors gracefully — if one data source fails, try fallbacks; if one asset in a list fails, skip it and warn rather than crash. 5. Cleans the data: - Ensure dates are `datetime64`. - Handle NaN: forward-fill small gaps (≤3 days); drop assets with >30% missing. - Align data on common trading dates when comparing multiple series. 6. Saves intermediate results to a temporary workspace CSV (or keeps in memory if the script does everything in one pass). Run the script using the user's specified Python/conda environment. If not specified, use the system default. ## Step 5 — Quantitative Analysis Compute metrics **appropriate to the user's question**. Do not blindly apply a fixed set of metrics. Choose what makes sense: - For return comparison: total return, annualized return, excess return / alpha. - For risk analysis: max drawdown, annualized volatility, Sharpe ratio. - For tracking analysis: tracking error, information ratio, correlation. - For valuation: PE/PB percentiles, dividend yield. - For prediction accuracy: predicted vs actual, RMSE, hit rate. - For portfolio analysis: asset allocation, sector exposure, concentration. The agent should determine which metrics are relevant based on the question context. ## Step 5.1 — Pro-Tips for Reliable Analysis Apply these principles to avoid common pitfalls in quantitative reporting: - **Unit Consistency**: Verify if your backtesting engine treats trade inputs as **shares** or **cash value**. Inconsistent handling (e.g., selling "100 shares" but interpreting it as "100 dollars") is a frequent cause of hidden performance erosion. - **Cross-Verification**: Don't trust high-level CAGR/NAV metrics blindly. Periodically reconcile the final liquidity by summing raw cash flows from underlying trade logs. - **Contextual Visuals**: Use timeline mapping (Gantt-style) to show when a strategy was "Active" vs. "Hedged/In Cash". This explains *why* a performance gap occurred, providing more insight than a simple NAV curve. - **Document Framework Patches**: If you apply a library-specific fix (e.g., a monkey-patch or a non-standard initialization) to bypass a known bug, document it clearly in the methodology notes. ## Step 6 — Generate the HTML Report This is the most important step for user experience. The report must be: ### Interactive & Self-Contained - **All CSS, JS logic, and raw data are embedded.** The HTML file must render perfectly when opened directly in any browser. - **Interactive Charts**: Load high-quality interactive charting libraries (specifically Apache ECharts, or alternatively Chart.js) from highly reliable public CDNs (e.g., `cdn.jsdelivr.net` or `cdnjs.cloudflare.com`). ECharts is strongly recommended for financial data because of its native support for data zoom sliders, interactive legends, crosshairs, CJK tooltips, and high-quality styling. - **Data Serialization**: In the python script, serialize raw timeseries data (dates, returns, alpha curves, drawdowns) as a JSON object, and inject it inside the HTML in a `