- [2025/12/14] We have updated the MCP authentication mechanism for the Notion official MCP server. Please refer to the [How to Register Accounts](./global_preparation/how2register_accounts.md) for more details. - [2025/12/12] The original single factor authentication mechanism for Snowflake is not working if the account finishes the trial period. We need to use RSA public key authentication instead.= and we have updated our codebases and documentation to support this new authentication mechanism. Please refer to the renewed [How to Register Accounts](./global_preparation/how2register_accounts.md) for more details. - [2025/12/12] Using Claude models via AWS Bedrock backend will sometimes leads to errors like this, we suggeet using Anthropic's official endpoint instead. ``` {"title":"Bad request","detail":"{\"message\":\"Expected toolResult blocks at messages.2.content for the following Ids: tooluse_z1WOsx8aThCAfWyK14mgEQ\"}","status":400} ``` - [2025/12/12] We add a new model provider `openai_stateful_responses` to support the Responses API for official OpenAI models. We suggest using this model provider instead of the `unified` model provider (based on ChatCompletions API) when using official OpenAI models to keep the consistency of the thinking contents generated by the model. ## Task Review Records The following temporary local review trackers were consolidated here on 2026-06-18 before pushing `toolathlon-tasks-verified-toreview`. # PR 48 Task Review Tracker Branch: `toolathlon-tasks-verified-toreview` Diff basis: `origin/main..HEAD` Line churn is insertions plus deletions from `git diff --numstat`. Binary file changes are noted separately and are not included in line churn. | Done | Bucket | Task | Files | Line churn | Binary changes | Notes | | --- | --- | --- | ---: | ---: | ---: | --- | | [x] | Shared | `_shared_utils` | 1 | 109 | 0 | Reviewed by user on 2026-06-17; shared grade_with_retry helper accepted | | [x] | Tiny | `meeting-assign` | 1 | 3 | 0 | Reviewed by user before 2026-06-08 | | [x] | Tiny | `woocommerce-new-product` | 1 | 3 | 0 | Reviewed by user before 2026-06-08 | | [x] | Tiny | `canvas-homework-grader-python` | 1 | 5 | 0 | Reviewed by user before 2026-06-08 | | [x] | Tiny | `payable-invoice-checker` | 1 | 5 | 0 | Reviewed by user before 2026-06-08 | | [x] | Tiny | `course-assistant` | 1 | 6 | 0 | Reviewed by user before 2026-06-08 | | [x] | Tiny | `notion-hr` | 1 | 6 | 0 | Reviewed by user on 2026-06-08 | | [x] | Tiny | `task-tracker` | 1 | 7 | 0 | Reviewed by user on 2026-06-08 | | [x] | Tiny | `notion-find-job` | 1 | 8 | 0 | Reviewed by user on 2026-06-08 | | [x] | Tiny | `notion-movies` | 1 | 8 | 0 | Reviewed by user on 2026-06-08 | | [x] | Tiny | `git-bug-hunt` | 1 | 9 | 0 | Reviewed by user on 2026-06-08 | | [x] | Tiny | `inter-final-performance-analysis` | 1 | 9 | 0 | Reviewed by user on 2026-06-08; tuple-return fix applied | | [x] | Tiny | `notion-personal-website` | 2 | 9 | 0 | Reviewed by user on 2026-06-08 | | [x] | Tiny | `quantitative-financial-analysis` | 1 | 10 | 0 | Reviewed by user on 2026-06-08 | | [x] | Tiny | `woocommerce-customer-survey` | 2 | 10 | 0 | Reviewed by user on 2026-06-08 | | [x] | Tiny | `woocommerce-stock-alert` | 1 | 10 | 0 | Reviewed by user on 2026-06-08 | | [x] | Tiny | `filter-low-selling-products` | 2 | 10 | 0 | Reviewed by user on 2026-06-08 | | [x] | Light | `sla-timeout-monitor` | 1 | 16 | 0 | Reviewed by user on 2026-06-08; negative email polling fix applied | | [x] | Light | `language-school` | 1 | 19 | 0 | Reviewed by user on 2026-06-09; task prompt/evaluator updated to use current-year information. | | [x] | Light | `inventory-sync` | 1 | 20 | 0 | Reviewed by user on 2026-06-08; cleanup lifecycle fix applied | | [x] | Light | `ab-testing` | 1 | 20 | 0 | Reviewed by user on 2026-06-08 | | [x] | Light | `llm-training-dataset` | 1 | 23 | 0 | Reviewed by user on 2026-06-09 | | [x] | Light | `canvas-new-students-notification` | 1 | 23 | 0 | Reviewed by user on 2026-06-09 | | [x] | Light | `imagenet` | 1 | 26 | 0 | Reviewed by user on 2026-06-09 | | [x] | Light | `canvas-arrange-exam` | 1 | 27 | 0 | Reviewed by user on 2026-06-09 | | [x] | Light | `arrange-workspace` | 1 | 29 | 0 | Reviewed by user on 2026-06-09 | | [x] | Light | `student-interview` | 1 | 29 | 0 | Reviewed by user on 2026-06-11 | | [x] | Light | `apply-phd-email` | 2 | 30 | 0 | Reviewed by user on 2026-06-11 | | [x] | Moderate | `course-schedule` | 1 | 34 | 0 | Reviewed by user on 2026-06-11 | | [x] | Moderate | `canvas-do-quiz` | 2 | 34 | 0 | Reviewed by user on 2026-06-11 | | [x] | Moderate | `sync-todo-to-readme` | 1 | 36 | 0 | Reviewed by user on 2026-06-11; missing TODO groundtruth entries added | | [x] | Moderate | `gdp-cr5-analysis` | 1 | 39 | 0 | Reviewed by user on 2026-06-11 | | [x] | Moderate | `experiments-recordings` | 1 | 40 | 0 | Reviewed by user on 2026-06-11 | | [x] | Moderate | `k8s-deployment-cleanup` | 3 | 41 | 0 | Reviewed by user on 2026-06-12; negative email polling fix applied | | [x] | Moderate | `music-analysis` | 3 | 45 | 1 | Reviewed by user on 2026-06-12; 1942 summary section excluded from groundtruth | | [x] | Moderate | `update-material-inventory` | 2 | 55 | 0 | Reviewed by user on 2026-06-12 | | [x] | Moderate | `travel-expense-reimbursement` | 1 | 68 | 0 | Reviewed by user on 2026-06-12 | | [x] | Moderate | `k8s-redis-helm-upgrade` | 2 | 68 | 0 | Reviewed by user on 2026-06-12; return-code/runtime fixes applied | | [x] | Moderate | `nvidia-stock-analysis` | 1 | 71 | 0 | Reviewed by user on 2026-06-12; Sheet 1 snapshot GT, Sheet 3 alignment, MCP/tooling prompt fixes applied | | [x] | Moderate | `nhl-b2b-analysis` | 1 | 72 | 0 | Reviewed by user on 2026-06-12; spreadsheet discovery moved inside retry window | | [x] | Moderate | `wandb-best-score` | 1 | 79 | 0 | Reviewed by user on 2026-06-12 | | [x] | Moderate | `fillout-online-forms` | 1 | 85 | 0 | Reviewed by user on 2026-06-12 | | [x] | Moderate | `k8s-mysql` | 2 | 88 | 0 | Reviewed by user on 2026-06-12; podman/docker image preload and init failure propagation fixes applied | | [x] | Moderate | `canvas-art-manager` | 1 | 97 | 0 | Reviewed by user on 2026-06-12 | | [x] | Heavier | `canvas-submit-late-work` | 2 | 105 | 0 | Reviewed by user on 2026-06-12 | | [x] | Heavier | `k8s-safety-audit` | 3 | 113 | 0 | Reviewed by user on 2026-06-12; podman/docker image preload and script failure propagation fixes applied | | [x] | Heavier | `investment-decision-analysis` | 1 | 117 | 0 | Reviewed by user on 2026-06-12; sheet fetch error handling fix applied | | [x] | Heavier | `live-transactions` | 1 | 117 | 0 | Reviewed by user on 2026-06-12; Cloud Logging filter grouping fix applied | | [x] | Heavier | `wandb-shortest-length` | 2 | 124 | 0 | Reviewed by user on 2026-06-13; final-checkpoint row comparison fix applied | | [x] | Heavier | `k8s-pr-preview-testing` | 2 | 125 | 0 | Reviewed by user on 2026-06-13; exact kind cleanup, runtime preload, and k8s instance-suffix alignment fixes applied | | [x] | Heavier | `academic-warning` | 1 | 130 | 0 | Reviewed by user on 2026-06-15; negative CRITICAL log polling fix applied and stale FIXME removed | | [x] | Heavier | `detect-revised-terms` | 2 | 138 | 0 | Reviewed by user on 2026-06-15; prompt clarified, GT aligned, exact-match evaluator accepts full/without-final-row variants | | [x] | Heavier | `canvas-art-quiz` | 1 | 147 | 0 | Reviewed by user on 2026-06-15 | | [x] | Heavier | `find-alita-paper` | 1 | 150 | 0 | Reviewed by user on 2026-06-15 | | [x] | Heavier | `woocommerce-update-cover` | 1 | 162 | 0 | Reviewed by user on 2026-06-15; evaluation failure summary aligned with 100% requirement | | [x] | Big | `privacy-desensitization` | 3 | 240 | 1 | Reviewed by user on 2026-06-15; archive GT corrected, macOS metadata removed, preprocess/evaluator tar handling aligned | | [x] | Big | `woocommerce-product-recall` | 6 | 502 | 0 | Reviewed by user on 2026-06-16; strict form-template/content validation, email form-link body check, retry fallback, and prompt/template alignment applied | | [x] | Big | `oil-price` | 2 | 419 | 0 | Reviewed by user on 2026-06-16; summary field checks, retry log refresh, and backtest docstring alignment applied. Caveat: this task was not checked very carefully or exhaustively because it involves financial-domain assumptions. | | [x] | Big | `vlm-history-completer` | 2 | 390 | 0 | Reviewed by user on 2026-06-16; URL-source matching, retry wrapper, and groundtruth source updates accepted | | [x] | Big | `search-ca-school` | 3 | 434 | 0 | Reviewed by user on 2026-06-17; evaluator aligned with live CSRankings displayed ranking semantics, including default-off KDD, US institutions from institutions.csv, competition-rank ties, updated fallback JSON, and stale-warning fallback | | [x] | Big | `merge-hf-datasets` | 2 | 1037 | 0 | Reviewed by user on 2026-06-17; XLAM groundtruth schema wrapper accepted and evaluator type aliases extended to optional type strings | # PR54 Task Review Tracker Source PR: https://github.com/hkust-nlp/Toolathlon/pull/54 Branch under review: `toolathlon-tasks-verified-toreview` Initial scope is taken from the PR #54 body and changed-file list. As of the initial tracker setup, the GitHub connector returned no PR comments/review comments, so review progress starts from the PR body entries until thread-level feedback is available. Status legend: - `[ ]` not reviewed - `[~]` in review / needs discussion - `[x]` reviewed and accepted - `[!]` reviewed and needs follow-up fix | Status | Area | PR commit | Changed files | Claimed fix | Progress notes | | --- | --- | --- | --- | --- | --- | | [x] | `canvas-do-quiz` | `71a876bf` | `tasks/finalpool/canvas-do-quiz/evaluation/check_remote.py` | Read `kept_score` instead of latest-attempt `score` for pre-completed quizzes with leftover untaken attempts. | Accepted and applied; all configured quizzes use `keep_highest`, so the evaluator now prefers gradebook `kept_score` and falls back to latest-attempt `score`. | | [x] | `university-course-selection` | `d762e4f2` | `tasks/finalpool/university-course-selection/evaluation/check_local.py` | Add softer field comparison for instructor titles, restrictions, `无`/empty cells, and class-time week ranges. | Reviewed and rejected; the reference Excel explicitly specifies `周x 第x-y节` class-time formatting and `无` for unrestricted courses, and the existing `normalize_str` comparison already absorbs whitespace/punctuation differences in restrictions without accepting empty cells, reordered tokens, or copied PDF titles/week ranges. | | [x] | `update-material-inventory` | `2c934bd5` | `tasks/finalpool/update-material-inventory/preprocess/main.py`, `tasks/finalpool/update-material-inventory/preprocess/sheets_setup.py` | Capture real spreadsheet ID, write `config.json` into agent workspace, and add Sheets helper. | Accepted and applied; Sheets evaluation now has the missing `GoogleSheetsClient`, preprocess stores the copied spreadsheet ID in `spreadsheet_id`, and generated config is also emitted to the agent workspace. | | [x] | `latex-prompt-box` | `18e57f4f` | `tasks/finalpool/latex-prompt-box/evaluation/main.py` | Accept `ff9100` and normalize `\textless/\textbar/\textgreater` to literal `<>|`. | Accepted and applied with stricter color semantics; M-STAR's source redefines `proxYellow` from `ffbb00` to `ff9100`, so the evaluator now requires `ff9100` for `lightProxYellow` directly or through `colorlet` source tracing, and it treats LaTeX text commands for `<`, `>`, and `|` as prompt-token equivalents. | | [x] | `sync-todo-to-readme` | `a1914277` | `tasks/finalpool/sync-todo-to-readme/docs/task.md`, `tasks/finalpool/sync-todo-to-readme/docs/task_cn.md` | Clarify README baseline and main-to-dev TODO diff scope. | Reviewed and rejected; this task was already fixed by making the README ground truth match a full scan of TODOs in the dev workspace, while the PR54 wording narrows the instruction to applying only main-to-dev TODO diffs from the existing README baseline and would miss unrelated TODO entries that the evaluator expects. | | [x] | `music-analysis` | `4b0796fd` | `tasks/finalpool/music-analysis/initial_workspace/music_analysis_result_example.xlsx` | Blank out `Weeks 1-1` for streak-zero example row to match ground truth. | Accepted and applied; the example workbook now leaves `Top 3 Week Range` blank when `Longest Consecutive Top 3 Weeks` is `0`, while preserving the `Longest Consecutive Top 1 Weeks` value. | | [x] | `identify-all-songs` | `b6724574` | `tasks/finalpool/identify-all-songs/evaluation/check_content.py` | Accept `Nothin' on You` and `Nothing On You` as aliases. | Reviewed and rejected; the current task has already been aligned to the Jin Lyrics `NAys76UOlpI` 17-song version, whose ground truth does not include `Nothing On You`, so this alias only applies to the obsolete 20-song version and would add stale evaluator behavior. | | [x] | `oil-price` | `22c61407` | `tasks/finalpool/oil-price/evaluation/main.py`, `tasks/finalpool/oil-price/initial_workspace/detail.md` | Make max drawdown sign-agnostic, tolerate cost regex variants, and clarify YYYY-MM. | Partially accepted and applied; kept strict Max Drawdown sign semantics by documenting `Max Drawdown %` as a non-negative percentage magnitude, added explicit `YYYY-MM` month-label guidance, documented that the `Cost Assumption` percentage should use two decimal places, and relaxed the evaluator to require only the two-decimal `0.40` rate with an optional percent sign in that field. | | [x] | `notion-hr` | `5fc224ba` | `tasks/finalpool/notion-hr/evaluation/main.py` | Compare Highest Degree case-insensitively. | Accepted and applied; `Highest Degree` is a normalized category value (`master`/`bachelor`) and now follows the same case-insensitive comparison style already used for candidate names, positions, and schools, without changing the required degree category. | | [x] | `k8s-mysql` GT | `4f077c26` | `tasks/finalpool/k8s-mysql/groundtruth_workspace/gtq2.csv` | Regenerate `gtq2.csv` from natural-reading SQL, changing driver count from 3 to 104. | Accepted and applied; recomputing Q2 from the bundled F1 CSVs by `(driverId, year)` over 1950-1959 seasons yields the same 104 sorted driver IDs, while the previous `{501, 554, 579}` answer was only a valid subset. | | [x] | `notion` utility | `b5adff73` | `utils/app_specific/notion/ops.py` | Paginate `get_page_content_as_text` with `has_more` / `start_cursor`. | Accepted and applied with a small guard for missing `next_cursor`; Notion block children are paginated at 100 results, so the previous single-request helper could truncate long pages and miss tail sections. | | [x] | `wandb-shortest-length` | `9018ae1e` | `tasks/finalpool/wandb-shortest-length/evaluation/main.py` | Accept either strict multiples of 100 or inclusive final checkpoint. | Reviewed and rejected as already covered; the current evaluator already accepts both the full GT with final step `499` and the strict every-100-steps form ending at `400`, so no PR54 code change is needed. | | [x] | `detect-revised-terms` | `e6ba91b7` | `tasks/finalpool/detect-revised-terms/evaluation/check_content.py`, `tasks/finalpool/detect-revised-terms/groundtruth_workspace/revised_terms.csv` | Add semantic citation matching for punctuation, parentheses, Chinese digits, article multisets, and substring direction. | Reviewed and rejected; the current task prompt was already clarified to require quoted/reproduced original text, complete corresponding new-law provisions, and separate rows when one original clause maps to multiple new-law clauses, while the PR54 GT/evaluator changes would loosen or collapse parts of that clarified requirement. | | [x] | `email-paper-homepage` | `046553d1` | `tasks/finalpool/email-paper-homepage/docs/task.md`, `tasks/finalpool/email-paper-homepage/evaluation/main.py` | Broaden scope to all accepted/published papers and make URL-shape grading more robust. | Reviewed and rejected; the current task was already intentionally realigned to update code open-sourcing for newly accepted papers only, with evaluator expectations matching that scope (`Enhancing` and workshop no codeurl, `Ipsum` required, `LLM Adaptive` checked, `Optimizing` optional-but-correct if present), while PR54 expands the prompt and file checks back to all accepted/published papers. | | [x] | `vlm-history-completer` | `30f1d57c` | `tasks/finalpool/vlm-history-completer/groundtruth_workspace/groundtruth.json` | Expand SD-1.5 ground truth with canonical aliases. | Reviewed and rejected; the actual diff only removes the restored official Meta Make-a-Scene source URL rather than expanding SD-1.5 aliases, conflicting with the earlier accepted fix that requires that source. | | [x] | `nvidia-stock-analysis` | `511beb69` | `tasks/finalpool/nvidia-stock-analysis/evaluation/main.py` | Add cached yfinance share-count snapshot to stabilize revised historical share counts. | Reviewed and rejected as already covered; the current evaluator already contains the 2026-06-12 Basic Trend snapshot from the earlier alignment fix, while the PR54 file would remove that snapshot/live-fallback structure and loosen the top-holder comparison. | | [x] | `privacy-desensitization` | `d13d85b3` | `tasks/finalpool/privacy-desensitization/docs/task.md`, `tasks/finalpool/privacy-desensitization/evaluation/main.py` | Replace exact redaction equality with F1 span grading and clarify vanity phone/debit/bank scope. | Reviewed and rejected; this task was already corrected by reverting the F1 span grader, fixing archive/metadata handling, and clarifying IP addresses with directly attached ports, while PR54 reintroduces the reverted F1 grader and removes that accepted IP-port prompt clarification. | | [x] | `merge-hf-datasets` | `69105207` | `tasks/finalpool/merge-hf-datasets/evaluation/main.py` | Add JSON-Schema type aliases, semantic tool-message comparison, and `required: null` tolerance. | Reviewed and rejected as already covered; the current evaluator already has canonical type aliases, semantic tool-message content comparison, and `required: null` tolerance, while the PR54 file would remove those accepted relaxations. | | [x] | Notion MCP schema data | `f58b59be` | `configs/notion-mcp-patches/notion-openapi.json`, `scripts/run_single_containerized.sh`, `scripts/run_single_decoupled.sh` | Widen Notion schema `parent.oneOf` and allow additional page properties for database row mutations. | Accepted and applied with sandbox-style runtime wiring; the patched OpenAPI file is now tracked under `configs/` and both current container launch scripts bind-mount it read-only over the image's `node_modules/@notionhq/notion-mcp-server/scripts/notion-openapi.json`, so the pinned Notion MCP v1.9.0 schema exposes database row create/update payloads without rebuilding the image. | # PR54 Refresh Review Tracker Source PR: https://github.com/hkust-nlp/Toolathlon/pull/54 Refresh range: `f58b59be..4a70aae0` after the initial PR54 review pass. Status legend: - `[ ]` not reviewed - `[~]` in review / needs discussion - `[x]` reviewed and accepted/rejected/no-op - `[!]` reviewed and needs follow-up fix | Status | Area | PR commit(s) | Changed files | Claimed fix | Progress notes | | --- | --- | --- | --- | --- | --- | | [x] | `oil-price` | `3ac2c3ea` | `tasks/finalpool/oil-price/evaluation/main.py` | Accept either `12/N` or `12/(N-1)` annualization convention. | Partially accepted and applied per user decision: evaluator now computes an alternate annualized return using the `12/(N-1)` convention and accepts either annualization value for `Annualized Return %`. Rejected the unrelated PR loosening that removed Summary derived-field checks (`spread`, `spread_mom_pct`, `z_score`, `regime`, `signal`) or changed other backtest checks. | | [x] | `nhl-b2b-analysis` | `f44b22a7` | `tasks/finalpool/nhl-b2b-analysis/evaluation/main.py` | Move spreadsheet discovery inside the retry loop. | Reviewed and rejected as a PR regression; the current branch already resolves `find_spreadsheet_in_folder(...)` inside `_check()`, so each `grade_with_retry` attempt re-lists Drive after propagation lag, while the refreshed PR hoists the lookup outside the retry loop. | | [x] | `live-transactions` | `c8d41ddc`, `5ae43766` | No net files in refreshed PR diff. | Parenthesize log-filter OR and compare `related_transactions` order-independently, then revert it. | No-op in refreshed PR; the later revert removes the earlier change from the current PR head. | | [x] | `inventory-sync` | `f47ed36e` | `tasks/finalpool/inventory-sync/evaluation/main.py` | Preserve config file across retry cleanup iterations. | Reviewed and rejected; the current branch's `finally` runs only after the whole `grade_with_retry(...)` call returns, so the WooCommerce config persists across retry attempts and is cleaned up only after success/final failure. The config is regenerated by preprocess each run from the initialized WooCommerce product mapping, so keeping the PR's failure-path leftover file is unnecessary and less clean. | | [x] | `language-school` | `ef964659` | `tasks/finalpool/language-school/evaluation/check_local.py` | Fix CMU deadline mismatch between grader and `gt_record`. | Reviewed and rejected; `gt_record.md` is an old record for this task, while the current prompt/evaluator were already updated for current-year requirements in `bc00275b`, including the CMU `2026-12-09` deadline. Do not roll the evaluator back to the stale `Dec. 10, 2025` record. | | [x] | `machine-operating` | `096a2da7` | `tasks/finalpool/machine-operating/evaluation/main.py` | Add Layer-2 retry around GCS fetch and drop redundant log-file check. | Accepted and applied; GCS bucket/file lookup plus download now runs under `grade_with_retry` to absorb upload propagation lag, and the unrelated `res_log_file` JSON/messages shape check was removed so grading focuses on the uploaded anomaly report content. | | [x] | `search-ca-school` | `53ef4bcf` | `tasks/finalpool/search-ca-school/docs/task.md`, `tasks/finalpool/search-ca-school/evaluation/main.py`, `tasks/finalpool/search-ca-school/groundtruth_workspace/*` | Use static 2016-2026 GT and drop CSRankings rank from agent schema. | Reviewed and rejected; PR rewrites the task from 2024 live CSRankings with required `cs_ranking_rank` to a static 2016-2026 GT without rank. A fresh recomputation from CSRankings source CSVs using both the current grader's AI+Vision+ML+NLP ex-IR口径 and a narrower `ai`-only口径 keeps Arizona State University in the US top30 near-LA set and does not include UC Riverside (UCR ranks about 57 under the current composite口径 and about 90 under `ai`-only), so the PR's new static GT (`+UCR`, `-ASU`) is inconsistent with the task口径. | | [x] | `quantitative-financial-analysis` | `40838e1c` | `tasks/finalpool/quantitative-financial-analysis/groundtruth_workspace/groundtruth_data.csv` | Drop four spurious 2025-07-31 GT rows. | Reviewed and rejected; PR deletes the four `2025-07-31` rows for AAPL/META/NVDA/TSLA because the idiomatic `yfinance.download(..., end='2025-07-31')` call treats `end` as exclusive and returns through `2025-07-30`. Current prompt asks for each trading day in June and July 2025, and `2025-07-31` is a valid trading day, so current branch intentionally keeps the July 31 rows from `9da97f37`, giving 42 trading dates × 4 tickers = 168 rows. | | [x] | `sync-todo-to-readme` | `a1914277`, `f039a9a5` | `tasks/finalpool/sync-todo-to-readme/docs/task.md`, `tasks/finalpool/sync-todo-to-readme/docs/task_cn.md`, `tasks/finalpool/sync-todo-to-readme/groundtruth_workspace/README.md` | Scope task to main→dev TODO diff and drop baseline `vllm_v_0_6_3` / stale entries from GT. | Reviewed and rejected; PR changes the task口径, not just GT. Under the current full-scan prompt, dev source has 240 TODOs and current GT exactly matches all 240, including TODOs that existed in code but were missing from the original README. Under PR's new diff-based prompt ("use existing README as baseline, only apply main→dev TODO additions/removals/line updates"), recomputing from `Toolathlon-Archive/LUFFY` gives 18 additions and 29 removals from the initial README, producing exactly PR's 215-entry GT; however that prompt rewrite is not accepted because the original task asks for complete dev TODO synchronization. | | [x] | `sla-timeout-monitor` | `168b8e00` | `tasks/finalpool/sla-timeout-monitor/evaluation/main.py` | Preserve intentional no-negative-poll design. | Reviewed and rejected; PR removes the current `check_email_not_sent` polling loop and returns to a single INBOX read for users who should not receive emails. Current branch added the negative polling in `7bf19293` to observe the same short Layer-2 window used for IMAP propagation lag (`3` attempts, `5s` gap) before passing a negative assertion; this can catch forbidden emails that were sent but not visible on the first IMAP read. PR's claim that waiting cannot help a "should be empty" check is only true for turning failures into passes, but it can prevent false passes, so current negative polling is kept. | | [x] | `detect-revised-terms` docs | `db8e6e2d` | `tasks/finalpool/detect-revised-terms/docs/task.md` | Keep sandbox docs version in PR. | Reviewed and rejected; refresh commit reverts the clarified prompt back to the older broad wording ("conflict/inconsistent/revised/repealed"). Current prompt was intentionally clarified in `d2707e93` to require substantive content differences, quoted/reproduced original legal text, complete corresponding new-law provisions, and separate rows for one-to-many mappings. The full PR head also still differs from current evaluator/GT due to the earlier rejected semantic-match/GT rewrite, so accepting the PR version would roll back the current exact normalized GT alignment. | | [x] | `imagenet` docs | `db8e6e2d` | `tasks/finalpool/imagenet/docs/task.md`, `tasks/finalpool/imagenet/docs/task_cn.md` | Keep sandbox docs version in PR. | Reviewed and rejected; PR reverts the prompt from the clarified `e0b01833` wording ("must exactly follow `format.tex`: same LaTeX structure, column names, column order, and parameter unit style") back to the older softer "example structure is `format.tex`" wording. Current evaluator still compares normalized `survey.tex` against the exact GT file, whose structure mirrors `format.tex`, so the stricter prompt is needed to align the task instructions with the grader. | | [x] | `notion-personal-website` | `db8e6e2d` | `tasks/finalpool/notion-personal-website/evaluation/check_remote.py` | Keep sandbox evaluator version in PR. | Narrowly accepted and applied; `check_remote.py` now imports the shared `utils.app_specific.notion.ops.get_page_content_as_text` instead of maintaining a duplicate local pagination helper. The shared helper is already paginated and uses the Notion retry wrapper, while the current guarded `has_more=True` / missing-`next_cursor` protection in `utils/app_specific/notion/ops.py` was kept rather than taking the PR's shared-util regression. Verified with `python3 -m py_compile`. | | [x] | `merge-hf-datasets` | `1479559a` | `tasks/finalpool/merge-hf-datasets/evaluation/main.py` | Restore grader with three tolerances missing from sandbox. | Reviewed as no-op; the current evaluator already matches the refreshed PR version and includes all three claimed tolerances: comma-joined/cross-language type canonicalization, `required: null` equivalence with absent `required`, and semantic comparison for tool-role message content. Verified `HEAD..origin/pr/54` has no diff for this task and `python3 -m py_compile` passes. | | [x] | `university-course-selection` | `fb2df5a3` | `tasks/finalpool/university-course-selection/evaluation/check_local.py` | Revert to sandbox version and drop strict column-set check. | Reviewed and rejected; PR removes the strict required-column/agent-vs-GT column-set checks and switches to comparing only the intersection of available required columns, while also reintroducing softer instructor-title, class-time week-range, and restriction-token comparisons. Current prompt explicitly asks to follow `reference_format.xlsx` and retain the original English headers, and the current evaluator's strict column check was previously accepted to enforce that shape. PR's intersection logic could let outputs with missing required columns pass as long as the remaining common columns match, so current strict column-set behavior is kept. | | [x] | `privacy-desensitization` | `fd400d91` | `tasks/finalpool/privacy-desensitization/preprocess/main.py`, `tasks/finalpool/privacy-desensitization/initial_workspace/files.tar.gz`, `tasks/finalpool/privacy-desensitization/groundtruth_workspace/gt_files.tar.gz` | Ship sandbox version in PR. | Rejected per user decision; PR restores the sandbox/F1 stack and simpler archive extraction/layout, while current branch intentionally keeps the reverted exact GT comparison, safer archive handling, macOS metadata filtering, top-level `files/` flattening, and accepted IP-with-attached-port prompt clarification. | | [x] | `nvidia-stock-analysis` | `fd400d91` | `tasks/finalpool/nvidia-stock-analysis/initial_workspace/data.txt`, `tasks/finalpool/nvidia-stock-analysis/initial_workspace/tips.txt`, `tasks/finalpool/nvidia-stock-analysis/task_config.json` | Ship sandbox version in PR. | Rejected per user decision; PR shrinks Sheet 3 instructions by dropping Yahoo/yfinance source wording, all-available-if-fewer-than-10 behavior, descending sort requirement, and `percentage change`, adds a NaN hint relative to current branch, and removes `web_search` / `playwright_with_chunk`. Full PR branch still also differs in `evaluation/main.py` via earlier rejected `511beb69`, which would remove the saved 2026-06-12 Basic Trend snapshot and loosen top-holder matching. Keep current accepted `d8670bed` evaluator plus current task instructions/tooling. | | [x] | `latex-prompt-box` | `f67d3b28`, `4a70aae0` | `tasks/finalpool/latex-prompt-box/evaluation/main.py`, `tasks/finalpool/latex-prompt-box/evaluation/complex_prompt.txt` | Replace with sandbox full grader and align runtime prompt escaping before `boxed`. | Partially accepted and applied per user decision; reference `complex_prompt.txt` now uses one `\textbackslash` before `boxed` to match the runtime `qwen-boxed` prompt, while the evaluator accepts both one- and two-`\textbackslash` renderings for this LaTeX display escape ambiguity. Rejected the refreshed PR's broader color changes for now; current branch keeps the accepted final-effective M-STAR `ff9100` color check from `e1576f98`. | | [x] | `email-paper-homepage` | `459cea61`, `ad109b32` | `tasks/finalpool/email-paper-homepage/docs/task.md`, `tasks/finalpool/email-paper-homepage/evaluation/main.py` | PR ultimately ships broader original scope and strict grader. | Partially accepted and applied per user decision; rejected the broader-scope prompt/grader rollback and kept current `0bb7ca2a` task口径 (`newly accepted papers only`, Enhancing/workshop no codeurl, Ipsum required, LLM Adaptive checked, Optimizing optional-but-correct if present). Accepted only the small `check_modified_files` robustness fix: handle `modified_files is None` and extract filenames from either GitHub compare API dicts or object-like records. | | [x] | `student-interview` | `98eaeacf`, `45b7a203` | `tasks/finalpool/student-interview/evaluation/main.py` | Convert event time to `+08:00` before working-hours check. | Accepted and applied per user decision; evaluator now defines `TASK_TZ=+08:00`, normalizes returned Google Calendar datetimes before reading task-local dates/working hours, and evaluates the hardcoded conflict windows (`15:00-17:00`, `09:00-11:00`) in the same task timezone so UTC-normalized API responses do not false-fail valid Hong Kong-time interviews or miss local conflicts. | | [x] | `notion-movies` | `c99925e7` | `tasks/finalpool/notion-movies/docs/task.md`, `tasks/finalpool/notion-movies/docs/user_system_prompt.md`, `tasks/finalpool/notion-movies/evaluation/check_remote.py` | Ship sandbox permissive grader. | Accepted and applied per user decision; PR removes the current prompt/user-system requirement that the Star Wars trailer be embedded/follow existing subpage format and changes the grader from requiring an actual Notion `embed`/`video` block to accepting the correct YouTube video ID in a Trailer/YouTube property or page content. The fixed Notion MCP v1.9.0 schema and our mounted patch both expose only paragraph/bulleted-list child creation with rich-text links for `Append block children`, not `embed`, `video`, or `bookmark`, so the permissive check better matches the available tool surface. | | [x] | `canvas-arrange-exam` | `8dc10e05` | `tasks/finalpool/canvas-arrange-exam/evaluation/check_local.py`, `tasks/finalpool/canvas-arrange-exam/groundtruth_workspace/exam_schedule.xlsx` | Use full instructor names in GT, token-subset proctor compare, and mark CS301 closed-book. | Accepted and applied per user decision, with follow-up alignment from local review: prompt now explicitly asks for full proctor names, the CS301 announcement identifies the replacement proctor as `Black Smith`, and GT uses `Black Smith` plus `Closed-book`. Evidence matches the rationale: Canvas users expose full instructor names for default instructors, and CS301's generated welcome announcement publishes `Exam Type: Closed Book` before the January proctor replacement announcement. Evaluator keeps token-subset proctor matching without the duplicate import/misleading PR comment. | # PR54 Second Refresh Review Tracker Source PR: https://github.com/hkust-nlp/Toolathlon/pull/54 Refresh range: `4a70aae0..bd8609b7` after the PR54 refresh review pass. Status legend: - `[ ]` not reviewed - `[~]` in review / needs discussion - `[x]` reviewed and accepted/rejected/no-op - `[!]` reviewed and needs follow-up fix | Status | Area | PR commit(s) | Changed files | Claimed fix | Progress notes | | --- | --- | --- | --- | --- | --- | | [x] | `vlm-history-completer` | `8276b934` | `tasks/finalpool/vlm-history-completer/groundtruth_workspace/groundtruth.json` | Revert SD-1.5 canonical aliases. | No-op / accepted as PR-head alignment; the commit restores the Make-a-Scene Meta source URL that current branch already keeps, so `HEAD..origin/pr/54` has no diff for this task. | | [x] | `university-course-selection` | `2985bdc0` | `tasks/finalpool/university-course-selection/evaluation/check_local.py` | Combine strict column checks with lenient field comparisons. | Accepted in full and applied per user decision: replaced the branch's grader with the PR-head version. Compared to the prior branch grader (plain `normalize_str` equality on every column + exact column set, any extra/missing column fails, any blank required cell drops the whole row), the PR version ADDS: instructor-title stripping (`陈彤兵 讲师`≈`陈彤兵`, `龚金平 教授(教学为主型)`≈`龚金平`), order-independent restriction multiset compare with `级`-suffix and program/quality-tag (`思政B`, `上海市精品课程团队`/`精品课程`) stripping, tolerance of extra columns (only required-column *presence* enforced), Course-ID-only row dropping, and Course-ID-keyed row sorting. Blank restriction cell still fails (reference requires `无`) and Class Time stays strict (`[1-16]` not stripped). Supersedes the branch's `871090dc Require course selection columns to match` exact-column stance. prompt / GT / initial_workspace are byte-identical across branches, so the decision was purely grader leniency. | | [ ] | `identify-all-songs` | `2ee41422` | `tasks/finalpool/identify-all-songs/evaluation/check_content.py` | Revert `Nothin' on You` / `Nothing On You` alias tolerance. | Pending review. | | [x] | `oil-price` | `5bbd6215` | `tasks/finalpool/oil-price/evaluation/main.py`, `tasks/finalpool/oil-price/initial_workspace/detail.md` | Port T's grader while keeping annualized-return leniency. | Accepted in full and applied per user decision. The branch had already absorbed the substantive "port T's grader" content via earlier `85f20239`/`03bd332d`, so the only remaining diff vs PR head was `evaluation/main.py` (detail.md / prompt / GT already identical). That remaining diff is a behavior-neutral refactor of the Annualized-Return comparison block: `cmp_pairs` restructured to carry the expected alternatives as a list (instead of an in-loop `if name ==` special case), expanded N=12 (12/N) vs N=11 (12/(N-1)) convention docs, and a `0`→`0.0` literal. Verified equivalent: same 5 metrics, same tolerances (0.05/0.05/0.05/0.01/0.05), same accept-either-annualization-convention dedup+min-delta logic. No grading-behavior change. | | [ ] | `merge-hf-datasets` | `15d043c2`, `33e11edc` | No net files in refreshed PR diff. | Revert earlier merge-hf grader tolerance commits. | Pending review / likely no-op; verify final head against current branch. | | [ ] | `privacy-desensitization` | `4db21313` | `tasks/finalpool/privacy-desensitization/docs/task.md`, `evaluation/main.py`, `preprocess/main.py`, `initial_workspace/files.tar.gz`, `groundtruth_workspace/gt_files.tar.gz` | Port T's privacy-desensitization version. | Pending review. | | [ ] | `nhl-b2b-analysis` | `cc48d732` | `tasks/finalpool/nhl-b2b-analysis/evaluation/main.py` | Revert moving spreadsheet discovery inside retry loop. | Pending review. | | [ ] | `inventory-sync` | `5c345540` | `tasks/finalpool/inventory-sync/evaluation/main.py` | Use `try/finally` cleanup after `grade_with_retry`. | Pending review. | | [ ] | `language-school` | `8ddf286e` | `tasks/finalpool/language-school/evaluation/check_local.py` | Revert CMU deadline mismatch fix. | Pending review. | | [x] | `quantitative-financial-analysis` | `7ea89fe4`, `dec71525` | `tasks/finalpool/quantitative-financial-analysis/evaluation/main.py`, `groundtruth_workspace/groundtruth_data.csv` | Restore July 31 GT rows and increase Sheets retry budget. | Accepted and applied per user decision. The July-31 GT rows (`7ea89fe4`) were already on the branch via `9da97f37`, so the csv matched and only `evaluation/main.py` (`dec71525`) differed. That diff is a pure evaluation-retry robustness change: `grade_with_retry` `max_attempts` 4→12 and added `poll_s=10` (≈11×10s ≈ 110s, to span Google Sheets' 60s/min read-quota window and stop a single 429 burst from false-failing the check), plus an explanatory comment. No grading-criteria change. Verified the branch's `grade_with_retry` already accepts `poll_s`/`max_attempts`, so no breakage. | | [ ] | `sla-timeout-monitor` | `4c17de88` | `tasks/finalpool/sla-timeout-monitor/evaluation/main.py` | Revert no-negative-poll design preservation. | Pending review. | | [x] | `latex-prompt-box` | `486ce189`, `bd8609b7` | `tasks/finalpool/latex-prompt-box/evaluation/main.py` (+ orphan `evaluation/complex_prompt.txt`) | Restore strict color with dual-`\textbackslash` acceptance and remove stray GT prompt escape. | Accepted and applied per user decision. `main.py` is a GT correctness fix: removes the stray `\\` between `\boxed\{\}.` and `<|im\_end|>`. Verified against the real source prompt (qwen-boxed template inside `initial_workspace/codes.tar.gz` → `qwen_math_eval_toolkit/utils.py:136`): `\boxed{}.` is immediately followed by `<|im_end|>` with NO newline, so per the task's "use `\\` for new lines" rule there must be no `\\` there — the branch's prior GT wrongly required it and would false-fail faithful agents (and pass agents that inserted a spurious line break). Both versions still accept one-or-two `\textbackslash` before `boxed` (runtime `\boxed`=1 vs source `\\boxed`=2 readings). Also aligned the orphan `evaluation/complex_prompt.txt` (trailing-newline only; not referenced by the grader, not part of commits 486ce189/bd8609b7) so the task dir is byte-identical to PR head. | | [x] | `woocommerce-customer-survey` | `ac3c9ebd`, `011093a3` | `tasks/finalpool/woocommerce-customer-survey/initial_workspace/form_requirement.md` (grader `evaluation/main.py` already aligned on branch) | Reconcile delivery question wording between `our` and `the`. | Accepted and applied per user decision. spec/grader consistency fix: the branch grader (`main.py:581`) already expected "Are you satisfied with **the** delivery service?", but the spec (`form_requirement.md`) still said "**our**", so a faithful agent would write "our" and false-fail. Adopted PR's spec wording ("the"). Only `form_requirement.md` needed applying — the grader main.py was already identical between branches (the `ac3c9ebd`→`011093a3` the→our→the churn netted to "the", which the branch grader already had). | | [x] | `reimbursement-form-filler` | `5570b637` | `tasks/finalpool/reimbursement-form-filler/evaluation/check_local.py` | Bridge `datetime` and string values for month cells. | Accepted and applied per user decision. grader correctness fix: adds a `_to_year_month` bridge in `compare_element` so a `datetime` month cell (the `Bill_Format.xlsx` template pre-fills [5,0]/[6,0]/[7,0] as `datetime(2025,4/5/6,1)`, which a faithful "don't change the layout" agent keeps) matches the GT's `"2025-MM"` strings instead of failing on type alone. Over-leniency risk (the YYYY-MM normalization drops the day) verified NOT applicable: every date cell in GT `department_expenses.xlsx` is month-level (`2025-04/05/06`), no day-level dates exist. Bridge only fires when both sides parse as a month-date; non-date cells fall through to the original compare. | # PR54 Branch-vs-PR Task Reconciliation Tracker Tracks the tasks under `tasks/` that differ between our verified branch and PR #54, so every difference is a recorded decision (adopt PR / keep ours / reject PR) rather than an accidental divergence. - **Our branch:** `toolathlon-tasks-verified-toreview` - **PR #54 branch:** `claude/task-fixes-2026-06-24` (head `bd8609b7`, a.k.a. `origin/pr/54`) - **Snapshot:** 2026-06-25. Originally 14 differing tasks under `tasks/`. This section complements the PR54 per-commit review trackers above (PR54 / PR54 Refresh / PR54 Second Refresh); where a decision here reverses one of those, this section is the authoritative record. ## Decision legend - `[PR]` adopt PR's version (our branch made equal to PR) - `[OURS]` keep our version (PR's change not adopted, but we changed the files ourselves) - `[REJECT]` reject PR's change (we keep the pre-PR version on purpose) `你改 / PR改` = did each side change the task relative to the merge-base `acb8e54a`. ## Tasks | # | Decision | Task | 你改 | PR改 | Differing files (ours↔PR) | Notes | |---|---|---|---|---|---|---| | 1 | `[PR]` ✅ applied | `imagenet` | no | yes | `docs/task.md`, `docs/task_cn.md` | **Adopt PR per user decision (2026-06-25) — reverses the earlier reject in `UpdateLogs` (rows 49/152/217).** PR reverts the prompt from the strict "must exactly follow `format.tex`" wording back to the softer "example structure". Now aligned to PR. | | 2 | `[PR]` ✅ applied | `sync-todo-to-readme` | no | yes | `docs/task.md`, `docs/task_cn.md`, `groundtruth_workspace/README.md` | **Adopt PR per user decision (2026-06-25) — reverses the earlier reject in `UpdateLogs` (rows 56/111/149).** PR narrows the task to a main→dev TODO diff (215-entry README GT) instead of the full-scan口径 (240). Now aligned to PR. | | 3 | `[PR]` ✅ applied | `canvas-arrange-exam` | yes | yes | `docs/task.md`, `evaluation/check_local.py`, `files/course_config.json`, `groundtruth_workspace/exam_schedule.xlsx` | **Adopt PR per user decision (2026-06-25) — reverses our earlier "Black Smith" choice (`a95e67d7` / `2d7f0cdd`).** PR keeps the CS301 replacement proctor as "Professor Smith" (no fabricated first name): announcement says "Professor Smith", GT Proctor cell `[7,2]` = "Smith", and the prompt drops the "fill in full names" requirement. grader logic (token-subset compare, strip "Professor") is unchanged between versions. Now aligned to PR. | | 4 | `[PR]` ✅ applied | `canvas-do-quiz` | yes | yes | `evaluation/check_remote.py` | **Adopt PR per user decision (2026-06-25).** Comment-only difference — both versions share the identical `kept_score`-over-`score` grading logic (our `67c44db9`); PR just has a more detailed comment (DB101 attempt-2 `untaken`/`score=null` example). No behavioral change. Now aligned to PR. | | 5 | `[PR]` ✅ applied | `email-paper-homepage` | yes | yes | `docs/task.md`, `evaluation/main.py` | **Adopt PR per user decision (2026-06-25) — reverses the earlier reject `d170d91b`.** PR broadens the task scope from "newly-accepted papers only" to "all accepted/published papers": Enhancing LLMs now requires its codeurl, Optimizing LLMs becomes `to_be_released` (no codeurl), the Workshop check is dropped, and the allowed-modify file set expands. Now aligned to PR. | | 6 | `[PR]` ✅ applied | `k8s-mysql` | yes | yes | `groundtruth_workspace/gtq2.csv` | **Adopt PR per user decision (2026-06-25).** Trailing-newline-only difference — driver_id data (through 789) is row-identical; grader reads via `csv.DictReader` which ignores the trailing newline, so no behavioral change. Now aligned to PR. | | 7 | `[PR]` ✅ applied | `notion-hr` | yes | yes | `evaluation/main.py` | **Adopt PR per user decision (2026-06-25).** Whitespace-only difference (trailing whitespace on a blank line + final newline); the degree/school case-insensitive comparison (our `f9eaf3b8`) is present in both. No behavioral change. Now aligned to PR. | | 8 | `[PR]` ✅ applied | `notion-personal-website` | yes | yes | `evaluation/check_remote.py` | **Adopt PR per user decision (2026-06-25).** Only difference is a dead/unused `import requests` (used in neither version) + whitespace; our pagination/shared-helper changes (`a1c9c087`/`6352a6dd`) are present in both. Equivalent (ours was marginally cleaner without the dead import). Now aligned to PR. | | 9 | `[OURS]` | `student-interview` | yes | yes | `evaluation/main.py` | **Keep ours (timezone-robust).** Both versions run in +08:00, but Google Calendar returns `dateTime` normalized to UTC (`Z`) on read — documented by PR author jxhe in the `iso_times_equal` docstring (commit `65bcd78c`, 2026-06-03) and the very reason that function exists. Our version converts every event to +08:00 before checking working hours AND meeting conflicts, so it is correct whether the API returns +08:00 or UTC. PR's Check 4 stamps the +08:00 conflict hours (e.g. 15:00–17:00 Academic Committee Meeting) with the event's returned tz (UTC) **without converting**, placing the conflict window 8h off and missing real conflicts — contradicting jxhe's own docstring. User deferred final sign-off (2026-06-26); leaning keep-ours. | | 10 | `[OURS]` | `woocommerce-product-recall` | yes | no | `evaluation/check_remote_recall.py`, `initial_workspace/recall_form_template.json` | **Keep ours.** PR did not touch this task; the entire diff is our `943be37d` simplifying the recall form to the `google_forms` MCP's capabilities — dropping `form_description`, per-field `description`, the `date`/`paragraph`/`checkbox` types and the `settings` block, and removing the grader's description checks. "Use PR/base" would revert this and re-require form features the MCP likely can't create. The descriptions part is well-corroborated: both our branch (`943be37d` + sibling `f88a7af3` "Remove customer survey form description requirement") and jxhe's own PR (the customer-survey `011093a3` we adopted in this round, which removed the form Description line) independently dropped form descriptions. Caveat: whether the MCP also can't do `date`/`checkbox`/`settings` is not statically verified. Decision: keep ours. | | 11 | `[OURS]` | `detect-revised-terms` | yes | yes | `docs/task.md`, `evaluation/check_content.py`, `groundtruth_workspace/revised_terms.csv` | **Keep ours, with user refinements (2026-06-26).** Kept our strict exact-match grader (NOT PR's recall-only/broad rewrite). User hardened `normalize_legal_clause`: drop the `中华人民共和国` prefix (《中华人民共和国合同法》≡《合同法》), strip parenthetical sub-clause qualifiers ((第三款), full/half-width), and truncate at the first `条` so clause identity is compared at the article level (paragraph/item text still verified via the content columns); also added a case4 GT row (合同法第二百三十二条 → 民法典第七百三十条). PR-reject rationale unchanged: PR reverts the clarified "substantive-change + quoted original + complete new-law + one-to-many rows" prompt and swaps to a broader口径/GT (incl. carried-over-unchanged clauses). | | 12 | `[OURS]` | `nvidia-stock-analysis` | no | yes | `evaluation/main.py`, `initial_workspace/data.txt`, `initial_workspace/tips.txt`, `task_config.json` | **Keep ours per user decision (2026-06-26).** Ours is the reproducibility-hardened version: Sheet 1 uses a frozen 2026-06-12 yfinance snapshot (PR reverts to live fetch → flaky when Yahoo revises historical share-count data); Sheet 3 does a strict top-holder comparison — name-matched (norm: lowercase + strip punctuation + inc/corp/llc), descending-sort enforced, exact row count — whereas PR loosens it to positional compare where a name mismatch is only a `Warning`; ours also keeps the `web_search`/`playwright_with_chunk` tools and the explicit "from Yahoo/yfinance, sorted desc, incl. percentage change" Sheet 3 contract that PR drops. No file change needed — already on our version. | | 13 | `[OURS]` | `search-ca-school` | no | yes | `docs/task.md`, `evaluation/main.py`, `groundtruth_workspace/AI_univ_LA_500miles_Top30.json`, `…_2024.json` | **Keep ours per user decision (2026-06-26; user re-ran the task, no issues).** Ours grades by live-CSRankings recompute (fetch `generated-author-info.csv` + `institutions.csv`, recompute AI-ex-IR US top-30 with CSRankings' own scoring) so agent and grader read the same source; PR replaces this with a static GT. A fresh recompute (same algorithm, top-40) for single years 2022 and 2023 confirms UC Riverside is NOT in the top-30 (2022 rank 39; 2023 not in top-40) while ASU is (2022 rank 25), so PR's static GT (`+UCR`/`-ASU`, no `cs_ranking_rank`, 2016–2026 window, dropped "name = CSRankings exact" rule) is inconsistent with the task口径. No file change needed — already on our version. | | 14 | `[PR]` ✅ applied | `wandb-shortest-length` | no | yes | `evaluation/main.py` | **Adopt PR per user decision (2026-06-26).** Behaviorally equivalent: both versions accept either the full GT (final step `499`) or the strict every-100-steps form (ending `400`) via the identical `n_ag == n_gt` / `n_ag == n_gt - 1` logic. PR only adds an explanatory comment and refactors the compare loop (drops the `compare_rows=0` init and the `.equals()` fast-path), with the same cell tolerance — no grading change. Reverses the earlier "already covered" reject; now byte-aligned to PR. | ## Summary - **9** adopt PR (`imagenet`, `sync-todo-to-readme`, `email-paper-homepage`, `canvas-arrange-exam`, `canvas-do-quiz`, `k8s-mysql`, `notion-hr`, `notion-personal-website`, `wandb-shortest-length`) — applied 2026-06-25/26, now aligned to PR. - **5** keep ours (own/refined version differs from PR): `student-interview`, `woocommerce-product-recall`, `detect-revised-terms`, `nvidia-stock-analysis`, `search-ca-school`. - **0** set aside — all 14 tasks decided. - After the 9 adoptions, **5** tasks under `tasks/` still differ from PR (all intentional). - Non-task files that also differ (out of scope here): `UpdateLogs_CommonIssues.md`, `scripts/run_single_containerized.sh`, `scripts/run_single_decoupled.sh`, `utils/app_specific/notion/ops.py`. # PR38 Task Review Tracker Source PR: https://github.com/hkust-nlp/Toolathlon/pull/38 Status legend: - `[ ]` not reviewed - `[~]` in review / needs discussion - `[x]` reviewed and accepted - `[!]` reviewed and needs follow-up fix Task scope is taken from PR #38, but each task review compares the PR #38 version against the current branch `toolathlon-tasks-verified-toreview`. Change-size categories are only an initial reference based on PR #38's own non-`SUGGESTED_CHANGES.md` churn; final decisions should be based on the current-branch comparison. | Status | Category | Task | Non-note churn | Changed files | Notes | | --- | --- | --- | ---: | ---: | --- | | [x] | medium | apply-phd-email | 59 | 3 | Accepted exact-subject latest-email selection; current branch already clears receiver INBOX; rejected retry removal and preprocess rewrite. | | [x] | tiny | canvas-arrange-exam | 5 | 3 | Accepted only preprocess exit(0) removal; prompt email mention rejected; do not restore PR38 memory.json. | | [x] | medium | email-paper-homepage | 83 | 3 | Rejected PR38 dynamic 2026 year; aligned prompt to newly accepted papers only; Enhancing and Workshop must not have codeurl, Ipsum must, LLM Adaptive remains checked, Optimizing codeurl is optional but must be correct if present. | | [x] | large | game-statistics | 194 | 3 | Accepted date alignment fix: preprocess/evaluation use launch_time task date, generated timestamps stay within task date, and historical stats are based on that date. | | [x] | medium | identify-all-songs | 93 | 2 | Accepted Jin Lyrics/NAys76UOlpI 17-song version only; removed old 20-song groundtruth and aligned prompts. | | [x] | large | imagenet | 264 | 2 | Rejected parser relaxation; clarified prompt to require exact format.tex structure, columns, order, and parameter unit style. | | [x] | light | latex-prompt-box | 20 | 2 | Accepted evaluator support for direct ffbb00 color or colorlet alias to a defined ffbb00 source color. | | [x] | large | llm-training-dataset | 232 | 2 | Kept retry; clarified shared datasets must be model-specific rows; evaluator now uses missing dataset names, model-specific sizes, and accepts GPT-Neo via The Pile, The Pile+22, or 22 components. | | [x] | medium | merge-hf-datasets | 115 | 2 | Rejected PR38 groundtruth/schema-wrapper removal; accepted narrow evaluator relaxations for required:null and tool-role JSON content semantics. | | [x] | tiny | notion-find-job | 3 | 2 | Accepted prompt salary boundary >=3000 after AHC confirmed 3000-3200; rejected PR38 evaluation/main.py retry removal. | | [x] | medium | notion-personal-website | 67 | 3 | Kept existing retry and accepted Notion block pagination; prompt already includes Exhibitions. | | [x] | large | oil-price | 168 | 2 | Rejected PR38; current branch already has clearer adjacent-month backtest spec, summary checks, tolerances, and retry. | | [x] | light | privacy-desensitization | 11 | 2 | Accepted prompt clarification that IP addresses include directly attached ports; rejected PR38 evaluator normalization and tar/preprocess regressions. | | [x] | light | quantitative-financial-analysis | 17 | 3 | Accepted missing 2025-07-31 issue; added verified Yahoo Finance values for AAPL/META/NVDA/TSLA; rejected PR38 blank rows, skipped numeric checks, and retry removal. | | [x] | medium | stock-build-position | 53 | 2 | Accepted with conservative stock-code alias checks; excludes forms like 700 and 2594. | | [x] | light | student-interview | 20 | 2 | Kept current branch semantic ISO time comparison; accepted PR38 conflict-message strftime fix only. | | [x] | medium | university-course-selection | 92 | 2 | Accepted strict column-match fix only; rejected PR38 instructor/time/restriction normalization. | | [x] | medium | vlm-history-completer | 74 | 3 | Restored Make-a-Scene official Meta source URL; rejected PR38 evaluator rewrite and SD1.5 alias removal. | | [x] | light | woocommerce-customer-survey | 26 | 5 | Accepted docs filename typo fixes and form requirement question wording to our delivery service; rejected evaluator alternative question matching, retry removal, and preprocess seed removal. | | [x] | light | youtube-repo | 32 | 2 | Accepted evaluator alias for HazyResearch/flash-attention as equivalent to Dao-AILab/flash-attention. | ## Review Log - Pending.