{ "cells": [ { "cell_type": "markdown", "metadata": {}, "source": [ "# 🏏 Cricket with `sportsdataverse-py`\n", "\n", "**ESPN** carries live scorecards, standings, and full match summaries for the world's\n", "most widely-played bat-and-ball game. `sportsdataverse-py` wraps that surface through\n", "the `sportsdataverse.cricket` module β€” one `league=` slug away from IPL, England's\n", "county circuit, the ICC World Cup, and every other tournament ESPN indexes.\n", "\n", "### Cricket in 30 seconds\n", "\n", "If you're new to cricket, the key numbers to know:\n", "\n", "| Concept | What it means in the data |\n", "|---|---|\n", "| **Innings** | A team's turn to bat; T20 matches have one per team, Tests have two |\n", "| **Score string** | `\"161/5 (18/20 ov, target 156)\"` β€” runs / wickets (overs used / overs allowed, target) |\n", "| **Wickets** | Dismissals; ten wickets = all out, innings ends |\n", "| **Overs** | Six-ball delivery sets; T20 = 20 overs, ODI = 50, Tests = open |\n", "| **Partnership** | Runs scored by two batters sharing the crease |\n", "\n", "The score string (`home_score` / `away_score`) is returned verbatim from ESPN so\n", "downstream analysis retains the full cricket context rather than stripping it to\n", "a bare integer.\n", "\n", "### What this notebook covers\n", "\n", "1. Setup and the `safe()` guard helper\n", "2. Scoreboard β€” live and recent matches for a league\n", "3. Standings β€” the `children` hierarchy, flattened\n", "4. Match summary β€” all 8 matchcard sections, with emphasis on the three\n", " heterogeneous batting / bowling / partnerships scorecard shapes\n", "5. Bonus endpoints β€” news, injuries, calendar\n", "6. Caveats and the full reference\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 🧰 The toolbox\n", "\n", "Everything returns a tidy **polars** `DataFrame` by default β€” pass\n", "`return_as_pandas=True` for pandas, or `return_parsed=False` for the raw JSON `dict`.\n", "\n", "The three cricket-specific parsers are:\n", "\n", "| Parser | Input | Output |\n", "|---|---|---|\n", "| `parse_cricket_scoreboard` | `espn_cricket_scoreboard` payload | One row per match; score strings in cricket format |\n", "| `parse_cricket_standings` | `espn_cricket_standings` payload | One row per team per group; flattened `group` column |\n", "| `parse_cricket_summary` | `espn_cricket_summary` payload | Dict of 8 section DataFrames (or single section) |\n", "\n", "All other `espn_cricket_*` wrappers reuse the universal parsers (`parse_news`,\n", "`parse_items`, `parse_single_entity`, etc.) shared across all ESPN-backed sports.\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## πŸ”Œ Setup\n", "\n", "```sh\n", "pip install sportsdataverse\n", "```\n", "\n", "No API key required." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "import polars as pl\n", "import sportsdataverse.cricket as cricket\n", "\n", "# IPL (Indian Premier League) league slug β€” used throughout this notebook.\n", "# Other common slugs: 'eng.1' (England domestic), 'icc.worldcup' (ODI World Cup)\n", "IPL = \"8048\"\n", "\n", "print(\"polars\", pl.__version__)" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "The ESPN cricket feed is live and occasionally rate-limited, so a small `safe()`\n", "helper runs every network call defensively. Any exception is caught, printed, and\n", "`None` is returned β€” downstream cells check for `None` before proceeding." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "from sportsdataverse.errors import AssetFetchError, NoDataError\n", "\n", "def safe(label, thunk):\n", " try:\n", " out = thunk()\n", " print(f\"βœ… {label}\")\n", " return out\n", " except (NoDataError, AssetFetchError) as e:\n", " print(f\"\\u23ed\\ufe0f {label}: {type(e).__name__}: {e}\")\n", " return None" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## πŸ“‘ Scoreboard β€” today's (and recent) matches\n", "\n", "[`espn_cricket_scoreboard`](../cricket/reference/site.md#espn_cricket_scoreboard)\n", "hits the Site v2 scoreboard endpoint for the given `league=` slug.\n", "By default (``) it routes the payload through\n", "`parse_cricket_scoreboard` and returns a tidy polars frame.\n", "Pass `return_parsed=False` to get the raw ESPN JSON dict instead.\n", "\n", "Each row is one match. The `home_score` / `away_score` columns carry the full\n", "cricket score string β€” ESPN doesn't expose a clean integer run-count separately,\n", "and the wickets + overs context matters." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "board = safe(\n", " \"IPL scoreboard\",\n", " lambda: cricket.espn_cricket_scoreboard(league=IPL),\n", ")\n", "board" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### What the columns mean\n", "\n", "| Column | Description |\n", "|---|---|\n", "| `event_id` | ESPN event identifier β€” pass this to `espn_cricket_summary` |\n", "| `date` | ISO-8601 match start time |\n", "| `name` / `short_name` | Full and abbreviated match name |\n", "| `home_team` / `away_team` | Display names |\n", "| `home_score` / `away_score` | Cricket score string, e.g. `\"161/5 (18/20 ov, target 156)\"` |\n", "| `status` | `\"Final\"`, `\"In Progress\"`, `\"Scheduled\"`, etc. |\n", "| `status_detail` | Human-readable detail, e.g. `\"Chennai Super Kings won by 5 wickets\"` |\n", "| `venue` | Ground name |\n", "| `neutral_site` | Boolean β€” neutral ground match |" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# If the scoreboard returned data, show the match-status breakdown.\n", "if board is not None and board.height:\n", " keep = [c for c in [\"name\", \"home_score\", \"away_score\", \"status\", \"status_detail\"] if c in board.columns]\n", " print(board.select(keep))\n", "else:\n", " print(\"scoreboard unavailable right now β€” try again outside an off-season window\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Raw payload mode\n", "\n", "Pass `return_parsed=False` to skip the parser entirely and work with the raw\n", "ESPN JSON. This is useful when you want to explore the full payload structure,\n", "or when you need a field the parser doesn't yet surface." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "raw_board = safe(\n", " \"IPL scoreboard (raw)\",\n", " lambda: cricket.espn_cricket_scoreboard(league=IPL, return_parsed=False),\n", ")\n", "if isinstance(raw_board, dict):\n", " print(\"Top-level keys:\", list(raw_board.keys()))\n", " events = raw_board.get(\"events\") or []\n", " print(f\"Events in payload: {len(events)}\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## πŸ† Standings\n", "\n", "[`espn_cricket_standings`](../cricket/reference/site.md#espn_cricket_standings)\n", "returns the league table for a given season. The ESPN cricket standings payload\n", "uses a `children` hierarchy (groups/divisions) rather than the flat `groups` shape\n", "used in most other ESPN sports.\n", "\n", "`parse_cricket_standings` flattens that hierarchy β€” each row is one team in one\n", "group, with a `group` column so you can split multi-group tournaments (e.g. ICC\n", "World Cup group stages) with a single `.filter()` call.\n", "\n", "Optional parameters: `season=`, `group=`, `standings_type=`." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "standings = safe(\n", " \"IPL standings\",\n", " lambda: cricket.espn_cricket_standings(league=IPL),\n", ")\n", "standings" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "if standings is not None and standings.height:\n", " print(\"Columns:\", standings.columns)\n", " print(\"\\nGroups present:\", standings[\"group\"].unique().to_list() if \"group\" in standings.columns else \"(none)\")\n", " # Sort by points (or net run rate when available).\n", " sort_col = next((c for c in [\"points\", \"wins\"] if c in standings.columns), None)\n", " if sort_col:\n", " print(f\"\\nTop 4 by {sort_col}:\")\n", " print(\n", " standings\n", " .sort(sort_col, descending=True)\n", " .select([c for c in [\"team\", \"group\", \"wins\", \"losses\", \"points\", \"net_run_rate\"]\n", " if c in standings.columns])\n", " .head(4)\n", " )\n", "else:\n", " print(\"standings unavailable right now\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Why the `children` hierarchy matters\n", "\n", "Most ESPN standings payloads have a flat `standings.groups[]` block. Cricket uses\n", "`standings.children[]` instead β€” each child is a group/division with its own\n", "entries array. `parse_cricket_standings` walks that nesting and stitches a `group`\n", "column onto every team row, so the resulting frame is directly filterable:\n", "\n", "```python\n", "# Keep only Group A in a World Cup-style tournament\n", "standings.filter(pl.col(\"group\") == \"Group A\")\n", "```\n", "\n", "The numeric stat columns (wins, losses, points, net run rate, etc.) are snake-cased\n", "versions of whatever ESPN ships β€” they vary by tournament format." ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## πŸƒ Match summary β€” the full scorecard\n", "\n", "[`espn_cricket_summary`](../cricket/reference/site.md#espn_cricket_summary) is the\n", "richest endpoint: a single call returns **8 frames** that collectively make up the\n", "entire matchcard.\n", "\n", "| Section key | Content |\n", "|---|---|\n", "| `header` | Match metadata β€” teams, status, venue, toss result |\n", "| `matchcards_batting` | Per-batter innings rows (runs, balls, fours, sixes, strike rate) |\n", "| `matchcards_bowling` | Per-bowler innings rows (overs, maidens, runs, wickets, economy) |\n", "| `matchcards_partnerships` | Partnership pairs (runs, balls, each batter's contribution) |\n", "| `rosters` | Full squad lists for both teams |\n", "| `game_info` | Match-level metadata (series name, match type, result method) |\n", "| `leaders` | Stat leaders for the match |\n", "| `standings` | In-match standings snapshot |\n", "\n", "The three `matchcards_*` frames have **different schemas** β€” batting rows carry\n", "`runs`/`balls`/`strike_rate`; bowling rows carry `overs`/`wickets`/`economy`;\n", "partnership rows carry `total_runs`/`total_balls` plus per-batter run splits.\n", "They are returned as **separate frames** so callers can work with each schema cleanly.\n", "\n", "Event ID `1535465` is an IPL match (Chennai Super Kings vs. Mumbai Indians) that\n", "is used as the worked example throughout." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "EVENT_ID = 1535465 # IPL β€” Chennai Super Kings vs. Mumbai Indians\n", "\n", "summary_raw = safe(\n", " f\"match summary {EVENT_ID}\",\n", " lambda: cricket.espn_cricket_summary(league=IPL, event_id=EVENT_ID, return_parsed=False),\n", ")\n", "if isinstance(summary_raw, dict):\n", " print(\"Top-level keys:\", list(summary_raw.keys()))" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Parsing all 8 sections at once\n", "\n", "Call `parse_cricket_summary` with `section=None` (the default) to get a\n", "`dict[str, pl.DataFrame]` keyed by section name." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "from sportsdataverse.cricket.cricket_espn_parsers import parse_cricket_summary\n", "\n", "if summary_raw is not None:\n", " frames = parse_cricket_summary(summary_raw)\n", " for name, df in frames.items():\n", " print(f\"{name:30s} {df.shape[0]:>4d} rows Γ— {df.shape[1]:>3d} cols\")\n", "else:\n", " frames = {}\n", " print(\"summary payload unavailable β€” frames dict is empty\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### 🏏 Section: `matchcards_batting`\n", "\n", "One row per batter per innings. Each innings is identified by `innings_number`\n", "and `team`. The `summary` column carries the batter's dismissal description\n", "(e.g. `\"c Rohit b Bumrah\"`). Batters who haven't faced a ball yet appear with\n", "null run / ball counts." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "batting = frames.get(\"matchcards_batting\", pl.DataFrame())\n", "if batting.height:\n", " print(\"Batting columns:\", batting.columns)\n", " show_cols = [c for c in [\"innings_number\", \"team\", \"athlete_display_name\",\n", " \"runs\", \"balls\", \"fours\", \"sixes\", \"strike_rate\", \"summary\"]\n", " if c in batting.columns]\n", " print(batting.select(show_cols).head(10))\n", "else:\n", " print(\"batting scorecard unavailable for this match\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### 🎯 Section: `matchcards_bowling`\n", "\n", "One row per bowler per innings. Key columns: `overs`, `maidens`, `runs_conceded`,\n", "`wickets`, `economy`. Economy rate is the average runs conceded per over β€”\n", "lower is better." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "bowling = frames.get(\"matchcards_bowling\", pl.DataFrame())\n", "if bowling.height:\n", " print(\"Bowling columns:\", bowling.columns)\n", " show_cols = [c for c in [\"innings_number\", \"team\", \"athlete_display_name\",\n", " \"overs\", \"maidens\", \"runs_conceded\", \"wickets\", \"economy\"]\n", " if c in bowling.columns]\n", " print(bowling.select(show_cols).head(10))\n", "else:\n", " print(\"bowling scorecard unavailable for this match\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### 🀝 Section: `matchcards_partnerships`\n", "\n", "One row per partnership (pair of batters sharing the crease) per innings.\n", "This frame has a **different schema** from batting and bowling β€” it carries\n", "total runs and balls for the partnership, plus per-batter run splits.\n", "Partnership data is cricket-specific and has no direct analogue in ball-sport\n", "box scores." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "partnerships = frames.get(\"matchcards_partnerships\", pl.DataFrame())\n", "if partnerships.height:\n", " print(\"Partnerships columns:\", partnerships.columns)\n", " show_cols = [c for c in [\"innings_number\", \"team\", \"total_runs\", \"total_balls\",\n", " \"batter1_display_name\", \"batter1_runs\",\n", " \"batter2_display_name\", \"batter2_runs\"]\n", " if c in partnerships.columns]\n", " print(partnerships.select(show_cols).head(8))\n", "else:\n", " print(\"partnerships scorecard unavailable for this match\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### 🏟️ Section: `header`\n", "\n", "Match-level metadata: teams, status, venue, and the competition context. This\n", "is typically the first frame you'd use to confirm match identity and outcome." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "header = frames.get(\"header\", pl.DataFrame())\n", "if header.height:\n", " show_cols = [c for c in [\"name\", \"status_type_name\", \"status_type_detail\",\n", " \"home_team\", \"home_score\",\n", " \"away_team\", \"away_score\", \"venue_full_name\"]\n", " if c in header.columns]\n", " print(header.select(show_cols).head())\n", "else:\n", " print(\"header unavailable for this match\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### πŸ“‹ Section: `game_info`\n", "\n", "Match-level metadata that doesn't fit the header: toss winner, match type\n", "(T20, ODI, Test), series name, playing conditions, and the result method if\n", "weather interruption applied (D/L method)." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "game_info = frames.get(\"game_info\", pl.DataFrame())\n", "if game_info.height:\n", " print(game_info.head())\n", "else:\n", " print(\"game_info unavailable for this match\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### πŸ”Ž Requesting a single section\n", "\n", "When you only need one frame, pass `section=` to `parse_cricket_summary` to\n", "avoid deserializing all 8 sections. The wrapper also accepts the section via\n", "`` + `section=` if you want to skip the intermediate raw dict." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "if summary_raw is not None:\n", " just_batting = parse_cricket_summary(summary_raw, section=\"matchcards_batting\")\n", " print(type(just_batting), just_batting.shape)\n", "\n", " # pandas interop β€” same one-liner as every other sdv-py endpoint\n", " just_batting_pd = parse_cricket_summary(summary_raw, section=\"matchcards_batting\",\n", " return_as_pandas=True)\n", " print(type(just_batting_pd))\n", "else:\n", " print(\"no payload to parse\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## πŸ“° News and injuries\n", "\n", "[`espn_cricket_news`](../cricket/reference/site.md#espn_cricket_news) and\n", "[`espn_cricket_injuries`](../cricket/reference/site.md#espn_cricket_injuries)\n", "both follow the universal wrapper contract β€” they return a polars frame by default,\n", "using the shared `parse_news` and `parse_injuries` parsers from\n", "`_common_espn_parsers`." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "news = safe(\n", " \"IPL news\",\n", " lambda: cricket.espn_cricket_news(league=IPL, limit=5),\n", ")\n", "if news is not None and news.height:\n", " show_cols = [c for c in [\"headline\", \"published\", \"type\"] if c in news.columns]\n", " print(news.select(show_cols).head(5))\n", "else:\n", " print(\"news unavailable right now\")" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "injuries = safe(\n", " \"IPL injuries\",\n", " lambda: cricket.espn_cricket_injuries(league=IPL),\n", ")\n", "if injuries is not None and injuries.height:\n", " print(injuries.head())\n", "else:\n", " print(\"injuries feed unavailable or empty right now\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## πŸ“… Calendar\n", "\n", "[`espn_cricket_calendar`](../cricket/reference/site.md#espn_cricket_calendar)\n", "returns the competition calendar β€” matchdays, rounds, or phases depending on the\n", "tournament format. It uses the universal `parse_items` parser." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "cal = safe(\n", " \"IPL calendar\",\n", " lambda: cricket.espn_cricket_calendar(league=IPL),\n", ")\n", "if cal is not None and cal.height:\n", " print(cal.shape)\n", " print(cal.head())\n", "else:\n", " print(\"calendar unavailable right now\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 🍳 Cookbook: common cricket tasks\n", "\n", "A handful of patterns you'll reach for constantly when working with the cricket surface." ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Recipe 1 β€” Top run-scorers from a batting scorecard 🏏\n", "\n", "Filter to the highest individual scores from a batting matchcard. Useful for\n", "building a match-by-match batting leaderboard." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "if batting.height and \"runs\" in batting.columns:\n", " (\n", " batting\n", " .filter(pl.col(\"runs\").is_not_null())\n", " .sort(\"runs\", descending=True)\n", " .select([c for c in [\"innings_number\", \"athlete_display_name\", \"runs\",\n", " \"balls\", \"fours\", \"sixes\", \"strike_rate\"]\n", " if c in batting.columns])\n", " .head(5)\n", " )\n", "else:\n", " print(\"batting frame not available\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Recipe 2 β€” Economy leaders from the bowling card 🎯\n", "\n", "Bowlers who took wickets and kept a tight economy rate β€” the T20 game-changers." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "if bowling.height and \"economy\" in bowling.columns:\n", " (\n", " bowling\n", " .filter(pl.col(\"wickets\").is_not_null() & (pl.col(\"wickets\").cast(pl.Float64, strict=False) > 0))\n", " .sort(\"economy\", descending=False)\n", " .select([c for c in [\"innings_number\", \"athlete_display_name\",\n", " \"overs\", \"wickets\", \"runs_conceded\", \"economy\"]\n", " if c in bowling.columns])\n", " .head(5)\n", " )\n", "else:\n", " print(\"bowling frame not available\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Recipe 3 β€” Largest partnerships 🀝\n", "\n", "Identify which batting pairs put on the biggest stands in a given innings.\n", "A large partnership is often the turning point in a T20 match." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "if partnerships.height and \"total_runs\" in partnerships.columns:\n", " (\n", " partnerships\n", " .filter(pl.col(\"total_runs\").is_not_null())\n", " .sort(\"total_runs\", descending=True)\n", " .select([c for c in [\"innings_number\", \"batter1_display_name\", \"batter2_display_name\",\n", " \"total_runs\", \"total_balls\"]\n", " if c in partnerships.columns])\n", " .head(5)\n", " )\n", "else:\n", " print(\"partnerships frame not available\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Recipe 4 β€” Standings: current top-4 playoff picture πŸ†\n", "\n", "In the IPL, the top 4 teams after the group stage advance to the playoffs.\n", "Sort the standings frame by points (then net run rate as a tiebreaker) to\n", "see where each franchise stands." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "if standings is not None and standings.height:\n", " sort_cols = [c for c in [\"points\", \"net_run_rate\"] if c in standings.columns]\n", " if sort_cols:\n", " (\n", " standings\n", " .sort(sort_cols, descending=[True] * len(sort_cols))\n", " .select([c for c in [\"team\", \"wins\", \"losses\", \"points\", \"net_run_rate\"]\n", " if c in standings.columns])\n", " .head(4)\n", " )\n", " else:\n", " print(\"expected sort columns not present:\", standings.columns)\n", "else:\n", " print(\"standings not available\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Recipe 5 β€” pandas interop: grouping rosters by role 🐼\n", "\n", "Every sdv-py endpoint accepts `return_as_pandas=True`, so dropping into the\n", "pandas world is a single keyword. Here we pull the match rosters as a pandas\n", "DataFrame and count players by position/type." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "if summary_raw is not None:\n", " rosters_pd = parse_cricket_summary(summary_raw, section=\"rosters\",\n", " return_as_pandas=True)\n", " if rosters_pd is not None and len(rosters_pd):\n", " print(type(rosters_pd))\n", " print(rosters_pd.columns.tolist())\n", " # Count players by position if the column exists\n", " pos_col = next((c for c in [\"position_name\", \"position\", \"type\"]\n", " if c in rosters_pd.columns), None)\n", " if pos_col:\n", " print(rosters_pd.groupby(pos_col, dropna=False).size().sort_values(ascending=False))\n", " else:\n", " print(\"rosters section empty\")\n", "else:\n", " print(\"no payload\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## ⚠️ Caveats and known limitations\n", "\n", "**No teams endpoint for IPL.** \n", "`espn_cricket_teams_site(league=\"8048\")` returns HTTP 404 β€” ESPN does not expose\n", "a teams listing for the IPL through the Site v2 API. Use the `season_teams`\n", "endpoint if you need franchise metadata for a specific season:\n", "\n", "```python\n", "teams_seasonal = safe(\n", " \"IPL season teams\",\n", " lambda: cricket.espn_cricket_season_teams(league=IPL),\n", ")\n", "```\n", "\n", "**`event_id` is required for `espn_cricket_summary`.** \n", "Unlike the scoreboard (which returns today's slate without an ID), the summary\n", "endpoint needs a specific event identifier. Obtain `event_id` from the\n", "`espn_cricket_scoreboard` output (`event_id` column).\n", "\n", "**Off-season scoreboards may be empty.** \n", "The IPL runs April–May; calling `espn_cricket_scoreboard(league=\"8048\")` in\n", "December returns an empty `events` list. The parser returns a zero-row frame\n", "(never raises), so your code doesn't need to guard against exceptions β€”\n", "only against `.height == 0`.\n", "\n", "**League slugs vary.** \n", "ESPN doesn't publish a canonical slug list. Common IPL slug is `\"8048\"`;\n", "England's county T20 Blast uses `\"eng.t20\"`. Use\n", "`espn_cricket_league_root(league=slug, return_parsed=False)` to verify that\n", "a slug resolves before building a pipeline around it.\n", "\n", "**`parse_cricket_summary` section `standings` reflects in-tournament state.** \n", "The standings embedded inside a match summary are a snapshot at match time.\n", "For the current full-tournament standings table, use `espn_cricket_standings`\n", "directly." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# Demonstrating the safe empty-frame contract β€” no exception even for an empty payload.\n", "from sportsdataverse.cricket.cricket_espn_parsers import parse_cricket_scoreboard\n", "\n", "empty_df = parse_cricket_scoreboard({})\n", "print(\"Empty payload β†’ zero-row frame:\", empty_df.shape)\n", "\n", "empty_frames = parse_cricket_summary({})\n", "print(\"Empty summary β†’ dict of zero-row frames:\")\n", "for name, df in empty_frames.items():\n", " print(f\" {name}: {df.shape}\")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## πŸŽ‰ Where to next\n", "\n", "- πŸ“‘ **Full endpoint reference** β€” every `espn_cricket_*` wrapper is documented\n", " on the [Cricket reference](../cricket/reference/) pages, grouped by Site v2,\n", " Web v3, and Core v2 families.\n", "- πŸ”‘ **`event_id` lookup** β€” the scoreboard frame's `event_id` column is the\n", " key that unlocks the full summary. Build a pipeline:\n", " `scoreboard β†’ filter completed β†’ event_id β†’ summary β†’ batting/bowling`.\n", "- 🌍 **Other leagues** β€” swap the `league=` slug to explore England county\n", " (`\"eng.1\"`), ICC Men's/Women's World Cup, PSL, BBL, and more. The\n", " parsers and workflow are identical across leagues.\n", "- 🐼 Pass `return_as_pandas=True` for pandas, or `return_parsed=False` on\n", " any `espn_cricket_*` wrapper for the raw ESPN JSON.\n", "- 🎯 **Player depth** β€” `espn_cricket_player_info`, `espn_cricket_player_gamelog`,\n", " and `espn_cricket_player_stats` accept `league=` + `athlete_id=` and follow the\n", " same `return_parsed` / `return_as_pandas` contract.\n", "- πŸŸ₯ The sister R package ecosystem is covered by\n", " [cfbfastR](https://cfbfastR.sportsdataverse.org) (American football),\n", " [hoopR](https://hoopR.sportsdataverse.org) (basketball),\n", " [baseballr](https://baseballr.sportsdataverse.org) (baseball), and\n", " [fastRhockey](https://fastRhockey.sportsdataverse.org) (hockey) β€” cricket\n", " sits in the Python-only surface for now.\n" ] } ], "metadata": { "kernelspec": { "display_name": "Python 3 (ipykernel)", "language": "python", "name": "python3" }, "language_info": { "codemirror_mode": { "name": "ipython", "version": 3 }, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "version": "3.11.0" } }, "nbformat": 4, "nbformat_minor": 5 }