{ "nbformat": 4, "nbformat_minor": 5, "metadata": { "kernelspec": { "display_name": "Python 3", "language": "python", "name": "python3" }, "language_info": { "name": "python", "version": "3.11.0" } }, "cells": [ { "cell_type": "markdown", "id": "soccer-title", "metadata": {}, "source": [ "# ⚽ Soccer with `sportsdataverse-py`\n", "\n", "From the Premier League to the World Cup, `sportsdataverse.soccer` gives you the global\n", "game in tidy **polars** DataFrames — no API key, no config, just pip and import.\n", "\n", "The entire surface is built on ESPN's public Site v2, Web v3, and Core v2 endpoints and\n", "surfaced through a single family of **`espn_soccer_*` wrappers** that accept a `league=` slug\n", "— so one call covers Premier League, MLS, Champions League, La Liga, or any of the other\n", "leagues ESPN tracks. Twelve league aliases (`espn_epl_*`, `espn_mls_*`, `espn_ucl_*`, …) let\n", "you drop the `league=` argument entirely when you only ever work one competition.\n", "\n", "R user? The closest companion for the European game is\n", "[worldfootballR](https://jaseziv.github.io/worldfootballR/).\n", "For women's basketball orbiting the same ESPN platform, see\n", "[wehoop](https://wehoop.sportsdataverse.org).\n", "\n", "Let's kick it off! ⚽" ] }, { "cell_type": "markdown", "id": "soccer-toolbox", "metadata": {}, "source": [ "## 🧰 The toolbox\n", "\n", "Everything returns a tidy **polars** `DataFrame` by default — pass\n", "`return_as_pandas=True` for pandas. The wrappers return the raw ESPN JSON by\n", "default; pass `` to run the built-in parser. ⭐ marks the\n", "most commonly used entry points.\n", "\n", "### Core wrappers (pass `league=` slug)\n", "\n", "| Function | What it gives you |\n", "|---|---|\n", "| [`espn_soccer_scoreboard`](../soccer/reference/site.md#espn_soccer_scoreboard) | ⭐ Match results / live scores for a date |\n", "| [`espn_soccer_standings`](../soccer/reference/site.md#espn_soccer_standings) | ⭐ League / conference / group standings |\n", "| [`espn_soccer_summary`](../soccer/reference/site.md#espn_soccer_summary) | ⭐ Full match summary — lineups, key events, team stats, commentary |\n", "| [`espn_soccer_teams_site`](../soccer/reference/site.md#espn_soccer_teams_site) | ⭐ Every team in the league (grab `team_id`s) |\n", "| [`espn_soccer_team_roster`](../soccer/reference/site.md#espn_soccer_team_roster) | One team's current squad |\n", "| [`espn_soccer_team_schedule`](../soccer/reference/site.md#espn_soccer_team_schedule) | A team's fixtures & results |\n", "| [`espn_soccer_news`](../soccer/reference/site.md#espn_soccer_news) | Latest news headlines for a league |\n", "| [`espn_soccer_injuries`](../soccer/reference/site.md#espn_soccer_injuries) | Current injury list |\n", "| [`espn_soccer_leaders`](../soccer/reference/web.md#espn_soccer_leaders) | Stat leaders by category |\n", "| [`espn_soccer_player_info`](../soccer/reference/core.md#espn_soccer_player_info) | Player bio & metadata |\n", "| [`espn_soccer_player_stats`](../soccer/reference/site.md#espn_soccer_player_stats) | Player season statistics |\n", "| [`espn_soccer_game_probabilities`](../soccer/reference/site.md#espn_soccer_game_probabilities) | In-game win probabilities |\n", "\n", "### Parsers (turn raw JSON into polars)\n", "\n", "| Parser | Paired with |\n", "|---|---|\n", "| `parse_soccer_scoreboard` | `espn_soccer_scoreboard` |\n", "| `parse_soccer_standings` | `espn_soccer_standings` |\n", "| `parse_soccer_summary` | `espn_soccer_summary` — 11-section dispatcher |\n", "| `parse_soccer_teams` | `espn_soccer_teams_site` |\n", "| `parse_soccer_team_roster` | `espn_soccer_team_roster` |\n", "\n", "### League slugs quick reference\n", "\n", "| `league=` | Competition |\n", "|---|---|\n", "| `eng.1` | Premier League |\n", "| `usa.1` | MLS |\n", "| `uefa.champions` | Champions League |\n", "| `esp.1` | La Liga |\n", "| `ger.1` | Bundesliga |\n", "| `ita.1` | Serie A |\n", "| `fra.1` | Ligue 1 |\n", "| `uefa.europa` | Europa League |\n", "| `usa.nwsl` | NWSL |\n", "| `mex.1` | Liga MX |\n", "| `fifa.world` | FIFA World Cup |\n", "| `fifa.wwc` | FIFA Women's World Cup |" ] }, { "cell_type": "markdown", "id": "soccer-setup", "metadata": {}, "source": [ "## 🔌 Setup\n", "\n", "```sh\n", "pip install sportsdataverse\n", "```\n", "\n", "No API key required. All calls go to ESPN's public endpoints." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-imports", "metadata": {}, "outputs": [], "source": [ "import polars as pl\n", "import sportsdataverse.soccer as soccer\n", "from sportsdataverse.soccer.soccer_espn_parsers import (\n", " parse_soccer_scoreboard,\n", " parse_soccer_standings,\n", " parse_soccer_summary,\n", " parse_soccer_teams,\n", " parse_soccer_team_roster,\n", ")\n", "\n", "# The league aliases live in sub-modules;\n", "# import them for the alias-demo section later.\n", "from sportsdataverse.soccer import epl, mls, ucl, laliga, bundesliga\n", "\n", "print('polars version:', pl.__version__)" ] }, { "cell_type": "markdown", "id": "soccer-safe-md", "metadata": {}, "source": [ "ESPN's live endpoints are seasonal and occasionally rate-limited, so a tiny\n", "`safe()` helper runs them defensively — you get the frame when the feed is up,\n", "and a friendly one-liner when it isn't (never a scary traceback). 🛟" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-safe", "metadata": {}, "outputs": [], "source": [ "from sportsdataverse.errors import AssetFetchError, NoDataError\n", "\n", "def safe(label, thunk):\n", " try:\n", " out = thunk()\n", " print(f'✅ {label}')\n", " return out\n", " except (NoDataError, AssetFetchError) as e:\n", " print(f\"\\u23ed\\ufe0f {label}: {type(e).__name__}: {e}\")\n", " return None" ] }, { "cell_type": "markdown", "id": "soccer-scoreboard-md", "metadata": {}, "source": [ "## 📅 Scoreboard — a day's results\n", "\n", "[`espn_soccer_scoreboard`](../soccer/reference/site.md#espn_soccer_scoreboard)\n", "returns every match for a league on a given date. Pass `dates=YYYYMMDD`; omit\n", "it to get today's slate. The raw payload is a nested ESPN JSON dict — pass\n", "`` to flatten it into a tidy polars frame, or call\n", "`parse_soccer_scoreboard` explicitly on the raw dict.\n", "\n", "We'll start with a **Premier League** match-day." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-scoreboard-raw", "metadata": {}, "outputs": [], "source": [ "# Raw payload (the default) — useful when you need the full nested structure.\n", "raw_board = safe(\n", " 'EPL scoreboard (raw)',\n", " lambda: soccer.espn_soccer_scoreboard(league='eng.1', dates=20240310),\n", ")\n", "type(raw_board), list(raw_board.keys()) if isinstance(raw_board, dict) else 'unavailable'" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-scoreboard-parsed", "metadata": {}, "outputs": [], "source": [ "# Parsed frame — one row per match.\n", "board = safe(\n", " 'EPL scoreboard (parsed)',\n", " lambda: parse_soccer_scoreboard(\n", " soccer.espn_soccer_scoreboard(league='eng.1', dates=20240310)\n", " ),\n", ")\n", "if board is not None and getattr(board, 'height', 0):\n", " keep = [c for c in board.columns\n", " if c in ('game_id', 'name', 'short_name', 'status_type_description',\n", " 'home_team_abbreviation', 'away_team_abbreviation',\n", " 'home_score', 'away_score', 'date')]\n", " out = board.select(keep).head()\n", "else:\n", " out = 'no scoreboard data for that date'\n", "out" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-scoreboard-shape", "metadata": {}, "outputs": [], "source": [ "print('board shape:', getattr(board, 'shape', 'N/A'))\n", "print('columns:', getattr(board, 'columns', []))" ] }, { "cell_type": "markdown", "id": "soccer-standings-md", "metadata": {}, "source": [ "## 🏆 Standings\n", "\n", "[`espn_soccer_standings`](../soccer/reference/site.md#espn_soccer_standings)\n", "flattens the league table into one row per team per group/conference. The\n", "`group` column is what makes this multi-competition friendly:\n", "\n", "- **Single-table leagues** (EPL, La Liga, Bundesliga): one group, one table.\n", "- **Conference leagues** (MLS — Eastern/Western Conferences): filter on\n", " `group` to isolate a conference.\n", "- **Group-stage tournaments** (Champions League, World Cup): each\n", " group/group-stage pod gets its own `group` label.\n", "\n", "We'll pull the **Premier League** table and an **MLS** table side by side." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-standings-epl", "metadata": {}, "outputs": [], "source": [ "epl_table = safe(\n", " 'EPL standings',\n", " lambda: parse_soccer_standings(\n", " soccer.espn_soccer_standings(league='eng.1', season=2023)\n", " ),\n", ")\n", "if epl_table is not None and getattr(epl_table, 'height', 0):\n", " keep = [c for c in epl_table.columns\n", " if c in ('rank', 'team_name', 'games_played', 'wins', 'losses',\n", " 'draws', 'goals_for', 'goals_against', 'goal_difference',\n", " 'points', 'group')]\n", " out = epl_table.select(keep).head(10)\n", "else:\n", " out = 'standings unavailable right now'\n", "out" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-standings-mls", "metadata": {}, "outputs": [], "source": [ "# MLS standings — multiple groups (Eastern/Western Conference)\n", "mls_table = safe(\n", " 'MLS standings',\n", " lambda: parse_soccer_standings(\n", " soccer.espn_soccer_standings(league='usa.1', season=2023)\n", " ),\n", ")\n", "if mls_table is not None and getattr(mls_table, 'height', 0):\n", " keep = [c for c in mls_table.columns\n", " if c in ('rank', 'team_name', 'wins', 'losses', 'draws', 'points', 'group')]\n", " print('MLS groups:', mls_table['group'].unique().to_list() if 'group' in mls_table.columns else 'n/a')\n", " out = mls_table.select(keep).head(10)\n", "else:\n", " out = 'MLS standings unavailable right now'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-summary-intro-md", "metadata": {}, "source": [ "## 🎬 Match summary — the 11-section dispatcher\n", "\n", "[`espn_soccer_summary`](../soccer/reference/site.md#espn_soccer_summary) fetches\n", "the full ESPN Site v2 summary payload for a single match (~500 KB–1 MB of\n", "nested JSON). `parse_soccer_summary` turns that into a **dict of polars\n", "DataFrames**, one key per section:\n", "\n", "| Section | Content |\n", "|---|---|\n", "| `header` | Match header — teams, score, status, venue |\n", "| `lineups` | Starting XI + substitutes, one row per player |\n", "| `key_events` | Goals, cards, own goals, substitutions |\n", "| `team_stats` | Per-team aggregate stats (shots, possession, passes …) |\n", "| `commentary` | Live commentary log, one row per broadcast call |\n", "| `leaders` | Statistical leaders (top scorers, etc.) |\n", "| `standings` | In-payload mini standings snapshot |\n", "| `head_to_head` | H2H history between the two clubs |\n", "| `last_five` | Each team's most recent 5 results |\n", "| `game_info` | Venue, attendance, referee, season details |\n", "| `shootout` | Penalty shootout rows (when applicable) |\n", "\n", "We'll use a **Chelsea vs. Manchester City** EPL match\n", "(event `656009` — 12 Nov 2023) to walk through each section." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-summary-fetch", "metadata": {}, "outputs": [], "source": [ "EVENT_ID = 656009 # Chelsea vs Man City, EPL, 12 Nov 2023\n", "\n", "raw_summary = safe(\n", " f'EPL summary event {EVENT_ID}',\n", " lambda: soccer.espn_soccer_summary(league='eng.1', event_id=EVENT_ID),\n", ")\n", "# Parse all 11 sections at once\n", "if raw_summary is not None:\n", " frames = parse_soccer_summary(raw_summary)\n", " print('sections parsed:', list(frames.keys()))\n", " print('rows per section:', {k: v.height for k, v in frames.items()})\n", "else:\n", " frames = {}" ] }, { "cell_type": "markdown", "id": "soccer-summary-header-md", "metadata": {}, "source": [ "### Section: `header` — match overview" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-summary-header", "metadata": {}, "outputs": [], "source": [ "header = frames.get('header')\n", "if header is not None and header.height:\n", " keep = [c for c in header.columns\n", " if c in ('name', 'home_team_name', 'away_team_name',\n", " 'home_score', 'away_score', 'status_type_description',\n", " 'venue_full_name', 'date')]\n", " out = header.select(keep)\n", "else:\n", " out = 'header section unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-summary-lineups-md", "metadata": {}, "source": [ "### Section: `lineups` — starting XIs and substitutes" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-summary-lineups", "metadata": {}, "outputs": [], "source": [ "lineups = frames.get('lineups')\n", "if lineups is not None and lineups.height:\n", " keep = [c for c in lineups.columns\n", " if c in ('team_name', 'athlete_display_name', 'position_name',\n", " 'jersey', 'starter', 'subbedIn', 'subbedOut')]\n", " out = lineups.select(keep).head(10)\n", "else:\n", " out = 'lineups section unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-summary-key-events-md", "metadata": {}, "source": [ "### Section: `key_events` — goals, cards, substitutions" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-summary-key-events", "metadata": {}, "outputs": [], "source": [ "key_events = frames.get('key_events')\n", "if key_events is not None and key_events.height:\n", " keep = [c for c in key_events.columns\n", " if c in ('clock_display_value', 'team_name', 'athlete_display_name',\n", " 'type_text', 'text', 'score_value')]\n", " out = key_events.select(keep).head()\n", "else:\n", " out = 'key_events section unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-summary-team-stats-md", "metadata": {}, "source": [ "### Section: `team_stats` — possession, shots, passes …" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-summary-team-stats", "metadata": {}, "outputs": [], "source": [ "team_stats = frames.get('team_stats')\n", "if team_stats is not None and team_stats.height:\n", " print('team_stats columns:', team_stats.columns)\n", " out = team_stats.head()\n", "else:\n", " out = 'team_stats section unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-summary-commentary-md", "metadata": {}, "source": [ "### Section: `commentary` — live match log" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-summary-commentary", "metadata": {}, "outputs": [], "source": [ "commentary = frames.get('commentary')\n", "if commentary is not None and commentary.height:\n", " keep = [c for c in commentary.columns\n", " if c in ('clock_display_value', 'type_id', 'text')]\n", " out = commentary.select(keep).head(6)\n", "else:\n", " out = 'commentary section unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-summary-other-md", "metadata": {}, "source": [ "### Remaining sections at a glance\n", "\n", "The remaining five sections follow the same pattern — each is a tidy\n", "DataFrame keyed off `frames[\"
\"]`." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-summary-other", "metadata": {}, "outputs": [], "source": [ "# Quick peek at game_info, head_to_head, last_five, leaders, shootout\n", "for section in ('game_info', 'head_to_head', 'last_five', 'leaders', 'shootout'):\n", " df = frames.get(section)\n", " if df is not None:\n", " print(f'{section:20s} shape={df.shape} cols={df.columns[:5]}')\n", " else:\n", " print(f'{section:20s} not in parsed frames')" ] }, { "cell_type": "markdown", "id": "soccer-summary-single-section-md", "metadata": {}, "source": [ "### Requesting a single section\n", "\n", "Pass `section=\"\"` to `parse_soccer_summary` when you only need one\n", "slice — the parser skips the rest and returns a single DataFrame directly." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-summary-single-section", "metadata": {}, "outputs": [], "source": [ "if raw_summary is not None:\n", " ke = parse_soccer_summary(raw_summary, section='key_events')\n", " print(type(ke).__name__, ke.shape)\n", " out = ke.head(3)\n", "else:\n", " out = 'summary unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-teams-md", "metadata": {}, "source": [ "## 🏟️ Teams — the master lookup\n", "\n", "[`espn_soccer_teams_site`](../soccer/reference/site.md#espn_soccer_teams_site)\n", "lists every team in a league. The `team_id` column is the key you feed into\n", "every team-scoped call (roster, schedule, injuries …)." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-teams", "metadata": {}, "outputs": [], "source": [ "epl_teams = safe(\n", " 'EPL teams',\n", " lambda: parse_soccer_teams(\n", " soccer.espn_soccer_teams_site(league='eng.1')\n", " ),\n", ")\n", "if epl_teams is not None and epl_teams.height:\n", " keep = [c for c in epl_teams.columns\n", " if c in ('team_id', 'display_name', 'abbreviation',\n", " 'location', 'name', 'short_display_name', 'is_active')]\n", " out = epl_teams.select(keep).head(10)\n", "else:\n", " out = 'teams unavailable right now'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-team-roster-md", "metadata": {}, "source": [ "## 👥 Team roster\n", "\n", "[`espn_soccer_team_roster`](../soccer/reference/site.md#espn_soccer_team_roster)\n", "pulls the current squad for a single team. Pass `team_id=` from the teams\n", "frame above. We'll use **Arsenal** (`team_id=359`)." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-team-roster", "metadata": {}, "outputs": [], "source": [ "ARSENAL_ID = 359 # Arsenal FC\n", "\n", "roster = safe(\n", " f'Arsenal roster (team_id={ARSENAL_ID})',\n", " lambda: parse_soccer_team_roster(\n", " soccer.espn_soccer_team_roster(league='eng.1', team_id=ARSENAL_ID)\n", " ),\n", ")\n", "if roster is not None and getattr(roster, 'height', 0):\n", " keep = [c for c in roster.columns\n", " if c in ('athlete_id', 'display_name', 'jersey',\n", " 'position_name', 'age', 'birth_country')]\n", " out = roster.select(keep).head(10)\n", "else:\n", " out = 'roster unavailable right now'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-team-schedule-md", "metadata": {}, "source": [ "## 📆 Team schedule\n", "\n", "[`espn_soccer_team_schedule`](../soccer/reference/site.md#espn_soccer_team_schedule)\n", "returns the fixtures and results for a team in a given season. Pass `season=`\n", "as an integer year to scope to a specific campaign." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-team-schedule", "metadata": {}, "outputs": [], "source": [ "team_sched = safe(\n", " f'Arsenal schedule 2023/24',\n", " lambda: soccer.espn_soccer_team_schedule(\n", " league='eng.1', team_id=ARSENAL_ID, season=2024,\n", " ),\n", ")\n", "if isinstance(team_sched, dict):\n", " # Raw payload — show top-level keys as an orientation\n", " print('payload keys:', list(team_sched.keys()))\n", "elif team_sched is not None:\n", " print('shape:', team_sched.shape)\n", " print(team_sched.head())\n", "else:\n", " print('schedule unavailable right now')" ] }, { "cell_type": "markdown", "id": "soccer-news-injuries-md", "metadata": {}, "source": [ "## 🗞️ News & injuries\n", "\n", "Two lightweight feeds round out the live suite:\n", "\n", "- [`espn_soccer_news`](../soccer/reference/site.md#espn_soccer_news) — latest\n", " editorial headlines for a league, optionally limited by `limit=`.\n", "- [`espn_soccer_injuries`](../soccer/reference/site.md#espn_soccer_injuries) —\n", " current injury list for a league (all teams)." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-news", "metadata": {}, "outputs": [], "source": [ "news = safe(\n", " 'EPL news',\n", " lambda: soccer.espn_soccer_news(league='eng.1', limit=10),\n", ")\n", "if isinstance(news, dict):\n", " articles = news.get('articles', [])\n", " for a in articles[:3]:\n", " if isinstance(a, dict):\n", " print('•', a.get('headline', a.get('title', str(a)))[:100])\n", "else:\n", " print('news unavailable right now')" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-injuries", "metadata": {}, "outputs": [], "source": [ "injuries = safe(\n", " 'EPL injuries',\n", " lambda: soccer.espn_soccer_injuries(league='eng.1'),\n", ")\n", "if isinstance(injuries, dict):\n", " print('injury payload keys:', list(injuries.keys()))\n", "elif injuries is not None:\n", " print('shape:', injuries.shape)\n", "else:\n", " print('injuries feed unavailable right now')" ] }, { "cell_type": "markdown", "id": "soccer-leaders-md", "metadata": {}, "source": [ "## 📊 Stat leaders\n", "\n", "[`espn_soccer_leaders`](../soccer/reference/web.md#espn_soccer_leaders) returns\n", "the statistical leaderboard for a league, optionally scoped to a season,\n", "stat category, and page. Categories include `goals`, `assists`, `yellowCards`,\n", "`redCards`, and others ESPN surfaces." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-leaders", "metadata": {}, "outputs": [], "source": [ "goal_leaders = safe(\n", " 'EPL goal leaders 2023/24',\n", " lambda: soccer.espn_soccer_leaders(\n", " league='eng.1', category='goals', season=2024, limit=10,\n", " ),\n", ")\n", "if isinstance(goal_leaders, dict):\n", " print('leaders payload keys:', list(goal_leaders.keys()))\n", "elif goal_leaders is not None:\n", " print('shape:', goal_leaders.shape)\n", " print(goal_leaders.head())\n", "else:\n", " print('leaders unavailable right now')" ] }, { "cell_type": "markdown", "id": "soccer-aliases-md", "metadata": {}, "source": [ "## 🔗 League aliases — drop the `league=` argument\n", "\n", "When you only work with one competition, the twelve league sub-modules\n", "(`epl`, `mls`, `ucl`, `laliga`, `bundesliga`, `seriea`, `ligue1`, `uel`,\n", "`nwsl`, `ligamx`, `wc`, `wwc`) provide pre-bound aliases where `league=` is\n", "already wired in. Every `espn_soccer_*` function has a corresponding\n", "`espn__*` variant.\n", "\n", "```python\n", "# These three calls are exactly equivalent:\n", "soccer.espn_soccer_scoreboard(league='eng.1', dates=20240310)\n", "epl.espn_epl_scoreboard(dates=20240310)\n", "from sportsdataverse.soccer.epl import espn_epl_scoreboard\n", "espn_epl_scoreboard(dates=20240310)\n", "```\n", "\n", "Here's a quick tour of three aliases side by side." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-aliases-epl", "metadata": {}, "outputs": [], "source": [ "# --- EPL alias ---\n", "epl_board = safe(\n", " 'epl.espn_epl_scoreboard',\n", " lambda: parse_soccer_scoreboard(\n", " epl.espn_epl_scoreboard(dates=20240310)\n", " ),\n", ")\n", "if epl_board is not None and getattr(epl_board, 'height', 0):\n", " keep = [c for c in epl_board.columns\n", " if c in ('short_name', 'home_score', 'away_score', 'status_type_description')]\n", " out = epl_board.select(keep).head()\n", "else:\n", " out = 'EPL alias unavailable'\n", "out" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-aliases-mls", "metadata": {}, "outputs": [], "source": [ "# --- MLS alias: standings ---\n", "mls_alias_table = safe(\n", " 'mls.espn_mls_standings',\n", " lambda: parse_soccer_standings(\n", " mls.espn_mls_standings(season=2023)\n", " ),\n", ")\n", "if mls_alias_table is not None and getattr(mls_alias_table, 'height', 0):\n", " keep = [c for c in mls_alias_table.columns\n", " if c in ('rank', 'team_name', 'wins', 'losses', 'draws', 'points', 'group')]\n", " out = mls_alias_table.select(keep).head(8)\n", "else:\n", " out = 'MLS alias unavailable'\n", "out" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-aliases-ucl", "metadata": {}, "outputs": [], "source": [ "# --- UCL alias: standings (shows group labels A–H during the group stage) ---\n", "ucl_table = safe(\n", " 'ucl.espn_ucl_standings',\n", " lambda: parse_soccer_standings(\n", " ucl.espn_ucl_standings(season=2024)\n", " ),\n", ")\n", "if ucl_table is not None and getattr(ucl_table, 'height', 0):\n", " keep = [c for c in ucl_table.columns\n", " if c in ('rank', 'team_name', 'wins', 'draws', 'losses', 'points', 'group')]\n", " print('UCL groups:', ucl_table['group'].unique().sort().to_list() if 'group' in ucl_table.columns else 'n/a')\n", " out = ucl_table.select(keep).head(10)\n", "else:\n", " out = 'UCL alias unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-cookbook-md", "metadata": {}, "source": [ "## 🍳 Cookbook: common soccer tasks\n", "\n", "The real fun is in the questions. Six recipes built from the frames we've\n", "already pulled — every one ends in a tidy, ready-to-read frame." ] }, { "cell_type": "markdown", "id": "soccer-recipe1-md", "metadata": {}, "source": [ "### Recipe 1 — Derive goal difference from the standings table 📈\n", "\n", "The standings frame already ships `goals_for` and `goals_against`;\n", "sort by goal difference to see who dominated the scoring charts." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-recipe1", "metadata": {}, "outputs": [], "source": [ "if (epl_table is not None and getattr(epl_table, 'height', 0)\n", " and {'goals_for', 'goals_against'}.issubset(epl_table.columns)):\n", " out = (\n", " epl_table\n", " .with_columns(\n", " (pl.col('goals_for') - pl.col('goals_against')).alias('gd')\n", " )\n", " .select(['rank', 'team_name', 'goals_for', 'goals_against', 'gd', 'points'])\n", " .sort('gd', descending=True)\n", " .head(10)\n", " )\n", "else:\n", " out = 'run the standings cell above first'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-recipe2-md", "metadata": {}, "source": [ "### Recipe 2 — Starter vs. substitute counts from a lineup 🧮\n", "\n", "Count the starters and bench players per team from the `lineups`\n", "section of the match summary we already fetched." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-recipe2", "metadata": {}, "outputs": [], "source": [ "lineup_df = frames.get('lineups')\n", "if (lineup_df is not None and lineup_df.height\n", " and {'team_name', 'starter'}.issubset(lineup_df.columns)):\n", " out = (\n", " lineup_df\n", " .group_by(['team_name', 'starter'])\n", " .agg(pl.len().alias('players'))\n", " .sort(['team_name', 'starter'])\n", " )\n", "else:\n", " out = 'run the match-summary cells above first'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-recipe3-md", "metadata": {}, "source": [ "### Recipe 3 — Goals, cards, and substitutions in the key-events log 🎯\n", "\n", "The `key_events` section has a `type_text` column. Group it to get a\n", "breakdown of event types for the match." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-recipe3", "metadata": {}, "outputs": [], "source": [ "ke_df = frames.get('key_events')\n", "if ke_df is not None and ke_df.height and 'type_text' in ke_df.columns:\n", " out = (\n", " ke_df\n", " .group_by('type_text')\n", " .agg(pl.len().alias('count'))\n", " .sort('count', descending=True)\n", " )\n", "else:\n", " out = 'run the match-summary cells above first'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-recipe4-md", "metadata": {}, "source": [ "### Recipe 4 — Position breakdown of a squad 🏃\n", "\n", "Use the parsed roster to count players by position — a quick depth-chart\n", "picture for squad assessment." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-recipe4", "metadata": {}, "outputs": [], "source": [ "if (roster is not None and getattr(roster, 'height', 0)\n", " and 'position_name' in roster.columns):\n", " out = (\n", " roster\n", " .group_by('position_name')\n", " .agg(pl.len().alias('players'))\n", " .sort('players', descending=True)\n", " )\n", "else:\n", " out = 'run the team-roster cell above first'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-recipe5-md", "metadata": {}, "source": [ "### Recipe 5 — Multi-league standings in one loop 🌍\n", "\n", "Because the league slug is just a string, looping over several competitions\n", "is trivial. Count the teams in the standings table for four top leagues." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-recipe5", "metadata": {}, "outputs": [], "source": [ "leagues = {\n", " 'eng.1': 'Premier League',\n", " 'esp.1': 'La Liga',\n", " 'ger.1': 'Bundesliga',\n", " 'ita.1': 'Serie A',\n", "}\n", "rows = []\n", "for slug, name in leagues.items():\n", " result = safe(\n", " f'{name} standings',\n", " lambda s=slug: parse_soccer_standings(\n", " soccer.espn_soccer_standings(league=s, season=2023)\n", " ),\n", " )\n", " rows.append({\n", " 'league': name,\n", " 'slug': slug,\n", " 'teams': result.height if result is not None else 0,\n", " 'groups': (result['group'].n_unique() if 'group' in result.columns else 1)\n", " if result is not None and result.height else 0,\n", " })\n", "\n", "pl.DataFrame(rows)" ] }, { "cell_type": "markdown", "id": "soccer-recipe6-md", "metadata": {}, "source": [ "### Recipe 6 — Pandas interop 🐼\n", "\n", "Every parser accepts `return_as_pandas=True`, and any polars frame\n", "converts with `.to_pandas()`. Once you're in pandas land the full\n", "pandas / NumPy / scikit-learn world opens up." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-recipe6", "metadata": {}, "outputs": [], "source": "if epl_table is not None and getattr(epl_table, 'height', 0):\n epl_pd = epl_table.to_pandas()\n print(type(epl_pd).__name__, '|', epl_pd.shape)\n numeric = [c for c in ('wins', 'losses', 'draws', 'goals_for', 'goals_against', 'points')\n if c in epl_pd.columns]\n out = epl_pd[numeric].describe().round(1) if numeric else epl_pd.head()\nelse:\n _epl_pd = safe(\n 'EPL standings (pandas)',\n lambda: parse_soccer_standings(\n soccer.espn_soccer_standings(league='eng.1', season=2023),\n return_as_pandas=True,\n ),\n )\n if _epl_pd is not None and not getattr(_epl_pd, 'empty', True):\n out = _epl_pd\n else:\n out = 'standings unavailable right now'\nout" }, { "cell_type": "markdown", "id": "soccer-chapions-league-md", "metadata": {}, "source": [ "## 🏅 Champions League deep-dive\n", "\n", "A quick tour of the same surface on the UCL — demonstrating that\n", "the exact same parser stack handles group-stage multi-table standings." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-ucl-scoreboard", "metadata": {}, "outputs": [], "source": [ "ucl_board = safe(\n", " 'UCL scoreboard',\n", " lambda: parse_soccer_scoreboard(\n", " soccer.espn_soccer_scoreboard(league='uefa.champions', dates=20231107)\n", " ),\n", ")\n", "if ucl_board is not None and getattr(ucl_board, 'height', 0):\n", " keep = [c for c in ucl_board.columns\n", " if c in ('short_name', 'home_score', 'away_score',\n", " 'status_type_description', 'date')]\n", " out = ucl_board.select(keep).head()\n", "else:\n", " out = 'UCL scoreboard unavailable'\n", "out" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-ucl-groups", "metadata": {}, "outputs": [], "source": [ "# UCL group-stage standings — one row per team per group\n", "if ucl_table is not None and getattr(ucl_table, 'height', 0):\n", " keep = [c for c in ucl_table.columns\n", " if c in ('group', 'rank', 'team_name', 'wins', 'draws', 'losses', 'points')]\n", " # Show Group A only\n", " if 'group' in ucl_table.columns:\n", " first_group = ucl_table['group'].sort()[0]\n", " out = (\n", " ucl_table\n", " .filter(pl.col('group') == first_group)\n", " .select(keep)\n", " .sort('rank')\n", " )\n", " else:\n", " out = ucl_table.select(keep).head(8)\n", "else:\n", " out = 'run the UCL standings cell above first'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-women-md", "metadata": {}, "source": [ "## 🌸 Women's soccer — NWSL & Women's World Cup\n", "\n", "The same wrappers cover women's football. Swap the league slug:\n", "\n", "- `usa.nwsl` — National Women's Soccer League (NWSL)\n", "- `fifa.wwc` — FIFA Women's World Cup\n", "- `eng.wsl` — FA Women's Super League\n", "- `usa.ncaa.w.soccer` — NCAA Women's Division I\n", "\n", "The `nwsl` sub-module alias mirrors the pattern of `epl`, `mls`, etc." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-nwsl", "metadata": {}, "outputs": [], "source": [ "from sportsdataverse.soccer import nwsl\n", "\n", "nwsl_table = safe(\n", " 'NWSL standings 2023',\n", " lambda: parse_soccer_standings(\n", " nwsl.espn_nwsl_standings(season=2023)\n", " ),\n", ")\n", "if nwsl_table is not None and getattr(nwsl_table, 'height', 0):\n", " keep = [c for c in nwsl_table.columns\n", " if c in ('rank', 'team_name', 'wins', 'losses', 'draws', 'points', 'group')]\n", " out = nwsl_table.select(keep).head(10)\n", "else:\n", " out = 'NWSL standings unavailable right now'\n", "out" ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-wwc", "metadata": {}, "outputs": [], "source": [ "# Women's World Cup 2023 — group-stage standings\n", "from sportsdataverse.soccer import wwc\n", "\n", "wwc_table = safe(\n", " \"WWC 2023 standings\",\n", " lambda: parse_soccer_standings(\n", " wwc.espn_wwc_standings(season=2023)\n", " ),\n", ")\n", "if wwc_table is not None and getattr(wwc_table, 'height', 0):\n", " keep = [c for c in wwc_table.columns\n", " if c in ('group', 'rank', 'team_name', 'wins', 'draws', 'losses', 'points')]\n", " print('WWC groups:', wwc_table['group'].unique().sort().to_list() if 'group' in wwc_table.columns else 'n/a')\n", " out = wwc_table.select(keep).head(8)\n", "else:\n", " out = 'WWC standings unavailable right now'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-kloppy-intro", "metadata": {}, "source": [ "## 📦 Event data from any provider with kloppy\n", "\n", "ESPN gives you the match log; for per-event coordinates (every pass, carry, shot) the open\n", "standard is [kloppy](https://kloppy.pysport.org), which reads ~15 providers (StatsBomb, Opta,\n", "Wyscout, Sportec, SkillCorner, …) into one event model. It is an **optional extra**:\n", "\n", "```bash\n", "pip install \"sportsdataverse[soccer]\"\n", "```\n", "\n", "`soccer_open_events('statsbomb', match_id)` loads one match of\n", "[StatsBomb open data](https://github.com/statsbomb/open-data) (free for research and\n", "non-commercial use) as a polars frame — one row per event, snake_case columns, coordinates on\n", "kloppy's 0–1 pitch by default (`coordinates='statsbomb'` keeps the 120 × 80 units).\n", "Match 8658 is the 2018 World Cup final, France v Croatia." ] }, { "cell_type": "code", "execution_count": null, "id": "soccer-kloppy-load", "metadata": {}, "outputs": [], "source": [ "from sportsdataverse.soccer import soccer_open_events\n", "\n", "events = safe(\n", " 'StatsBomb open data — 2018 World Cup final (match 8658)',\n", " lambda: soccer_open_events('statsbomb', 8658),\n", ")\n", "if events is not None:\n", " print('events:', events.shape)\n", " out = (\n", " events.filter(pl.col('event_type') == 'SHOT')\n", " .select('period_id', 'timestamp', 'team_id', 'player_id',\n", " 'coordinates_x', 'coordinates_y', 'result')\n", " .head(10)\n", " )\n", "else:\n", " out = 'kloppy not installed (pip install \"sportsdataverse[soccer]\") or StatsBomb open data unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "soccer-kloppy-next", "metadata": {}, "source": [ "Any other provider or file goes through kloppy itself, then\n", "`soccer_events_to_frame(dataset)` applies the same frame convention:\n", "\n", "```python\n", "from kloppy import statsbomb\n", "from sportsdataverse.soccer import soccer_events_to_frame\n", "ds = statsbomb.load(event_data='8658.json', lineup_data='lineups_8658.json')\n", "df = soccer_events_to_frame(ds)\n", "```\n", "\n", "To draw it, [sdvplot](https://github.com/sportsdataverse/sdvplot) puts the frame on a 105 × 68 m\n", "pitch in one line (sdvplot is not a dependency, so this is not executed here):\n", "\n", "```python\n", "from sdvplot import pitch_coords\n", "xy = pitch_coords(events, provider='statsbomb') # R: sdvplotR::sdv_pitch_coords(events, 'statsbomb')\n", "```" ] }, { "cell_type": "markdown", "id": "spadl-md", "metadata": {}, "source": [ "## 🧭 SPADL actions\n", "\n", "`soccer_spadl(dataset)` converts any kloppy event dataset into\n", "[SPADL](https://socceraction.readthedocs.io) (Soccer Player Action Description Language): one\n", "row per on-the-ball action, with every action oriented to attack left to right on a 105 × 68\n", "pitch. `soccer_open_dataset(...)` is the dataset-returning twin of `soccer_open_events(...)`." ] }, { "cell_type": "code", "execution_count": null, "id": "spadl-code", "metadata": {}, "outputs": [], "source": [ "from sportsdataverse.soccer import soccer_open_dataset, soccer_spadl\n", "\n", "ds = safe('statsbomb open dataset 8658', lambda: soccer_open_dataset('statsbomb', 8658))\n", "actions = safe('SPADL actions', lambda: soccer_spadl(ds)) if ds is not None else None\n", "actions.head() if actions is not None else 'kloppy not installed (pip install \"sportsdataverse[soccer]\")'" ] }, { "cell_type": "markdown", "id": "xthreat-md", "metadata": {}, "source": [ "## Valuing actions with Expected Threat\n", "\n", "`soccer_xthreat_rate(actions)` appends an `xt_value` column: the change in\n", "[Expected Threat](https://karun.in/blog/expected-threat.html) from each successful move\n", "(pass, dribble, cross), using a grid bundled with the package and fit on StatsBomb open data.\n", "\n", "SPADL coordinates are already 105 × 68 (origin bottom-left), so pass the `start_x`/`start_y` columns straight to a 105 × 68 pitch -- sdvplot's `pitch_coords` needs no provider conversion." ] }, { "cell_type": "code", "execution_count": null, "id": "xthreat-code", "metadata": {}, "outputs": [], "source": [ "from sportsdataverse.soccer import soccer_xthreat_rate\n", "\n", "rated = safe('xT values', lambda: soccer_xthreat_rate(actions)) if actions is not None else None\n", "rated.head() if rated is not None else 'kloppy not installed (pip install \"sportsdataverse[soccer]\")'" ] }, { "cell_type": "markdown", "id": "soccer-next-md", "metadata": {}, "source": [ "## 🎉 Where to next\n", "\n", "- 📡 **Full function list** — every `espn_soccer_*` wrapper is\n", " documented in the **Soccer → Reference** section of the sidebar.\n", "- 🔗 **League aliases** — `from sportsdataverse.soccer import epl, mls, ucl,\n", " laliga, bundesliga, seriea, ligue1, uel, nwsl, ligamx, wc, wwc` — each\n", " pre-binds the `league=` slug so your code reads cleaner.\n", "- 🐼 Pass `return_as_pandas=True` to any parser, or call `.to_pandas()` on\n", " the polars frame — your favourite pandas / scikit-learn tooling works as-is.\n", "- ⚙️ For the raw ESPN payload (nested dict), call the wrapper without\n", " `parse_soccer_*` wrapping.\n", "- 🌐 R user? The closest companion for European football data is\n", " [worldfootballR](https://jaseziv.github.io/worldfootballR/).\n", " For ESPN basketball on the same platform, see\n", " [wehoop](https://wehoop.sportsdataverse.org) (WNBA/WBB) and\n", " [hoopR](https://hoopR.sportsdataverse.org) (NBA/MBB).\n", "- Part of the [SportsDataverse](https://py.sportsdataverse.org/docs/ecosystem)\n", " ecosystem.\n", "\n", "Now go find the next *gol de placa* — the data's all here. ⚽🌟" ] } ] }