{ "cells": [ { "cell_type": "markdown", "id": "216d4e41", "metadata": {}, "source": [ "# ๐Ÿ€ NBA hoops with `sportsdataverse-py`\n", "\n", "Welcome to the hardwood! ๐ŸŽ‰ In just a few lines of Python you're about to\n", "pull a whole season of NBA data โ€” **teams, standings, rosters, play-by-play,\n", "box scores, schedules and statistical leaders** โ€” straight from ESPN and the\n", "SportsDataverse data releases. Everything comes back as a tidy **polars**\n", "DataFrame that's ready to slice, model, and chart. ๐Ÿš€\n", "\n", "We lead with the richest surface in the package: the **`espn_nba_*`** family,\n", "backed by ESPN's site / web / core APIs. If you know the R package\n", "[hoopR](https://hoopR.sportsdataverse.org), these names will feel like home.\n", "Python neighbor for the raw NBA Stats endpoints:\n", "[nba_api](https://github.com/swar/nba_api). Let's lace 'em up! ๐Ÿ‘Ÿ" ] }, { "cell_type": "markdown", "id": "ab3c82b2", "metadata": {}, "source": [ "## ๐Ÿงฐ The toolbox\n", "\n", "Here's the kit we'll reach for. The **`espn_nba_*`** wrappers (โญ our\n", "premium source) hit ESPN live and parse the JSON into polars for you; the\n", "**`load_nba_*`** loaders pull pre-built season parquets from the\n", "[sportsdataverse-data](https://github.com/sportsdataverse/sportsdataverse-data)\n", "releases โ€” fast and reliable. Click any name for the full reference.\n", "\n", "| Function | What it gives you | Source |\n", "|---|---|---|\n", "| [`espn_nba_teams`](../nba/reference/additional.md#espn_nba_teams) | All 30 NBA teams (grab `team_id`s here) | โญ ESPN |\n", "| [`espn_nba_scoreboard`](../nba/reference/site.md#espn_nba_scoreboard) | A day's slate โ€” scores, status, matchups | โญ ESPN |\n", "| [`espn_nba_schedule`](../nba/reference/additional.md#espn_nba_schedule) | Schedule for a date / date-range | โญ ESPN |\n", "| [`espn_nba_standings`](../nba/reference/site.md#espn_nba_standings) | Conference standings (W-L, win%, streak) | โญ ESPN |\n", "| [`espn_nba_team_roster`](../nba/reference/site.md#espn_nba_team_roster) | A team's active roster | โญ ESPN |\n", "| [`espn_nba_team_schedule`](../nba/reference/site.md#espn_nba_team_schedule) | One team's full-season schedule | โญ ESPN |\n", "| [`espn_nba_player_gamelog`](../nba/reference/web.md#espn_nba_player_gamelog) | A player's game-by-game log | โญ ESPN |\n", "| [`espn_nba_leaders`](../nba/reference/web.md#espn_nba_leaders) | League statistical leaders | โญ ESPN |\n", "| [`espn_nba_pbp`](../nba/reference/additional.md) | Full game payload (play-by-play, win prob, box) | โญ ESPN |\n", "| [`espn_nba_game_rosters`](../nba/reference/additional.md) | Both teams' rosters for one game | โญ ESPN |\n", "| [`load_nba_schedule`](../nba/reference/loaders.md#load_nba_pbp) | Multi-season schedule parquet | ๐Ÿ“ฆ release |\n", "| [`load_nba_player_boxscore`](../nba/reference/loaders.md#load_nba_player_boxscore) | Player box scores, every game | ๐Ÿ“ฆ release |\n", "| [`load_nba_standings`](../nba/reference/loaders.md#load_nba_standings) | Historical standings | ๐Ÿ“ฆ release |\n", "| [`espn_nba_injuries`](../nba/reference/site.md#espn_nba_injuries) | League-wide injury report, one row per team | โญ ESPN |\n", "| [`load_nba_team_boxscore`](../nba/reference/loaders.md#load_nba_team_boxscore) | Team box scores, every game (off/def, shooting) | ๐Ÿ“ฆ release |\n", "| [`load_nba_shots`](../nba/reference/loaders.md#load_nba_shots) | Every made shot with court coordinates | ๐Ÿ“ฆ release |\n", "| [`most_recent_nba_season`](../nba/reference/additional.md#most_recent_nba_season) | The current season year helper | ๐Ÿงฎ util |\n" ] }, { "cell_type": "markdown", "id": "c3c04921", "metadata": {}, "source": [ "## ๐Ÿ”Œ Setup\n", "\n", "```sh\n", "pip install sportsdataverse\n", "```\n", "\n", "No API key needed โ€” ESPN's public endpoints and the data releases are open. ๐Ÿ˜Š" ] }, { "cell_type": "code", "execution_count": null, "id": "99910ca5", "metadata": {}, "outputs": [], "source": [ "import polars as pl\n", "import sportsdataverse as sdv\n", "from sportsdataverse.nba import most_recent_nba_season\n", "\n", "pl.Config.set_tbl_rows(8)\n", "SEASON = most_recent_nba_season()\n", "print('current NBA season:', SEASON)" ] }, { "cell_type": "markdown", "id": "3abdff10", "metadata": {}, "source": [ "ESPN endpoints are live and seasonal, so we'll route every network call\n", "through a tiny `safe()` helper. When the feed is up you get the frame; when\n", "it's mid-offseason or briefly rate-limited you get a friendly one-liner\n", "instead of a scary traceback. ๐Ÿ›Ÿ" ] }, { "cell_type": "code", "execution_count": null, "id": "1492e5ba", "metadata": {}, "outputs": [], "source": [ "from sportsdataverse.errors import AssetFetchError, NoDataError\n", "\n", "def safe(label, thunk):\n", " try:\n", " out = thunk()\n", " n = out.height if isinstance(out, pl.DataFrame) else (len(out) if hasattr(out, '__len__') else '?')\n", " print(f'โœ… {label} โ€” {n} rows')\n", " return out\n", " except (NoDataError, AssetFetchError) as e:\n", " print(f\"\\u23ed\\ufe0f {label}: {type(e).__name__}: {e}\")\n", " return None" ] }, { "cell_type": "markdown", "id": "fb10ec38", "metadata": {}, "source": [ "## ๐ŸŸ๏ธ Teams\n", "\n", "Start with [`espn_nba_teams`](../nba/reference/additional.md#espn_nba_teams) โ€”\n", "one wide row per franchise. The `team_id` column is the key you'll pass into\n", "roster, schedule and standings calls everywhere else." ] }, { "cell_type": "code", "execution_count": null, "id": "25e0fdbb", "metadata": {}, "outputs": [], "source": [ "teams = safe('teams', sdv.nba.espn_nba_teams)\n", "(teams.select(['team_id', 'team_location', 'team_name', 'team_abbreviation', 'team_color']).head(10)\n", " if teams is not None else 'teams feed unavailable')" ] }, { "cell_type": "markdown", "id": "7d10b463", "metadata": {}, "source": [ "## ๐Ÿ“… Today on the slate (scoreboard)\n", "\n", "[`espn_nba_scoreboard`](../nba/reference/site.md#espn_nba_scoreboard) returns a\n", "tidy frame of every game for a date โ€” final scores, live status, and\n", "matchups. Pass `dates='YYYYMMDD'` for one day. Here's a slice of the 2024\n", "NBA Finals opener." ] }, { "cell_type": "code", "execution_count": null, "id": "cd77e5cf", "metadata": {}, "outputs": [], "source": [ "sb = safe('scoreboard', lambda: sdv.nba.espn_nba_scoreboard(dates='20240606'))\n", "keep = ['game_id', 'short_name', 'home_abbreviation', 'away_abbreviation',\n", " 'home_score', 'away_score', 'status_type_detail']\n", "(sb.select([c for c in keep if c in sb.columns]).head()\n", " if sb is not None and sb.height else 'no games on that date')" ] }, { "cell_type": "markdown", "id": "e5bed44e", "metadata": {}, "source": [ "## ๐Ÿ† Standings\n", "\n", "[`espn_nba_standings`](../nba/reference/site.md#espn_nba_standings) gives one\n", "row per team with wins, losses, win%, point differential and current streak.\n", "Pass `season=` (the end year of the season)." ] }, { "cell_type": "code", "execution_count": null, "id": "82e76d76", "metadata": {}, "outputs": [], "source": [ "standings = safe('standings', lambda: sdv.nba.espn_nba_standings(season=SEASON))\n", "cols = ['team_display_name', 'wins', 'losses', 'win_percent', 'games_behind',\n", " 'point_differential', 'streak']\n", "(standings.select([c for c in cols if c in standings.columns])\n", " .sort('win_percent', descending=True).head(10)\n", " if standings is not None and standings.height else 'standings unavailable')" ] }, { "cell_type": "markdown", "id": "1cf7bbea", "metadata": {}, "source": [ "## ๐Ÿณ Cookbook: common NBA tasks\n", "\n", "Now the fun part โ€” a handful of recipes you'll reach for again and again.\n", "Each one leans on the premium `espn_nba_*` wrappers." ] }, { "cell_type": "markdown", "id": "3bc7017f", "metadata": {}, "source": [ "### Recipe 1 โ€” A team and its roster ๐Ÿ‘ฅ\n", "\n", "Grab a `team_id` from [`espn_nba_teams`](../nba/reference/additional.md#espn_nba_teams),\n", "then pull the active roster with\n", "[`espn_nba_team_roster`](../nba/reference/site.md#espn_nba_team_roster)." ] }, { "cell_type": "code", "execution_count": null, "id": "5cc45447", "metadata": {}, "outputs": [], "source": [ "tid = None\n", "if teams is not None and teams.height:\n", " # Boston Celtics if present, else the first team\n", " row = teams.filter(pl.col('team_abbreviation') == 'BOS')\n", " tid = int((row if row.height else teams)['team_id'][0])\n", "\n", "roster = safe(f'roster {tid}', lambda: sdv.nba.espn_nba_team_roster(team_id=tid)) if tid else None\n", "cols = ['full_name', 'jersey', 'position_abbreviation', 'height', 'weight', 'age']\n", "(roster.select([c for c in cols if c in roster.columns]).head(10)\n", " if roster is not None and roster.height else 'roster unavailable')" ] }, { "cell_type": "markdown", "id": "ff4b839e", "metadata": {}, "source": [ "### Recipe 2 โ€” One team's season schedule ๐Ÿ“†\n", "\n", "[`espn_nba_team_schedule`](../nba/reference/site.md#espn_nba_team_schedule)\n", "returns every game on a team's calendar for a season โ€” perfect for building\n", "a results table or a strength-of-schedule view." ] }, { "cell_type": "code", "execution_count": null, "id": "9982e923", "metadata": {}, "outputs": [], "source": [ "tsched = safe(f'team schedule {tid}',\n", " lambda: sdv.nba.espn_nba_team_schedule(team_id=tid, season=SEASON)) if tid else None\n", "cols = ['id', 'date', 'name', 'short_name', 'season_year']\n", "(tsched.select([c for c in cols if c in tsched.columns]).head()\n", " if tsched is not None and tsched.height else 'team schedule unavailable')" ] }, { "cell_type": "markdown", "id": "01da4ca6", "metadata": {}, "source": [ "### Recipe 3 โ€” A player's game log โ›น๏ธ\n", "\n", "[`espn_nba_player_gamelog`](../nba/reference/web.md#espn_nba_player_gamelog)\n", "returns a game-by-game stat line for one athlete. The `stat_*` columns are\n", "positional (the ordered ESPN box categories); pair them with the opponent\n", "and result columns to see how the night went. (`1966` = LeBron James.)" ] }, { "cell_type": "code", "execution_count": null, "id": "5dbadafc", "metadata": {}, "outputs": [], "source": [ "gamelog = safe('LeBron gamelog',\n", " lambda: sdv.nba.espn_nba_player_gamelog(athlete_id=1966, season=SEASON))\n", "cols = ['event_date', 'opponent_abbreviation', 'home_away', 'game_result', 'score',\n", " 'stat_0', 'stat_1', 'stat_2']\n", "(gamelog.select([c for c in cols if c in gamelog.columns]).head()\n", " if gamelog is not None and gamelog.height else 'gamelog unavailable')" ] }, { "cell_type": "markdown", "id": "780fe101", "metadata": {}, "source": [ "### Recipe 4 โ€” Top scorers from the box-score release ๐Ÿฅ‡\n", "\n", "For a whole-season leaderboard the\n", "[`load_nba_player_boxscore`](../nba/reference/loaders.md#load_nba_player_boxscore)\n", "release is your friend โ€” it's a fast parquet download, no live API needed.\n", "Here we average points per game and rank the top 10 scorers." ] }, { "cell_type": "code", "execution_count": null, "id": "1bdf48da", "metadata": {}, "outputs": [], "source": [ "box = safe('player boxscore release', lambda: sdv.nba.load_nba_player_boxscore(seasons=[SEASON]))\n", "if box is not None and box.height:\n", " leaders = (\n", " box.filter(pl.col('minutes') > 0)\n", " .group_by(['athlete_display_name', 'team_abbreviation'])\n", " .agg(pl.len().alias('gp'),\n", " pl.col('points').mean().round(1).alias('ppg'),\n", " pl.col('rebounds').mean().round(1).alias('rpg'),\n", " pl.col('assists').mean().round(1).alias('apg'))\n", " .filter(pl.col('gp') >= 20)\n", " .sort('ppg', descending=True)\n", " .head(10)\n", " )\n", " out = leaders\n", "else:\n", " out = 'box-score release unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "ce5cee76", "metadata": {}, "source": [ "### Recipe 5 โ€” Offense vs defense, every team ๐Ÿ›ก๏ธ\n", "\n", "The [`load_nba_team_boxscore`](../nba/reference/loaders.md#load_nba_team_boxscore)\n", "release has one row per team-game with both `team_score` and\n", "`opponent_team_score` โ€” so points-for, points-against and net rating are a\n", "single `group_by` away." ] }, { "cell_type": "code", "execution_count": null, "id": "88b96c46", "metadata": {}, "outputs": [], "source": [ "tbox = safe('team boxscore release', lambda: sdv.nba.load_nba_team_boxscore(seasons=[SEASON]))\n", "if tbox is not None and tbox.height:\n", " netrtg = (\n", " tbox.group_by('team_abbreviation')\n", " .agg(pl.len().alias('gp'),\n", " pl.col('team_score').mean().round(1).alias('off_ppg'),\n", " pl.col('opponent_team_score').mean().round(1).alias('def_ppg'))\n", " .with_columns((pl.col('off_ppg') - pl.col('def_ppg')).round(1).alias('net'))\n", " .sort('net', descending=True)\n", " .head(10)\n", " )\n", " out = netrtg\n", "else:\n", " out = 'team box-score release unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "45fd23a9", "metadata": {}, "source": [ "### Recipe 6 โ€” Who lived behind the arc? ๐ŸŽฏ\n", "\n", "Sum makes and attempts across the season to get each team's true\n", "three-point percentage (game-level percentages can't just be averaged).\n", "Reuses the `tbox` frame from Recipe 5 โ€” no second download." ] }, { "cell_type": "code", "execution_count": null, "id": "93c840e2", "metadata": {}, "outputs": [], "source": [ "if tbox is not None and tbox.height:\n", " three_pt = (\n", " tbox.group_by('team_abbreviation')\n", " .agg(pl.col('three_point_field_goals_made').sum().alias('made'),\n", " pl.col('three_point_field_goals_attempted').sum().alias('att'))\n", " .with_columns((100 * pl.col('made') / pl.col('att')).round(1).alias('three_pt_pct'))\n", " .filter(pl.col('att') > 0)\n", " .sort('three_pt_pct', descending=True)\n", " .head(10)\n", " )\n", " out = three_pt\n", "else:\n", " out = 'team box-score release unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "22da2b4c", "metadata": {}, "source": [ "### Recipe 7 โ€” Double-double machines ๐Ÿ’ช\n", "\n", "A *double-double* is double digits in two of points / rebounds / assists.\n", "Count the categories per player-game, keep the ones that cleared two, then\n", "tally them up โ€” straight from\n", "[`load_nba_player_boxscore`](../nba/reference/loaders.md#load_nba_player_boxscore)." ] }, { "cell_type": "code", "execution_count": null, "id": "4feca44b", "metadata": {}, "outputs": [], "source": [ "pbox = safe('player boxscore release', lambda: sdv.nba.load_nba_player_boxscore(seasons=[SEASON]))\n", "if pbox is not None and pbox.height:\n", " dd = (\n", " pbox.filter(pl.col('minutes') > 0)\n", " .with_columns(\n", " ((pl.col('points') >= 10).cast(pl.Int8)\n", " + (pl.col('rebounds') >= 10).cast(pl.Int8)\n", " + (pl.col('assists') >= 10).cast(pl.Int8)).alias('cats10'))\n", " .filter(pl.col('cats10') >= 2)\n", " .group_by(['athlete_display_name', 'team_abbreviation'])\n", " .agg(pl.len().alias('double_doubles'))\n", " .sort('double_doubles', descending=True)\n", " .head(10)\n", " )\n", " out = dd\n", "else:\n", " out = 'player box-score release unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "bc4023ff", "metadata": {}, "source": [ "### Recipe 8 โ€” A tidy standings table ๐Ÿ†\n", "\n", "The [`load_nba_standings`](../nba/reference/loaders.md#load_nba_standings)\n", "release ships in **long** format (one row per team ร— stat). Pivot the stats\n", "you care about into columns to get a classic standings grid, sorted by\n", "win percentage." ] }, { "cell_type": "code", "execution_count": null, "id": "ef355e16", "metadata": {}, "outputs": [], "source": [ "stload = safe('standings release', lambda: sdv.nba.load_nba_standings(seasons=[SEASON]))\n", "wanted = ['wins', 'losses', 'winPercent', 'playoffSeed', 'pointDifferential']\n", "if stload is not None and stload.height and {'stat_name', 'value'}.issubset(stload.columns):\n", " table = (\n", " stload.filter(pl.col('stat_name').is_in(wanted))\n", " .select(['team_abbreviation', 'group_name', 'stat_name', 'value'])\n", " .pivot(values='value', index=['team_abbreviation', 'group_name'], on='stat_name')\n", " .sort('winPercent', descending=True)\n", " .head(12)\n", " )\n", " out = table\n", "else:\n", " out = 'standings release unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "5d98ec88", "metadata": {}, "source": [ "### Recipe 9 โ€” Built on threes (shot release + a join) ๐Ÿงฑ\n", "\n", "[`load_nba_shots`](../nba/reference/loaders.md#load_nba_shots) is one row per\n", "made shot with a `score_value`. Tally points from twos vs threes per team,\n", "then **join** team abbreviations from the box-score release to find who\n", "leaned hardest on the long ball." ] }, { "cell_type": "code", "execution_count": null, "id": "10cb7204", "metadata": {}, "outputs": [], "source": [ "shots = safe('shots release', lambda: sdv.nba.load_nba_shots(seasons=[SEASON]))\n", "if shots is not None and shots.height and tbox is not None and tbox.height:\n", " fg = shots.filter(pl.col('score_value').is_in([2, 3]))\n", " reliance = (\n", " fg.group_by('team_id')\n", " .agg(pl.col('score_value').filter(pl.col('score_value') == 3).len().alias('threes_made'),\n", " pl.col('score_value').sum().alias('points_from_fg'))\n", " .with_columns((3 * pl.col('threes_made')).alias('points_from_threes'))\n", " .with_columns((100 * pl.col('points_from_threes') / pl.col('points_from_fg'))\n", " .round(1).alias('pct_pts_from_3'))\n", " .filter(pl.col('threes_made') >= 500) # drop All-Star / special rosters\n", " )\n", " abbr = tbox.select(['team_id', 'team_abbreviation']).unique()\n", " out = (reliance.join(abbr, on='team_id', how='left')\n", " .select(['team_abbreviation', 'threes_made', 'pct_pts_from_3'])\n", " .sort('pct_pts_from_3', descending=True).head(10))\n", "else:\n", " out = 'shots / team box-score release unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "6d52f66d", "metadata": {}, "source": [ "### Recipe 10 โ€” Head-to-head, game by game ๐Ÿค\n", "\n", "Filter the team box-score release to one matchup and you get the full\n", "season series โ€” every meeting, the score, and who won. Swap the two\n", "abbreviations for any rivalry you like." ] }, { "cell_type": "code", "execution_count": null, "id": "61e53b26", "metadata": {}, "outputs": [], "source": [ "TEAM_A, TEAM_B = 'BOS', 'NY'\n", "if tbox is not None and tbox.height and 'opponent_team_abbreviation' in tbox.columns:\n", " series = (\n", " tbox.filter((pl.col('team_abbreviation') == TEAM_A)\n", " & (pl.col('opponent_team_abbreviation') == TEAM_B))\n", " .select([c for c in ['game_date', 'team_home_away', 'team_score',\n", " 'opponent_team_score', 'team_winner']\n", " if c in tbox.columns])\n", " .sort('game_date')\n", " )\n", " out = series if series.height else f'no {TEAM_A} vs {TEAM_B} games in {SEASON}'\n", "else:\n", " out = 'team box-score release unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "8b08a3a2", "metadata": {}, "source": [ "### Recipe 11 โ€” Who's banged up? ๐Ÿฉน (pandas interop)\n", "\n", "[`espn_nba_injuries`](../nba/reference/site.md#espn_nba_injuries) hits ESPN\n", "live for the league-wide injury report. Ask for a **pandas** frame with\n", "`return_as_pandas=True` (a handy interop point), count the listed players\n", "per team, then hand the result back to polars for the final sort." ] }, { "cell_type": "code", "execution_count": null, "id": "280d4a3c", "metadata": {}, "outputs": [], "source": [ "import ast\n", "\n", "inj = safe('injuries', lambda: sdv.nba.espn_nba_injuries(return_as_pandas=True))\n", "if inj is not None and getattr(inj, 'shape', (0,))[0] and 'injuries' in inj.columns:\n", " def _n_listed(s):\n", " try:\n", " v = ast.literal_eval(s) if isinstance(s, str) else s\n", " return len(v) if isinstance(v, list) else 0\n", " except Exception:\n", " return 0\n", " inj = inj.copy()\n", " inj['players_listed'] = inj['injuries'].apply(_n_listed)\n", " out = (pl.from_pandas(inj[['display_name', 'players_listed']])\n", " .filter(pl.col('players_listed') > 0)\n", " .sort('players_listed', descending=True)\n", " .head(12))\n", "else:\n", " out = 'injury report unavailable (off-season or feed down)'\n", "out" ] }, { "cell_type": "markdown", "id": "13324c9f", "metadata": {}, "source": [ "## ๐ŸŽฌ Play-by-play & game rosters\n", "\n", "Now for the granular stuff. [`espn_nba_pbp`](../nba/reference/additional.md)\n", "returns the **whole game payload** as a dict โ€” play-by-play, win probability,\n", "box score, and header โ€” keyed by `game_id` (an ESPN event id). Pair it with\n", "[`espn_nba_game_rosters`](../nba/reference/additional.md) for who actually\n", "suited up.\n", "\n", "We'll use Game 1 of the 2024 Finals (`game_id=401585660`)." ] }, { "cell_type": "code", "execution_count": null, "id": "3e88943f", "metadata": {}, "outputs": [], "source": [ "GAME_ID = 401585660\n", "pbp = safe('pbp payload', lambda: sdv.nba.espn_nba_pbp(game_id=GAME_ID))\n", "(list(pbp.keys())[:8] if isinstance(pbp, dict) else 'pbp unavailable')" ] }, { "cell_type": "code", "execution_count": null, "id": "40d0ebb1", "metadata": {}, "outputs": [], "source": [ "plays = (pl.DataFrame(pbp['plays'], infer_schema_length=None)\n", " if isinstance(pbp, dict) and pbp.get('plays') else None)\n", "cols = ['period.number', 'clock.displayValue', 'text', 'homeScore', 'awayScore', 'scoringPlay']\n", "(plays.select([c for c in cols if c in plays.columns]).head()\n", " if plays is not None and plays.height else 'no plays parsed')" ] }, { "cell_type": "markdown", "id": "879a80fd", "metadata": {}, "source": [ "### Slice it: every 3-pointer in the game ๐ŸŽฏ\n", "\n", "The `plays` frame is just polars โ€” so a scoring slice is one filter away.\n", "Here we pull made three-pointers in chronological order." ] }, { "cell_type": "code", "execution_count": null, "id": "e5158032", "metadata": {}, "outputs": [], "source": [ "if plays is not None and plays.height:\n", " threes = (\n", " plays.filter(pl.col('scoringPlay') == True)\n", " .filter(pl.col('text').str.contains('(?i)three point|3pt|three-point'))\n", " .select([c for c in ['period.number', 'clock.displayValue', 'text',\n", " 'homeScore', 'awayScore'] if c in plays.columns])\n", " )\n", " out = threes.head(10) if threes.height else 'no three-pointers matched the text filter'\n", "else:\n", " out = 'no plays to slice'\n", "out" ] }, { "cell_type": "markdown", "id": "1ce614d2", "metadata": {}, "source": [ "### Who played? Game rosters ๐Ÿ“‹\n", "\n", "[`espn_nba_game_rosters`](../nba/reference/additional.md) returns both teams'\n", "rosters for a single game, one row per athlete โ€” including the `starter`\n", "flag and jersey number." ] }, { "cell_type": "code", "execution_count": null, "id": "3e67195d", "metadata": {}, "outputs": [], "source": [ "grosters = safe('game rosters', lambda: sdv.nba.espn_nba_game_rosters(game_id=GAME_ID))\n", "cols = ['athlete_display_name', 'team_abbreviation', 'starter', 'jersey', 'position_name']\n", "(grosters.select([c for c in cols if c in grosters.columns]).head(10)\n", " if grosters is not None and grosters.height else 'game rosters unavailable')" ] }, { "cell_type": "markdown", "id": "6830ad0c", "metadata": {}, "source": [ "## ๐Ÿ“ฆ Bulk season data with the loaders\n", "\n", "When you want *everything* for a season at once โ€” not one game at a time โ€”\n", "the `load_nba_*` loaders pull pre-built parquet releases. They're fast,\n", "reliable, and don't depend on a live API being up.\n", "\n", "| Loader | Grain |\n", "|---|---|\n", "| [`load_nba_schedule`](../nba/reference/loaders.md#load_nba_schedule) | one row per game |\n", "| [`load_nba_player_boxscore`](../nba/reference/loaders.md#load_nba_player_boxscore) | one row per player-game |\n", "| [`load_nba_standings`](../nba/reference/loaders.md#load_nba_standings) | one row per team-season |" ] }, { "cell_type": "code", "execution_count": null, "id": "1f2d0e05", "metadata": {}, "outputs": [], "source": [ "sched = safe('schedule release', lambda: sdv.nba.load_nba_schedule(seasons=[SEASON]))\n", "cols = ['id', 'date', 'home_display_name', 'away_display_name', 'home_score', 'away_score']\n", "(sched.select([c for c in cols if c in sched.columns]).head()\n", " if sched is not None and sched.height else 'schedule release unavailable')" ] }, { "cell_type": "markdown", "id": "0628b87b", "metadata": {}, "source": [ "### Pipeline: the highest-scoring games of the season ๐Ÿ”ฅ\n", "\n", "With the schedule release in hand, a combined-points leaderboard is a quick\n", "polars pipeline โ€” cast the scores, sum them, sort descending." ] }, { "cell_type": "markdown", "id": "624ecf3c", "source": "## stats.nba.com surface (`nba_stats_*`)\n\nBeyond the ESPN wrappers, `sportsdataverse.nba` also ships **112 wrappers** for the official `stats.nba.com` API โ€” the same source powering [nba_api](https://github.com/swar/nba_api). The module is `nba_stats` and every function is named `nba_stats_`. These are the **capture-confirmed live, non-deprecated** endpoints; deprecated and never-populated endpoints are excluded (see `ENDPOINT_HEALTH.md` in `sdv-internal-refs` for the full active/dying/barren/dead matrix).\n\n**League routing** is a single `league_id` parameter: `\"00\"` โ†’ NBA, `\"20\"` โ†’ G-League, `\"15\"` โ†’ Summer League. One stem, one host, league chosen per call.\n\n**curl_cffi required for live calls:** `stats.nba.com` uses TLS/JA3 fingerprinting that silently blocks plain `requests`. Install the optional transport: `pip install \"sportsdataverse[all]\"` (or `pip install curl_cffi`). The wrappers run fully offline when fed a fixture โ€” live calls simply need the extra installed.\n\nWrappers return a tidy **polars DataFrame** by default (parsed via the generic `parse_nba_stats_result_sets` parser โ€” which also flattens the shot-location endpoints' grouped 2-level headers into composite columns and unrolls the `scoreboardv3` game feed). Pass `return_parsed=False` for the raw `Dict` or `return_as_pandas=True` for pandas. Multi-result-set endpoints (e.g. `playercareerstats`) return `dict[str, DataFrame]`.", "metadata": {} }, { "cell_type": "code", "id": "bb98203c", "source": "from sportsdataverse.nba import nba_stats\n\n# NBA player dashboard (polars DataFrame by default)\n# Live calls require: pip install \"sportsdataverse[all]\" (curl_cffi transport)\ndf = safe('nba_stats_leaguedashplayerstats (NBA)',\n lambda: nba_stats.nba_stats_leaguedashplayerstats(league_id=\"00\"))\n\n# Same endpoint, other leagues via league_id on the SAME stem:\ngleague = safe('nba_stats_leaguedashplayerstats (G-League)',\n lambda: nba_stats.nba_stats_leaguedashplayerstats(league_id=\"20\"))\nsummer = safe('nba_stats_leaguedashplayerstats (Summer League)',\n lambda: nba_stats.nba_stats_leaguedashplayerstats(league_id=\"15\"))\n\n# Multi result-set endpoint -> dict[str, polars.DataFrame]\ncareer = safe('nba_stats_playercareerstats (LeBron)',\n lambda: nba_stats.nba_stats_playercareerstats(player_id=\"2544\"))\n\nprint('career result sets:', list(career.keys()) if isinstance(career, dict) else career)\ndf", "metadata": {}, "execution_count": null, "outputs": [] }, { "cell_type": "code", "execution_count": null, "id": "4f8fd4c5", "metadata": {}, "outputs": [], "source": [ "if sched is not None and sched.height and {'home_score', 'away_score'}.issubset(sched.columns):\n", " hot = (\n", " sched.with_columns(\n", " (pl.col('home_score').cast(pl.Int64, strict=False)\n", " + pl.col('away_score').cast(pl.Int64, strict=False)).alias('total_points')\n", " )\n", " .filter(pl.col('total_points').is_not_null())\n", " .sort('total_points', descending=True)\n", " .select([c for c in ['date', 'home_display_name', 'away_display_name',\n", " 'home_score', 'away_score', 'total_points'] if c in sched.columns])\n", " .head(10)\n", " )\n", " out = hot\n", "else:\n", " out = 'schedule release unavailable'\n", "out" ] }, { "cell_type": "markdown", "id": "81e62b8a", "metadata": {}, "source": [ "## ๐ŸŽ‰ Where to next\n", "\n", "You just toured the **premium `espn_nba_*` surface** plus the season\n", "loaders โ€” teams, scoreboard, standings, rosters, schedules, player game\n", "logs, play-by-play, and bulk box scores. A few parting tips:\n", "\n", "- Pass `return_as_pandas=True` to any wrapper for a pandas frame instead of polars.\n", "- ESPN `espn_nba_*` wrappers also accept `return_parsed=False` for the raw JSON dict.\n", "- Full reference lives in the **NBA** section of the sidebar:\n", " [ESPN site API](../nba/reference/site.md) ยท\n", " [ESPN web API](../nba/reference/web.md) ยท\n", " [ESPN core API](../nba/reference/core.md) ยท\n", " [additional functions](../nba/reference/additional.md) ยท\n", " [loaders](../nba/reference/loaders.md)\n", "- R user? The same surface lives in [hoopR](https://hoopR.sportsdataverse.org).\n", "- Need raw NBA Stats endpoints? See [nba_api](https://github.com/swar/nba_api).\n", "\n", "Now go break down some film โ€” and may your jumper always find the bottom of\n", "the net! ๐Ÿ€๐Ÿ”ฅ" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## ๐Ÿฆ“ Officiating\n", "\n", "Last Two Minute reports and referee assignments, from `official.nba.com` and the\n", "NBA's own feeds.\n" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "from sportsdataverse.nba import nba_officiating\n", "\n", "print(\"nba_officiating:\", [n for n in dir(nba_officiating) if not n.startswith(\"_\")][:8])\n", "safe(\"referee assignments\", lambda: nba_officiating.nba_referee_assignments(game_date=\"2024-04-01\"))\n", "safe(\"L2M games\", lambda: nba_officiating.nba_l2m_games())\n" ] } ], "metadata": { "kernelspec": { "display_name": "Python 3", "language": "python", "name": "python3" }, "language_info": { "name": "python" } }, "nbformat": 4, "nbformat_minor": 5 }