{ "cells": [ { "cell_type": "markdown", "id": "2a16c70c", "metadata": {}, "source": [ "# 🏈 College football with `sportsdataverse-py`\n", "\n", "Saturdays in autumn, condensed into tidy DataFrames. πŸ‚ In a few lines of\n", "Python you're about to pull **a decade of play-by-play**, full **rosters**,\n", "**schedules**, **team info**, plus **live** ESPN scoreboards, standings,\n", "polls and recruiting boards β€” all as clean **polars** frames ready to model.\n", "\n", "CFB has no single native premium API, so our **premium path** is two-pronged:\n", "\n", "1. πŸ—„οΈ **Release loaders** (`load_cfb_*`) β€” pre-built, EPA/WPA-enriched\n", " datasets served straight from the\n", " [cfbfastR-data](https://github.com/sportsdataverse/cfbfastR-data) GitHub\n", " release. Fast, reliable, no key needed.\n", "2. πŸ“‘ **ESPN families** (`espn_cfb_*`) β€” live scoreboards, team pages,\n", " standings, rankings, recruiting and per-play participants.\n", "\n", "R user? Every verb here has a twin in\n", "[cfbfastR](https://cfbfastR.sportsdataverse.org). Let's kick off! 🏈" ] }, { "cell_type": "markdown", "id": "2d72edc7", "metadata": {}, "source": [ "## 🧰 The toolbox\n", "\n", "Everything returns a tidy **polars** `DataFrame` by default β€” pass\n", "`return_as_pandas=True` for pandas, or (on the `espn_cfb_*` wrappers)\n", "`return_parsed=False` for the raw JSON. ⭐ marks the **premium** path.\n", "\n", "| Function | What it gives you | Source |\n", "|---|---|---|\n", "| [`load_cfb_pbp`](../cfb/reference/loaders.md#load_cfb_pbp) | Full **play-by-play** with EPA/WPA, since 2003 | ⭐ release |\n", "| [`load_cfb_rosters`](../cfb/reference/loaders.md#load_cfb_rosters) | Season **rosters** (bio, position, hometown) | ⭐ release |\n", "| [`load_cfb_schedule`](../cfb/reference/loaders.md#load_cfb_schedule) | Season **schedule** + results | ⭐ release |\n", "| [`load_cfb_team_info`](../cfb/reference/loaders.md#load_cfb_team_info) | **Team** metadata: conference, colors, venue | ⭐ release |\n", "| [`load_cfb_ratings`](../cfb/reference/loaders.md#load_cfb_ratings) | Opponent-adjusted **EPA ratings** per team-season | ⭐ release |\n", "| [`load_cfb_betting_lines`](../cfb/reference/additional.md#load_cfb_betting_lines) | Historical **betting market** lines (spread/total/ML) | ⭐ release |\n", "| [`espn_cfb_scoreboard`](../cfb/reference/site.md#espn_cfb_scoreboard) | Live + recent **scoreboard** for a date/week | ⭐ ESPN |\n", "| [`espn_cfb_schedule`](../cfb/reference/additional.md#espn_cfb_schedule) | ESPN **schedule** frame for a date/week | ⭐ ESPN |\n", "| [`espn_cfb_teams`](../cfb/reference/additional/other.md#espn_cfb_teams) | Every FBS/FCS **team** (grab `team_id`s) | ⭐ ESPN |\n", "| [`espn_cfb_team_roster`](../cfb/reference/site.md#espn_cfb_team_roster) | One team's **roster** | ⭐ ESPN |\n", "| [`espn_cfb_team_schedule`](../cfb/reference/site.md#espn_cfb_team_schedule) | One team's **schedule** | ⭐ ESPN |\n", "| [`espn_cfb_standings`](../cfb/reference/site.md#espn_cfb_standings) | Conference / division **standings** | ⭐ ESPN |\n", "| [`espn_cfb_rankings`](../cfb/reference/site.md#espn_cfb_rankings) | AP / Coaches / CFP **polls** | ⭐ ESPN |\n", "| [`espn_cfb_leaders`](../cfb/reference/web.md#espn_cfb_leaders) | League **stat leaders** by category | ⭐ ESPN |\n", "| [`espn_cfb_recruits`](../cfb/reference/core.md#espn_cfb_recruits) | Season **recruiting** class | ⭐ ESPN |\n", "| [`espn_cfb_play_participants`](../cfb/reference/additional/other.md#espn_cfb_play_participants) | Per-play **athletes** (passer/rusher/tackler…) | ⭐ ESPN |\n", "| [`CFBPlayProcess`](../cfb/reference/additional.md#CFBPlayProcess) | Full ESPN **PBP pipeline** (EPA/WPA + box) | ⭐ ESPN |\n", "| [`most_recent_cfb_season`](../cfb/reference/additional.md#most_recent_cfb_season) | The current season year helper | helper |\n" ] }, { "cell_type": "markdown", "id": "f2620c5a", "metadata": {}, "source": [ "## πŸ”Œ Setup\n", "\n", "```sh\n", "pip install sportsdataverse\n", "```\n", "\n", "No API key required. The `load_cfb_*` loaders read public parquet from the\n", "cfbfastR-data release, and the `espn_cfb_*` wrappers hit ESPN's public\n", "endpoints." ] }, { "cell_type": "code", "execution_count": null, "id": "eb5e47be", "metadata": {}, "outputs": [], "source": [ "import polars as pl\n", "import sportsdataverse as sdv\n", "from sportsdataverse.cfb import most_recent_cfb_season\n", "\n", "SEASON = most_recent_cfb_season()\n", "print('most recent CFB season:', SEASON)" ] }, { "cell_type": "markdown", "id": "4c6f3709", "metadata": {}, "source": [ "ESPN's live endpoints are seasonal and occasionally rate-limited, so a tiny\n", "`safe()` helper runs the riskier calls defensively β€” you get the frame when\n", "the feed is up, and a friendly one-liner when it isn't (never a scary\n", "traceback). The release loaders are reliable, so we call those directly. πŸ›Ÿ" ] }, { "cell_type": "code", "execution_count": null, "id": "9eb295b0", "metadata": {}, "outputs": [], "source": [ "from sportsdataverse.errors import AssetFetchError, NoDataError\n", "\n", "def safe(label, thunk):\n", " try:\n", " out = thunk()\n", " print(f'βœ… {label}')\n", " return out\n", " except (NoDataError, AssetFetchError) as e:\n", " print(f\"\\u23ed\\ufe0f {label}: {type(e).__name__}: {e}\")\n", " return None" ] }, { "cell_type": "markdown", "id": "9e44ed76", "metadata": {}, "source": [ "## πŸ—„οΈ Premium loaders: a whole season in one call\n", "\n", "The `load_cfb_*` family is the fastest way to get **clean, complete**\n", "season data. Each takes a `seasons=` int or list (β‰₯ 2003) and returns one\n", "tidy frame. Let's start with the schedule β€” one row per game, with final\n", "scores, conference flags, and playoff fields baked in.\n", "\n", "| Function | Grain | Highlights |\n", "|---|---|---|\n", "| [`load_cfb_schedule`](../cfb/reference/loaders.md#load_cfb_schedule) | one row / game | scores, neutral-site & conference flags, playoff rounds |\n" ] }, { "cell_type": "code", "execution_count": null, "id": "91fda377", "metadata": {}, "outputs": [], "source": [ "schedule = sdv.cfb.load_cfb_schedule(seasons=[2023])\n", "print('schedule shape:', schedule.shape)\n", "schedule.select([\n", " 'game_id', 'week', 'home_team', 'away_team',\n", " 'home_points', 'away_points', 'home_conference', 'neutral_site',\n", "]).head()" ] }, { "cell_type": "markdown", "id": "b2070da9", "metadata": {}, "source": [ "## πŸ‘₯ Premium loaders: rosters\n", "\n", "[`load_cfb_rosters`](../cfb/reference/loaders.md#load_cfb_rosters) gives you\n", "every listed player for a season β€” name, position, jersey, physicals and\n", "hometown. Perfect for joining onto play-by-play or building depth tables." ] }, { "cell_type": "code", "execution_count": null, "id": "ee66dd2a", "metadata": {}, "outputs": [], "source": [ "rosters = sdv.cfb.load_cfb_rosters(seasons=[2023])\n", "print('rosters shape:', rosters.shape)\n", "rosters.select([\n", " 'athlete_id', 'first_name', 'last_name', 'team_location',\n", " 'position', 'jersey', 'cfbd_home_state',\n", "]).head()" ] }, { "cell_type": "markdown", "id": "a3f9b408", "metadata": {}, "source": [ "## 🏟️ Premium loaders: team info\n", "\n", "[`load_cfb_team_info`](../cfb/reference/loaders.md#load_cfb_team_info)\n", "carries the reference metadata you'll want to label every chart: school\n", "name, conference, classification (FBS/FCS), team colors, and venue." ] }, { "cell_type": "code", "execution_count": null, "id": "80cfee9d", "metadata": {}, "outputs": [], "source": [ "team_info = sdv.cfb.load_cfb_team_info(seasons=[2023])\n", "print('team_info shape:', team_info.shape)\n", "team_info.select([\n", " 'team_id', 'school', 'conference', 'classification',\n", " 'venue_name', 'city', 'state', 'dome',\n", "]).head()" ] }, { "cell_type": "markdown", "id": "6ea68087", "metadata": {}, "source": [ "## 🎬 Premium loaders: play-by-play with EPA\n", "\n", "The crown jewel. [`load_cfb_pbp`](../cfb/reference/loaders.md#load_cfb_pbp)\n", "returns **every play** of a season with hundreds of engineered columns β€”\n", "down & distance, win probability, and **Expected Points Added (EPA)**\n", "already computed. (It's a big pull, so we grab a single season and peek.) πŸ“Š" ] }, { "cell_type": "code", "execution_count": null, "id": "20cef1ee", "metadata": {}, "outputs": [], "source": [ "# The release serves PBP for whichever seasons are currently published.\n", "# Try a few recent-ish seasons and keep the first one that comes back full,\n", "# so the EPA recipes below always have real plays to chew on.\n", "pbp = pl.DataFrame()\n", "for yr in (2023, 2022, 2021, 2020):\n", " cand = safe(f'load_cfb_pbp {yr}', lambda yr=yr: sdv.cfb.load_cfb_pbp(seasons=[yr]))\n", " if cand is not None and cand.width > 0 and cand.height > 0:\n", " pbp, PBP_SEASON = cand, yr\n", " break\n", "else:\n", " PBP_SEASON = None\n", "\n", "print('pbp season:', PBP_SEASON, '| pbp shape:', pbp.shape)\n", "cols = ['game_id', 'start.pos_team.name', 'down', 'distance',\n", " 'play_type', 'EPA', 'wpa']\n", "have = [c for c in cols if c in pbp.columns]\n", "pbp.select(have).head() if have else 'pbp not published for these seasons right now'" ] }, { "cell_type": "markdown", "id": "986d7596", "metadata": {}, "source": [ "## πŸ“‘ Live from ESPN: the scoreboard\n", "\n", "When you need *today's* slate (or a specific date), the ESPN wrappers shine.\n", "[`espn_cfb_scoreboard`](../cfb/reference/site.md#espn_cfb_scoreboard) takes a\n", "`dates=YYYYMMDD` (or season year) and returns the games on the board. We wrap\n", "it in `safe()` since live endpoints can be quiet in the offseason." ] }, { "cell_type": "code", "execution_count": null, "id": "5c34fac8", "metadata": {}, "outputs": [], "source": [ "board = safe(\n", " 'ESPN scoreboard',\n", " lambda: sdv.cfb.espn_cfb_scoreboard(dates=20231125), # rivalry Saturday\n", ")\n", "if board is not None and getattr(board, 'height', 0):\n", " keep = [c for c in board.columns\n", " if c in ('game_id', 'name', 'short_name', 'status_type_description',\n", " 'home_team_abbreviation', 'away_team_abbreviation')]\n", " out = board.select(keep).head() if keep else board.head()\n", "else:\n", " out = 'no games on the board for that date'\n", "out" ] }, { "cell_type": "markdown", "id": "b2e90963", "metadata": {}, "source": [ "## 🏫 Live from ESPN: teams (and their `team_id`s)\n", "\n", "[`espn_cfb_teams`](../cfb/reference/additional/other.md#espn_cfb_teams) lists every\n", "team in a division (`groups=80` FBS, `groups=81` FCS). The `team_id` column\n", "is the key you feed into every team-scoped ESPN call below." ] }, { "cell_type": "code", "execution_count": null, "id": "8e0f31f4", "metadata": {}, "outputs": [], "source": [ "teams = safe('ESPN teams', sdv.cfb.espn_cfb_teams)\n", "if teams is not None and teams.height:\n", " cols = [c for c in ('team_id', 'team_location', 'team_name',\n", " 'team_abbreviation') if c in teams.columns]\n", " out = teams.select(cols).head(8)\n", "else:\n", " out = 'teams unavailable right now'\n", "out" ] }, { "cell_type": "markdown", "id": "aafefcf2", "metadata": {}, "source": [ "## 🍳 Cookbook: common CFB tasks\n", "\n", "Now the fun part β€” real questions, answered with a few expressions. The\n", "loaders are reliable so these recipes lean on them, reaching for ESPN where\n", "it adds something live." ] }, { "cell_type": "markdown", "id": "15091002", "metadata": {}, "source": [ "### Recipe 1 β€” Highest-scoring games of the season πŸ”₯\n", "\n", "Straight from the loaded schedule: add the two scores and sort. No casting\n", "needed β€” the release frame already stores points as integers." ] }, { "cell_type": "code", "execution_count": null, "id": "63ab1562", "metadata": {}, "outputs": [], "source": [ "(schedule\n", " .with_columns(\n", " (pl.col('home_points') + pl.col('away_points')).alias('total_points')\n", " )\n", " .sort('total_points', descending=True)\n", " .select(['week', 'home_team', 'away_team',\n", " 'home_points', 'away_points', 'total_points'])\n", " .head(10))" ] }, { "cell_type": "markdown", "id": "2c47842a", "metadata": {}, "source": [ "### Recipe 2 β€” Team offensive EPA/play leaderboard πŸ“ˆ\n", "\n", "This is what premium EPA-tagged play-by-play unlocks. Filter to real\n", "scrimmage plays, group by the offense, and average the EPA per play β€” a\n", "clean efficiency ranking in five lines." ] }, { "cell_type": "code", "execution_count": null, "id": "eefcba31", "metadata": {}, "outputs": [], "source": [ "team_col = 'start.pos_team.name' # human-readable offense on each play\n", "epa_cols = {team_col, 'EPA', 'play'}\n", "if epa_cols.issubset(pbp.columns):\n", " leaderboard = (\n", " pbp\n", " .filter(pl.col('play') & pl.col('EPA').is_not_null())\n", " .group_by(team_col)\n", " .agg(\n", " pl.len().alias('plays'),\n", " pl.col('EPA').mean().round(3).alias('epa_per_play'),\n", " )\n", " .filter(pl.col('plays') >= 500)\n", " .sort('epa_per_play', descending=True)\n", " .rename({team_col: 'offense'})\n", " .head(15)\n", " )\n", " out = leaderboard\n", "else:\n", " out = 'expected EPA columns not present in this pbp build'\n", "out" ] }, { "cell_type": "markdown", "id": "422b3771", "metadata": {}, "source": [ "### Recipe 3 β€” A team's roster, sorted by position 🧩\n", "\n", "Join the loaded roster against `team_info` to resolve a school name to its\n", "players, then count the depth at each position group." ] }, { "cell_type": "code", "execution_count": null, "id": "c94609e5", "metadata": {}, "outputs": [], "source": [ "team_name = 'Michigan'\n", "squad = (\n", " rosters\n", " .filter(pl.col('team_location') == team_name)\n", " .select(['first_name', 'last_name', 'position', 'jersey',\n", " 'height', 'weight', 'cfbd_home_state'])\n", ")\n", "if squad.height:\n", " depth = (squad.group_by('position')\n", " .agg(pl.len().alias('players'))\n", " .sort('players', descending=True))\n", " print(f'{team_name}: {squad.height} players')\n", " out = depth.head(10)\n", "else:\n", " out = f'no roster rows for {team_name} (try another school string)'\n", "out" ] }, { "cell_type": "markdown", "id": "da377940", "metadata": {}, "source": [ "### Recipe 4 β€” Who was on the field? Per-play participants πŸ•΅οΈ\n", "\n", "[`espn_cfb_play_participants`](../cfb/reference/additional/other.md#espn_cfb_play_participants)\n", "resolves the athletes involved in each play (passer, rusher, receiver,\n", "tackler…) straight from ESPN's authoritative `participants[]` array β€” far\n", "more reliable than regex-parsing the play text. Set `resolve_missing=False`\n", "to skip the per-athlete `$ref` fan-out and keep it snappy." ] }, { "cell_type": "code", "execution_count": null, "id": "acccf372", "metadata": {}, "outputs": [], "source": [ "gid = 401628334 # 2024 CFP National Championship\n", "participants = safe(\n", " f'play participants {gid}',\n", " lambda: sdv.cfb.espn_cfb_play_participants(\n", " game_id=gid, resolve_missing=False,\n", " ),\n", ")\n", "if participants is not None and getattr(participants, 'height', 0):\n", " name_cols = [c for c in participants.columns if c.endswith('_player_name')]\n", " show = ['play_id'] + name_cols[:4] if 'play_id' in participants.columns else name_cols[:5]\n", " out = participants.select([c for c in show if c in participants.columns]).head()\n", "else:\n", " out = 'participants feed quiet right now (offseason / rate limit)'\n", "out" ] }, { "cell_type": "markdown", "id": "727fdd4c", "metadata": {}, "source": [ "### Recipe 5 β€” Build a standings table from the schedule πŸ†\n", "\n", "No standings endpoint needed: stack each team's home and away results, count wins and losses, and you've got a win-percentage table for *any* season the loader serves." ] }, { "cell_type": "code", "execution_count": null, "id": "dffd9343", "metadata": {}, "outputs": [], "source": [ "completed = schedule.filter(pl.col('completed') == True)\n", "home = completed.select(\n", " pl.col('home_team').alias('team'),\n", " (pl.col('home_points') > pl.col('away_points')).alias('win'),\n", ")\n", "away = completed.select(\n", " pl.col('away_team').alias('team'),\n", " (pl.col('away_points') > pl.col('home_points')).alias('win'),\n", ")\n", "standings_tbl = (\n", " pl.concat([home, away])\n", " .group_by('team')\n", " .agg(\n", " pl.col('win').sum().alias('wins'),\n", " (~pl.col('win')).sum().alias('losses'),\n", " )\n", " .with_columns(\n", " (pl.col('wins') / (pl.col('wins') + pl.col('losses')))\n", " .round(3).alias('win_pct')\n", " )\n", " .sort(['wins', 'win_pct'], descending=True)\n", ")\n", "standings_tbl.head(10)" ] }, { "cell_type": "markdown", "id": "26f97abb", "metadata": {}, "source": [ "### Recipe 6 β€” End-of-season power ratings ⚑\n", "\n", "[`load_cfb_ratings`](../cfb/reference/loaders.md#load_cfb_ratings) ships opponent-adjusted EPA ratings for every team-season: `adj_net` is adjusted offensive EPA per play minus adjusted defensive EPA per play, with ranks alongside. Join the schedule's team names onto its `team_id` for a ready-to-rank power table β€” no model to fit." ] }, { "cell_type": "code", "execution_count": null, "id": "d437da60", "metadata": {}, "outputs": [], "source": [ "ratings = sdv.cfb.load_cfb_ratings(seasons=[2023])\n", "names = pl.concat([\n", " schedule.select(pl.col('home_id').alias('team_id'), pl.col('home_team').alias('team')),\n", " schedule.select(pl.col('away_id').alias('team_id'), pl.col('away_team').alias('team')),\n", "]).unique(subset=['team_id'])\n", "assert ratings.schema['team_id'] == names.schema['team_id'] # join keys share one dtype\n", "power = (\n", " ratings\n", " .join(names, on='team_id', how='left')\n", " .sort('net_rank')\n", " .select(['net_rank', 'team', 'adj_net', 'adj_off_epa', 'adj_def_epa', 'games'])\n", ")\n", "power.head(15)" ] }, { "cell_type": "markdown", "id": "608024ed", "metadata": {}, "source": [ "### Recipe 7 β€” One team's full game log πŸ“œ\n", "\n", "Filter the schedule to a single program, then flip the home/away columns so every row reads from *that team's* perspective β€” opponent, points for, points against, and the margin. Swap `team` to scout anyone." ] }, { "cell_type": "code", "execution_count": null, "id": "e41da625", "metadata": {}, "outputs": [], "source": [ "team = 'Michigan'\n", "gamelog = (\n", " schedule\n", " .filter((pl.col('home_team') == team) | (pl.col('away_team') == team))\n", " .unique(subset=['game_id'])\n", " .with_columns(\n", " pl.when(pl.col('home_team') == team)\n", " .then(pl.col('away_team')).otherwise(pl.col('home_team'))\n", " .alias('opponent'),\n", " pl.when(pl.col('home_team') == team)\n", " .then(pl.col('home_points')).otherwise(pl.col('away_points'))\n", " .alias('pts_for'),\n", " pl.when(pl.col('home_team') == team)\n", " .then(pl.col('away_points')).otherwise(pl.col('home_points'))\n", " .alias('pts_against'),\n", " )\n", " .with_columns(\n", " (pl.col('pts_for') - pl.col('pts_against')).alias('margin')\n", " )\n", " .select(['week', 'opponent', 'pts_for', 'pts_against', 'margin',\n", " 'neutral_site'])\n", " .sort('week')\n", ")\n", "gamelog.head(16) if gamelog.height else f'no games found for {team}'" ] }, { "cell_type": "markdown", "id": "3e8d93cd", "metadata": {}, "source": [ "### Recipe 8 β€” Rushing leaders, EPA included πŸƒ\n", "\n", "Premium play-by-play means leaderboards aren't just totals β€” they carry **efficiency**. Filter to designed runs, sum the yards, and average the EPA per carry to separate the bell-cows from the truly explosive backs." ] }, { "cell_type": "code", "execution_count": null, "id": "78e21388", "metadata": {}, "outputs": [], "source": [ "rush_cols = {'rush', 'rusher_player_name', 'statYardage', 'EPA'}\n", "if rush_cols.issubset(pbp.columns):\n", " rushers = (\n", " pbp\n", " .filter((pl.col('rush') == True)\n", " & pl.col('rusher_player_name').is_not_null())\n", " .group_by('rusher_player_name')\n", " .agg(\n", " pl.len().alias('carries'),\n", " pl.col('statYardage').sum().alias('rush_yds'),\n", " pl.col('EPA').mean().round(3).alias('epa_per_rush'),\n", " )\n", " .filter(pl.col('carries') >= 100)\n", " .sort('rush_yds', descending=True)\n", " .head(15)\n", " )\n", " out = rushers\n", "else:\n", " out = 'rushing columns not present in this pbp build'\n", "out" ] }, { "cell_type": "markdown", "id": "9d1b019a", "metadata": {}, "source": [ "### Recipe 9 β€” The closest finishes of the year 🎒\n", "\n", "Rank completed FBS games by final margin, breaking ties toward the highest-scoring shootouts, and you've got the season's white-knuckle finishes in one expression." ] }, { "cell_type": "code", "execution_count": null, "id": "77f428ae", "metadata": {}, "outputs": [], "source": [ "thrillers = (\n", " schedule\n", " .filter((pl.col('completed') == True) & (pl.col('home_division') == 'fbs'))\n", " .with_columns(\n", " (pl.col('home_points') - pl.col('away_points')).abs().alias('margin'),\n", " (pl.col('home_points') + pl.col('away_points')).alias('total_points'),\n", " )\n", " .sort(['margin', 'total_points'], descending=[False, True])\n", " .select(['week', 'home_team', 'away_team',\n", " 'home_points', 'away_points', 'margin'])\n", " .head(10)\n", ")\n", "thrillers" ] }, { "cell_type": "markdown", "id": "8557151a", "metadata": {}, "source": [ "### Recipe 10 β€” Where does the talent come from? πŸ—ΊοΈ\n", "\n", "Roll the season roster up by `home_state` to map the recruiting footprint of college football β€” a quick reminder of just how much of the sport flows out of a handful of states." ] }, { "cell_type": "code", "execution_count": null, "id": "2bd7fb78", "metadata": {}, "outputs": [], "source": [ "talent_map = (\n", " rosters\n", " .filter(pl.col('cfbd_home_state').is_not_null())\n", " .group_by('cfbd_home_state')\n", " .agg(pl.len().alias('players'))\n", " .sort('players', descending=True)\n", " .head(15)\n", ")\n", "talent_map" ] }, { "cell_type": "markdown", "id": "f146e581", "metadata": {}, "source": [ "### Recipe 11 β€” Conference vs. non-conference, by margin πŸ”€\n", "\n", "The schedule's `conference_game` flag lets you split the slate. Restrict to FBS, then compare the average final margin in league play versus the out-of-conference cupcakes β€” group games are (predictably) tighter." ] }, { "cell_type": "code", "execution_count": null, "id": "ee1095e9", "metadata": {}, "outputs": [], "source": [ "fbs = schedule.filter(\n", " (pl.col('home_division') == 'fbs') & (pl.col('completed') == True)\n", ")\n", "splits = (\n", " fbs\n", " .with_columns(\n", " (pl.col('home_points') - pl.col('away_points')).abs().alias('margin')\n", " )\n", " .group_by('conference_game')\n", " .agg(\n", " pl.len().alias('games'),\n", " pl.col('margin').mean().round(1).alias('avg_margin'),\n", " pl.col('home_points').add(pl.col('away_points'))\n", " .mean().round(1).alias('avg_total_points'),\n", " )\n", " .sort('conference_game')\n", ")\n", "splits" ] }, { "cell_type": "markdown", "id": "4d42218c", "metadata": {}, "source": [ "### Recipe 12 β€” Biggest betting favorites in history πŸ’Έ\n", "\n", "[`load_cfb_betting_lines`](../cfb/reference/additional.md#load_cfb_betting_lines) is a premium release frame of historical sportsbook lines. Average the spread across books per game and sort to surface the most lopsided favorites β€” the mismatches Vegas saw coming a mile away." ] }, { "cell_type": "code", "execution_count": null, "id": "903a119b", "metadata": {}, "outputs": [], "source": [ "lines = safe('load_cfb_betting_lines', sdv.cfb.load_cfb_betting_lines)\n", "if lines is not None and {'season', 'market_type', 'lines',\n", " 'game_desc', 'abbr'}.issubset(lines.columns):\n", " target = sorted(lines['season'].drop_nulls().unique().to_list())[-1]\n", " favorites = (\n", " lines\n", " .filter((pl.col('season') == target)\n", " & (pl.col('market_type') == 'spread')\n", " & pl.col('lines').is_not_null())\n", " .group_by(['game_desc', 'abbr'])\n", " .agg(pl.col('lines').mean().round(1).alias('avg_spread'))\n", " .filter(pl.col('avg_spread') < 0) # negative spread = favorite\n", " .sort('avg_spread')\n", " .head(10)\n", " )\n", " print(f'biggest favorites, {int(target)} season:')\n", " out = favorites\n", "else:\n", " out = 'betting-lines frame unavailable right now'\n", "out" ] }, { "cell_type": "markdown", "id": "c98e4a21", "metadata": {}, "source": [ "### Recipe 13 β€” Hand it to pandas 🐼\n", "\n", "Every loader takes `return_as_pandas=True`, and any polars frame converts with `.to_pandas()`. Once it's a pandas DataFrame the whole pandas/`numpy`/`scikit-learn` world opens up β€” here, a one-call `.describe()` of scoring across the season." ] }, { "cell_type": "code", "execution_count": null, "id": "fbc7fe90", "metadata": {}, "outputs": [], "source": [ "score_pd = (\n", " schedule\n", " .select(['home_points', 'away_points'])\n", " .to_pandas()\n", ")\n", "score_pd['total_points'] = score_pd['home_points'] + score_pd['away_points']\n", "print(type(score_pd).__module__)\n", "score_pd.describe().round(1)" ] }, { "cell_type": "markdown", "id": "02cbb3b1", "metadata": {}, "source": [ "## πŸ—žοΈ Live tour: standings, polls, leaders & recruits\n", "\n", "A quick lap through the rest of the live ESPN surface. Each is wrapped in\n", "`safe()` so the page renders cleanly whatever the feed is doing today.\n", "\n", "| Function | Use it for |\n", "|---|---|\n", "| [`espn_cfb_standings`](../cfb/reference/site.md#espn_cfb_standings) | conference / division standings |\n", "| [`espn_cfb_rankings`](../cfb/reference/site.md#espn_cfb_rankings) | AP / Coaches / CFP polls |\n", "| [`espn_cfb_leaders`](../cfb/reference/web.md#espn_cfb_leaders) | league stat leaders by category |\n", "| [`espn_cfb_recruits`](../cfb/reference/core.md#espn_cfb_recruits) | a season's recruiting class |\n" ] }, { "cell_type": "code", "execution_count": null, "id": "ce513067", "metadata": {}, "outputs": [], "source": [ "standings = safe('ESPN standings', sdv.cfb.espn_cfb_standings)\n", "rankings = safe('ESPN rankings (polls)', sdv.cfb.espn_cfb_rankings)\n", "(standings.head()\n", " if standings is not None and getattr(standings, 'height', 0)\n", " else (rankings.head()\n", " if rankings is not None and getattr(rankings, 'height', 0)\n", " else 'standings & rankings unavailable right now'))" ] }, { "cell_type": "code", "execution_count": null, "id": "ceda0f45", "metadata": {}, "outputs": [], "source": [ "leaders = safe(\n", " 'ESPN passing leaders',\n", " lambda: sdv.cfb.espn_cfb_leaders(category='passingYards', season=2023, limit=15),\n", ")\n", "recruits = safe(\n", " 'ESPN recruiting class',\n", " lambda: sdv.cfb.espn_cfb_recruits(season=2024, limit=25),\n", ")\n", "(leaders.head()\n", " if leaders is not None and getattr(leaders, 'height', 0)\n", " else (recruits.head()\n", " if recruits is not None and getattr(recruits, 'height', 0)\n", " else 'leaders & recruits unavailable right now'))" ] }, { "cell_type": "markdown", "id": "31f666ca", "metadata": {}, "source": [ "## πŸ§ͺ Bonus: process one game from scratch with `CFBPlayProcess`\n", "\n", "Want EPA/WPA on a single *live* game without loading a whole season?\n", "[`CFBPlayProcess`](../cfb/reference/additional.md#CFBPlayProcess) drives the\n", "full ESPN pipeline: `.espn_cfb_pbp()` fetches the raw summary, then\n", "`.run_processing_pipeline()` returns a dict whose `plays` key is the\n", "fully-featured play list (alongside an advanced box score and metadata)." ] }, { "cell_type": "code", "execution_count": null, "id": "7dc488ef", "metadata": {}, "outputs": [], "source": [ "from sportsdataverse.cfb import CFBPlayProcess\n", "\n", "def process_game(game_id):\n", " game = CFBPlayProcess(gameId=game_id)\n", " game.espn_cfb_pbp()\n", " processed = game.run_processing_pipeline()\n", " return pl.DataFrame(processed['plays'], infer_schema_length=None)\n", "\n", "plays = safe('CFBPlayProcess 401628334', lambda: process_game(401628334))\n", "if plays is not None and plays.height:\n", " cols = [c for c in ('period', 'pos_team', 'down', 'distance',\n", " 'play_type', 'EPA') if c in plays.columns]\n", " out = plays.select(cols).head()\n", "else:\n", " out = 'live PBP pipeline quiet right now'\n", "out" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## πŸ›οΈ stats.ncaa.org football β€” `parse_cfb_ncaa_pbp` + box parsers\n", "\n", "parsers for the **stats.ncaa.org** football surface.\n", "`parse_cfb_ncaa_pbp` takes the raw HTML of a `/contests/{id}/play_by_play`\n", "page and emits one tidy row per play, cfbfastR-style β€” drive context,\n", "down/distance/yard line, a classified `play_type`, extracted\n", "players/yards/kick details, and boolean flags β€” while `cfb_ncaa_box` adds the\n", "box-score / drives / officials tables:\n", "\n", "```python\n", "from sportsdataverse.cfb import parse_cfb_ncaa_pbp\n", "\n", "pbp = parse_cfb_ncaa_pbp(html, contest_id=6081276) # html = captured page\n", "pbp.filter(pl.col(\"touchdown\") == True).head()\n", "```\n", "\n", "stats.ncaa.org is rate-limit unfriendly, so this family is parser-first:\n", "feed it pages captured by your own fetch pipeline (the shared proxy-bound\n", "NCAA fetch layer used by the `ncaa_mbb_*` family) rather than scraping live\n", "in a notebook." ] }, { "cell_type": "markdown", "id": "920e4773", "metadata": {}, "source": [ "## πŸŽ‰ Where to next\n", "\n", "- πŸ—„οΈ **Loaders** are your premium fast-path β€” full reference on the\n", " [Loaders](../cfb/reference/loaders.md) page (`load_cfb_pbp`,\n", " `load_cfb_rosters`, `load_cfb_schedule`, `load_cfb_team_info`).\n", "- πŸ“‘ **ESPN families** live across the\n", " [Site](../cfb/reference/site.md), [Web](../cfb/reference/web.md),\n", " [Core](../cfb/reference/core.md) and\n", " [Additional](../cfb/reference/additional.md) reference pages.\n", "- 🐼 Pass `return_as_pandas=True` for pandas, or `return_parsed=False` on the\n", " `espn_cfb_*` wrappers for the raw JSON.\n", "- πŸŸ₯ R user? The same verbs live in\n", " [cfbfastR](https://cfbfastR.sportsdataverse.org).\n", "- Part of the [SportsDataverse](https://py.sportsdataverse.org/docs/ecosystem)\n", " ecosystem.\n", "\n", "Now go chart some chunk plays β€” and may your EPA always be positive! πŸ“ˆπŸˆ" ] } ], "metadata": { "kernelspec": { "display_name": "Python 3", "language": "python", "name": "python3" }, "language_info": { "name": "python" } }, "nbformat": 4, "nbformat_minor": 5 }