# f1verse [![PyPI](https://img.shields.io/pypi/v/f1verse.svg)](https://pypi.org/project/f1verse/) [![Python](https://img.shields.io/pypi/pyversions/f1verse.svg)](https://pypi.org/project/f1verse/) [![Tests](https://github.com/jinsim/f1verse/actions/workflows/test.yml/badge.svg)](https://github.com/jinsim/f1verse/actions/workflows/test.yml) [![Dependencies](https://img.shields.io/badge/dependencies-0-brightgreen)](https://github.com/jinsim/f1verse/blob/main/pyproject.toml) [![License](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE) **[Documentation](https://jinsim.github.io/f1verse/)** · [MCP server](https://jinsim.github.io/f1verse/f1-mcp-server/) · [llms.txt](https://jinsim.github.io/f1verse/llms.txt) · [Changelog](CHANGELOG.md) · [Contributing](CONTRIBUTING.md) · [Security](SECURITY.md) · [License](LICENSE) **The story layer for Formula 1 data.** Data libraries fetch and tidy — f1verse tells you *what happened*: lead changes, laps led, event timelines, stint strategy, race pace — and *what happens next*: title probabilities from a season simulation that ships with its own backtest. **Zero dependencies.** Standard library only. Full live-timing coverage from 2023; lap-by-lap racing back to 1996, pit stops to 2011, and results, qualifying and standings to 1950 — each answer stating which era it came from and what that era does not hold. ```bash pip install f1verse ``` ```python import f1verse race = f1verse.load(2026, 12) # year, round — no other library needed race.laps_led() # {'ANT': 32, 'NOR': 31, 'HAM': 9} race.leader_runs() # [{'abbr': 'NOR', 'from': 1, 'to': 4}, ...] race.results()[7] # {'abbr': 'HUL', 'gap': '+1 LAP', ...} race.race_pace() # median pace — pit/SC/VSC laps excluded by default race.story() # one call, whole story, plain JSON race.championship_prediction() # per-lap "if it ended now" title projection race.team_radio() # timestamped clip URLs (nothing downloaded) f1verse.championship_projection(2026) # who wins the title, 20,000 seasons f1verse.title_scenarios(2026) # and who is mathematically out ``` ### Give it to an AI agent f1verse ships its own MCP server. No install step, no dependencies: ```json {"mcpServers": {"f1verse": {"command": "uvx", "args": ["--from", "f1verse", "f1verse-mcp"]}}} ``` That is the whole setup — the server is standard library only, so it starts and answers `tools/list` in about 140 ms instead of unpacking a scientific stack into a throwaway environment first. Eight tools, not eighty: a model picks the right one. For any other LLM pipeline, the library describes itself: ```python f1verse.tools() # MCP-dialect JSON schemas f1verse.tools("openai") # function-calling dialect f1verse.call_tool("f1_race_story", {"year": 2026, "round": 12}) ``` Errors are written for the caller that has to fix them without reading this page: ```python f1verse.call_tool("f1_race_summary", {}) # LookupError: unknown tool 'f1_race_summary' — available: f1_race_story, ... f1verse.load_session(2026, 12, "Qualy") # LookupError: ... listing the sessions that weekend actually had ``` ### The whole weekend, not just the race ```python f1verse.sessions(2026, 12) # Practice 1 · Sprint Qualifying · Sprint · Qualifying · Race q = f1verse.load_session(2026, 12, "Qualifying") q.results()[0] # {'abbr': 'NOR', 'q1': 72.695, 'q1_gap': 0.085, 'q3': 71.163, 'q3_gap': 0.0, # 'best': 71.163, 'eliminated_in': None, ...} q.segments()["q1"] # {'fastest': 'PIA', 'advanced': [...16 codes...], 'eliminated': [...], # 'cut_margin': 0.022} ``` Each kind gets the classification it actually has. Qualifying gaps are to the fastest lap **of that segment** — the pole-sitter above was 0.085 s off in Q1 — because a single "gap to leader" column would misreport the session. A sprint loads as a `Race`; practice is a best-lap table. ### Is this data safe to publish? ```python race.quality_report() # {'state': 'final', # provisional → settled → final, or corrected # 'coverage': {'overall': 0.9955, 'sectors': 0.9824, 'compound': 1.0}, # 'missing': ['STR.lap_46.lap_duration', ...], # 'source_age_seconds': 312, # 'revisions': [], # source rewrites this install has observed # 'crosscheck': {...}, # 'publishable': True} ``` `crosscheck` answers *do independent sources agree*. `quality_report` adds the three things that verdict is silent about: how complete the data is, how old the copy is, and whether the classification is still provisional. The chequered flag is not the final classification — scrutineering disqualifications and penalties land hours later and **rewrite rows in place**. So the rows the stewards can change are not cached forever until the session is final, and any change that is seen is recorded: ```python before = race.snapshot() # hashed, comparable, JSON — you persist it ... f1verse.diff(before, race.refresh().snapshot()) # {'changed': True, # 'changes': [{'abbr': 'HAM', 'field': 'position', 'before': 4, 'after': None}, # {'abbr': 'HAM', 'field': 'gap', 'before': '+8.1s', 'after': 'DSQ'}]} f1verse.revisions() # every source rewrite observed, with the f1verse.vintage(rec) # superseded body when it was small enough ``` There is deliberately no `as_of=` time travel: f1verse can tell you what it sees now and when it saw a value change, not reconstruct a value nobody here ever fetched. ### How the race actually unfolded ```python race = f1verse.load(2025, 24) race.running_order()[30] # ['PIA', 'VER', 'NOR', 'LEC', 'RUS', ...] who was where, lap by lap race.position_changes()[1] # {'lap': 2, 'moves': 18, 'biggest': {'abbr': 'PIA', 'gained': 17, # 'from': 19, 'to': 2}} race.battles()[0] # {'ahead': 'HAD', 'behind': 'OCO', 'from': 2, 'to': 14, 'laps': 13, # 'closest': 0.529} <- a thirteen-lap fight the results table hides ``` `battles` finds pairs that held consecutive positions within 1.5 s for at least three laps. A scrap for eighth that ran a third of the race never shows up in a classification; it is often the best part of the afternoon. ### Races from before the live feeds ```python old = f1verse.load_archive(2008, 18) # Brazil, the last-corner title old.coverage # {'lap_times': True, 'pit_stops': False, 'stints': False, # 'note': 'lap times only'} old.leader_runs() # [{'abbr': 'MAS', 'from': 1, 'to': 9}, {'abbr': 'TRU', 'from': 10, 'to': 11}, # {'abbr': 'MAS', 'from': 12, 'to': 38}, ...] old.laps_led() # {'MAS': 64, 'TRU': 2, 'ALO': 2, 'RAI': 3} ``` The `coverage` block is not decoration. 2008 has no stint data anywhere, so `ArchiveRace` has no `stints()` — rather than returning an empty dict that reads like "no pit stops happened". What the era recorded, you get; what it did not, it says. ### Seasons against each other ```python f1verse.season_shape(2025) # {'rounds': 24, 'final_margin': 2.0, # 'lead_changes': [{'round': 5, 'from': 'NOR', 'to': 'PIA'}, # {'round': 20, 'from': 'PIA', 'to': 'NOR'}], ...} f1verse.title_margins(top=5) # 2025 NOR 423.0 vs VER 421.0 margin 2.0 (0.08 of a win) # 2008 HAM 98.0 vs MAS 97.0 margin 1.0 (0.10 of a win) # 2012 VET 281.0 vs ALO 278.0 margin 3.0 (0.12 of a win) ``` Points systems changed repeatedly, so a raw margin cannot compare eras — one point in 1958 was most of a win. Every row also carries the gap measured in wins, which is the comparison that survives the rule changes. ### Where the passing happened ```python f1verse.overtake_hotspots(race)[0] # {'from_s': 3510.0, 'to_s': 3540.0, 'signals': 69, 'drivers': [...]} ``` The timing feed publishes an `OvertakeState` per car that almost never changes — about a hundred transitions in nineteen thousand records. That sparsity is the value: the transitions are a free index of the moments worth looking at, from the same feed that times the race. Read it as "something happened here", then confirm against `running_order`. ### Beyond a single race ```python f1verse.career("max_verstappen") # {'starts': 245, 'wins': 71, 'podiums': 131, 'poles': 48, ...} 1950-present f1verse.milestones("max_verstappen") # [{'stat': 'poles', 'current': 48, 'target': 50, 'remaining': 2}] f1verse.circuit_profile(2026, 13) # corners, marshal sectors, track outline, and pit loss split by track state # {'normal': 25.43, 'sc': 16.11, 'vsc': 18.4} <- what an undercut costs here # plus historic record: 75 races held, pole-to-win rate 0.30 # Geometry is also interpreted, but never overclaimed: every corner carries # its lap position, preceding-run share and local heading deflection; mini- # and marshal-sector boundaries carry lap percentages. profile = f1verse.circuit_profile(2026, 13) profile["layout"]["corners"][0] # {'number': 1, 'progress_pct': 8.412, 'local_deflection_deg': 72.6, ...} # What a map cannot say, the cars can. This measures the circuit from the # session's own telemetry: the height profile, where overtakes actually # happen, how much of the lap is flat out and where it is braked, and every # numbered corner as the car experienced it. f1verse.circuit_survey(2026, 13)["corners"]["corners"][0] # {'corner': 1, 'apex_speed_kph': 103, 'radius_m': 40.0, # 'lateral_load_g': 2.09, 'gear_at_apex': 2, 'severity': 'medium'} # Published specifications are stored rather than derived, and the # measurement audits them rather than replacing them. f1verse.circuit_profile(2026, 13, measure=True)["audit"] # {'verdict': 'agrees', 'checked_age_days': 0, # 'checks': [{'field': 'length_m', 'published': 4259, 'measured': 4274.4, # 'off_by_percent': 0.36, 'state': 'agrees'}, ...]} # The published facts and the audit are reachable on their own, so a caller # can ask what is claimed without paying for a telemetry survey. f1verse.circuit_facts("Austin") # {'length_m': 5513, 'race_laps': 56, 'race_distance_km': 308.405, # 'provenance': {'length_m': 'curated', ...}, # 'source': 'formula1.com official circuit page, 2026 united-states', # 'checked': '2026-09-01'} # None for a circuit nobody has curated yet - it does not invent a length f1verse.circuit_audit(2026, 13) # the same verdict as profile(measure=True)["audit"], on its own f1verse.circuit_directory() # every venue recorded in F1 results, with stable id, city, country and # coordinates; unlike a current map, it does not pretend an old layout has # today's geometry f1verse.circuit_history("monza") # every F1 event at the venue, recent winners and starting positions, # pole-to-win conversion, and the drivers with the most wins f1verse.head_to_head(2026) # teammate quali/race scores per constructor f1verse.standings(2026) ``` ### Tyre life, and what the stewards struck out ```python f1verse.stint_degradation(race)[3] # {'driver': 'NOR', 'stint': 2, 'compound': 'HARD', 'tyre_age_at_start': 0, # 'clean_laps_used': 18, 'degradation_s_per_lap': 0.041} f1verse.circuit_abrasion(race) # {'factor': 1.4, 'verdict': 'abrasive', 'samples': 55} f1verse.tyre_outlook(race) # projected to the cliff ``` A car gets faster all race as it burns fuel, so raw lap times understate degradation on every stint. Rates are fitted to **fuel-normalised clean laps only**, and a stint with too few of them returns `{'degradation_s_per_lap': None, 'reason': 'too few clean laps'}` rather than a number fitted to noise. ```python f1verse.lap_deletions(session.race_control) # [{'car_number': 55, 'lap_time': '1:25.773', 'stands': True, # 'reason': 'TRACK LIMITS AT TURN 3 LAP 3', ...}] ``` A reinstated lap keeps the reversal visible instead of being quietly dropped. Check this before treating a fastest lap or a qualifying position as settled. ### Live timing, straight off the wire ```python from f1verse.sources import liveclient with liveclient.LiveFeed() as feed: for topic, patch, stamp in feed.messages(): ... liveclient.run("session-{n}.jsonl") # record, rotate on turnover, reconnect ``` The official SignalR feed over a WebSocket written in the standard library — no websocket package, no SignalR client. Connecting takes three undocumented courtesies (an affinity cookie only handed to a pre-flight request, the official application's user agent, an invocation id on the subscription); they live in one place so no caller rediscovers them. Recorded frames carry a millisecond arrival stamp — the difference between an archive and a screenshot — so `replay()` runs a session back at true speed. The stream is not a lap table, so `f1verse.sources.timing.laps_from_stream` rebuilds one: arrival order lies, sector times land after the next lap has begun, and qualifying carries phantom lap times that are really the gap between two runs. Every rebuilt lap row carries its provenance. ### Running this on a schedule ```python f1verse.status(2026) # {'latest_race': {'round': 12, 'meeting': 'Dutch Grand Prix'}, # 'next': {'round': 13, 'session': 'Practice 1'}, 'next_in_hours': 125.7} f1verse.due(2026, processed=[...session keys you already handled...]) # sessions finished, settled (45 min past the flag) and not yet processed — # nothing published twice, nothing missed after downtime ``` Caching is policy-driven, not blanket: lap and telemetry data is immutable and cached forever, schedules expire every few hours — a calendar cached for a season would hide a cancelled round for the rest of the year — and rows the stewards can still rewrite expire until the session is final. `f1verse.cache_info()` and `f1verse.clear_cache(older_than=...)` are there for operators; `clear_cache` never drops the revision journal. ### Telemetry, track position, conditions ```python f1verse.lap_telemetry(race, "NOR", 40) # per-sample speed, throttle, brake, gear, RPM and DRS state for one lap f1verse.lap_trace(race, "NOR", 40) # x/y/z coordinates of that lap f1verse.top_speeds(race) # fastest reading per driver f1verse.weather_summary(race) # {'track_c': {'min': 25.8, 'max': 38.2}, 'rain': True, 'samples': 191} ``` Telemetry is high-frequency, so these take a bounded window and filter server-side rather than downloading a session and trimming it locally. ### Grounded narration ```python facts = f1verse.race_facts(race) # all numbers computed and formatted here f1verse.brief(race) # deterministic text, no model required result = f1verse.narrate( race, generate=lambda prompt: my_model(prompt), cache_dir=".cache/narration", ) # {'text': '...', 'source': 'generated' | 'cache' | 'template', ...} ``` `narrate` accepts any text-generation callback; f1verse has no model SDK dependency. Drafts are checked against the structured fact sheet. Unknown numbers and driver codes are rejected, generation is retried at most twice, and a deterministic summary is returned if verification still fails. The optional cache is exact-match only and stores verified text. ### Who wins the championship — and who still can Two different questions, and mixing them is how a projection ends up implying somebody is out when the arithmetic says otherwise. ```python f1verse.title_scenarios(2026) # arithmetic, not a forecast — maximum points left is a fixed number # {'rounds_left': 11, 'sprints_left': 1, 'max_points_available': 283, # 'drivers': [{'driver': 'RUS', 'gap_to_leader': 59.0, # 'still_possible': True, 'needs_avg_per_round': 5.4}, ...], # 'still_alive': 23} f1verse.championship_projection(2026) # the rest of the season, played out 20,000 times # {'drivers': [{'driver': 'ANT', 'title_probability': 0.954, # 'points_now': 242.0, 'projected_points_median': 446, # 'projected_points_p10': 390, 'projected_points_p90': 490, # 'races_in_sample': 11, 'measured_dnf_rate': 0.08}, ...], # 'assumptions': {'ignores': ['car development', 'weather', ...], ...}} ``` A simulated finish is **resampled from the positions that driver has actually finished in**, not drawn from a curve around their average. The distinction decides championships: alternating wins and retirements is a different proposition from finishing fourth every weekend, and the mean is the same for both. Retirements fire at each driver's measured rate, sprints score on their own table, and ties break on wins then seconds, as the regulations do. Every run also re-draws the driver's own level first, bootstrapping their results before playing the season against that version of them — twelve races is a small sample, and treating it as settled truth is how a forecast becomes more confident than anyone should be. ```python f1verse.backtest_projection() # does the model deserve to be believed? # {'by_confidence': {'over_90': {'n': 5, 'correct': 5, 'rate': 1.0}, # 'under_60': {'n': 2, 'correct': 0, 'rate': 0.0}}, ...} ``` Replayed at round 12 of 2019-2025 it went 5/5 when it claimed 90%+; both misses were seasons it had itself called at 53% and 57% — the two that ran to the final round. Read the buckets, not the headline: a model that says 55% and is wrong has done nothing wrong. > `championship_prediction` (in `feeds`) is a different thing — F1's own > live "if it ended now" table. `championship_projection` is this model. > One reports, the other forecasts. ### Predictions, pit-stop verdicts, official documents ```python f1verse.win_probabilities({"NOR": 1, "ANT": 2, "RUS": 3}, year=2026, upto_round=12, circuit_id="monza") # every probability ships with its own evidence: # grid base rate measured over 233 real races (pole wins 54.1%) # blended with that circuit's pole-to-win conversion (Monza: 0.30) # scaled by recent form (average finish over the last 5 rounds) f1verse.pit_exchanges(race, pit_loss_s=22.74) # [{'lap': 17, 'driver': 'RUS', 'rival': 'PIA', 'verdict': 'worked', # 'gain_s': 3.54}, ...] # neutralised laps (red flag / SC / VSC) and same-lap covering stops are # excluded — calling those undercuts would be wrong f1verse.fia_documents(2026) # stewards' decisions, classified f1verse.power_unit_documents(2026) # "who changed which engine part" ``` ## Why this exists Raw timing data needs a lot of domain knowledge before it means anything: - **Classified gaps are not comparable across lapped cars.** A car one lap down can show a smaller number than one that finished ahead on the lead lap. → `format_gap` applies the broadcast convention (`+1 LAP`). - **"Who led the race" has to be derived.** Lead changes, laps led and the moments they happened are not published as such. → `leader_runs`, `laps_led`, `timeline`. - **Race pace needs rules**, not just a threshold: in/out laps and laps run under SC/VSC have to go, or the number is meaningless. → `race_pace` applies them by default. - **Web and video pipelines need plain JSON.** Every f1verse output is JSON-safe Python, ready to serialise. - **Data quality should be checkable in code**, not read from logs. → `crosscheck` and `quality_report` return structured verdicts. - **Results change after the flag.** A cache that treats a classification as immutable makes a stewards' decision invisible. → revisable rows expire until the session is final; changes are recorded. ## Additional live-timing feeds The official live-timing archive publishes several feeds that are rarely surfaced. f1verse parses three of them, with the same caching and rate-limit etiquette as the rest of the library: ```python f1verse.championship_prediction(session) # per-lap "if the race ended now" projection of both championships, # including the moments the projected champion changed f1verse.team_radio(session) # timestamped team-radio clips: [{'t', 'utc', 'driver_number', 'url'}] # URLs only — nothing is downloaded or redistributed f1verse.timing_stats(session) # personal bests, best sectors, speed-trap figures ``` ## Running continuously See **[OPERATIONS.md](OPERATIONS.md)** for caching policy, rate limits and scheduling. ## Tests ```bash pip install -e ".[test]" && pytest -q ``` ## Sources | Layer | What it gives | |---|---| | Race data | laps, stints, pits, positions, results, overtakes | | Live-timing archive | championship projection, team radio, timing stats | | Historic records | careers, circuit records, standings — 1950 onward | | Circuit geometry | track outline, corners, marshal sectors, pit loss | All are public endpoints, read at runtime. See `src/f1verse/sources/` for the exact hosts, and `NOTICE` for how upstream data rights relate to this project's own license. ## Design rules 1. **Zero required dependencies.** The native loader reads public REST endpoints and the official live-timing archive directly, with its own on-disk cache and polite pacing. 2. **Everything returned is plain JSON-safe Python.** 3. **F1 domain rules are defaults, not options.** 4. **Cross-checked where possible** — lapped-car gaps, for instance, are computed by convention *and* confirmed against a second source. 5. **Code only.** No timing data, media, or images are bundled or redistributed; data is fetched by the end user. ## License Apache-2.0. See [LICENSE](LICENSE) and [NOTICE](NOTICE). Use it privately or commercially; read it, change it, redistribute it, and build products or hosted services on it. Preserve the license and notices, state significant changes, and observe the license's patent terms. There is no source-disclosure requirement for software or services that use f1verse. This license covers f1verse's own source. The Formula 1 data it reads at runtime belongs to its respective rights holders under their own terms. ## Roadmap - Broader cross-validation coverage - Deviation detection: expected range, actual, evidence - Localisation packages --- *Unofficial fan project. Not affiliated with, endorsed by, or associated with Formula 1, FIA, FOM, or any F1 team. F1, FORMULA 1 and related marks are trademarks of Formula One Licensing BV. This library contains code only — no timing data, media, or images are included or redistributed; data is fetched by the end user from publicly accessible endpoints, subject to the respective providers' terms.*