oneproof.dev

Evidence infrastructure for AI systems

Show, don't tell, what your AI system did.

oneproof is an evidence layer for AI systems — a way to independently reconstruct, compare and verify what a system actually did, by someone who wasn't there when it did it.

Your AI system — what already exists The evidence layer — what oneproof adds A question someone needs answered "what notice is required to terminate early?" Agent and application logic planning · tool selection · orchestration Retrieval read · convert · clean · chunk · embed · index · search The model prompt assembled · generation · citation An answer someone acts on "ninety days" — or "not stated in retrieved evidence" question in — out 0211c293 cfg a19c plan · tools in 0211c293 out a6ad8533 cfg 7f31 retrieval — one record per stage in a6ad8533 out f46e4fd5 cfg 3e08 prompt · model in 16e73d45 out 9cc0cc2a cfg b204 answer digest 9cc0cc2a supported false digests, not copies — the record holds no document text what you have today: logs and traces — ran, returned OK, 340 ms what they cannot tell you: what it did it to Reconstruct rebuild a stage from its record Compare two runs, stage by stage Verify by someone who was not there

Where it sits. One record per stage — what went in, what came out, what ran. Digests, not copies. Nothing in the system above has to change its behaviour to be recorded.

Why a layer, and why nowObservability tells you it worked. It cannot tell you what it worked on.

Every stage logs success, because every stage succeeded. Nothing in that sequence has anything to say about the fact that one step quietly moved the sentence you needed out of reach.

A log is testimony: it tells you what a system says it did.
A record is evidence: it lets someone else check.

9
comparisons to locate a fault — the same nine at thirteen documents and at forty‑seven thousand
99.96%
of documents shown to be affected, computed without opening one of them
3
machines, two operating systems, identical digests — so a stage can be rebuilt instead of stored
0
network calls, keys or accounts a checker needs. Offline, or it isn't evidence

Every figure on this page is measured and traceable to a published file. Where each one comes from →

What it looks like when you have itTwo runs of the same question, laid side by side.

A routine upgrade changed what the system answered. No error, no alert, every log green. This is the record doing the work a log cannot.

one question · two recorded runs · one configuration line changed stage baseline candidate verdict document in0211c293efe00211c293efe0● same root converted0211c293efe00211c293efe0● same root cleaneda6ad8533e58aa6ad8533e58a● same root splitf46e4fd5c8475973d5fe20b7● FIRST DIFFERENCE embedded3317f5ba98bd8879c65c2fe3● downstream indexed5e023c0999bf72704d9fe7d5● downstream retrievedcff666115d20006e8fded426● downstream prompt built16e73d452a1964a63d68cc3d● downstream answer9cc0cc2a58619cc0cc2a5861● same digest nine comparisons — and the same nine at forty-seven thousand documents. you compare stages, not documents. the change that caused it oneproof-chunker 3.4 → 3.5 divide_paragraphs False → True max_chars 900 → 220 strategy paragraph-packing → sentence-packing what you can then do 01 · locate Which stage broke it the first row that disagrees. everything above it is identical; everything below is consequence, not cause. 02 · measure the reach 47,412 of 47,430 documents whose passages moved — computed from index digests, with no chunk text read. eighteen provably untouched. 03 · replay Rebuild that one stage and watch it produce the same wrong thing again — or, where a stage cannot honestly be rebuilt, have the record say so. 04 · read the caveat the answer digest matched on both sides because both runs refused. that is not robustness, and the report says so.

Real output. 47,430 SEC filings, 19.3 GB, one configuration line changed. The full report, and the passage that was present and never read →

The four doorsFour questions, at four moments in the life of a system.

One discipline across all of them: measured instead of asserted, and every answer checkable by someone who does not trust us. Use one alone, or all four together.

time → Before you build Choose oneground Is this the right retrieval setup for our documents? measured on your own data, against the right answers preview · free · on your machine At every action Prevent onedoor Was this action allowed, and who says so? per-action authorization, decided and receipted shipped · v0.7.0 on PyPI Continuously Detect onewatch Is it still the system we audited last month? verifiable baselines, so a change leaves a record that it happened public · open since 10 September After the fact Prove onetrace What did this answer actually rest on? a record at every stage, so two runs can be compared by a stranger format & verifier public · free
Before you build · Choose

Choose — oneground

Is this the right retrieval setup for our documents?

architecture choice8 option(s): 0 meets, 6 fails, 2 couldnt_check nothing is recommended

Measured on your own documents against the right answers, before you build on it. Its most common first output is a refusal to recommend.

Preview released 16 September · free · your data stays on your machine
After the fact · Prove

Prove — onetrace

What did this answer actually rest on?

stage receiptssplit f46e4fd5 → 5973d5fe FIRST DIFFERENCE · 9 comparisons reach 47,412 of 47,430 · 0 read

A record at every stage, so an answer traces to the exact passage that produced it — and when two runs disagree, so does the place they stopped agreeing.

Format, samples and verifier public · free · the SDK is not shipped
At every action · Prevent

Prevent — onedoor

Was this action allowed, and who says so?

policy ledger · payments pack transfer cap € 500 █████ refunds cap € 300 ███ version 29e85d2c…5166 ratified

Define caps, bounds and undo windows; ratify a version; every action is then checked against that ratified snapshot — allow, review or refuse — and every verdict names the exact policy version that decided it.

Shipped · v0.7.0 on PyPI · Apache-2.0 · source public
Continuously · Detect

Detect — onewatch

Is it still the system we audited last month?

change evidencebaseline byte-identical watch 74/74 pins hold a system that changes will say so

Verifiable baselines, so when something underneath you changes you hold a record that it did — evidence rather than a feeling.

Public since 10 September · Apache-2.0 · verification tooling open

Not mockupsFive real outputs. Click through them.

Every panel below is a tool's own output from a run that happened. Where a tool cannot answer it says so in its own words rather than rounding up — and each one names where you can go and read it.

8 option(s): 0 meets, 6 fails, 2 couldnt_check

No option meets every constraint, so nothing is
recommended. Recommending an option whose constraints
could not all be checked would be rounding
couldn't-check up to a verdict.

To decide latency_p95: run `oneground verify` against
a real engine in the environment the constraint
targets. Latency is never taken from simulation.

A tool that always answers is a tool you cannot cite.

Eight candidate setups, measured on a real corpus. None cleared every constraint, so none is recommended — and the two it could not check are named rather than folded in.

oneground · the published report it came from →

Go deeperThree ways in, depending on why you're here.

If you are buying

What it costs and what it refuses

Cost, risk, what leaves your building, and the limits stated before you find them.

If you will run it

Commands, formats, preconditions

What each tool needs, what it writes, and how to check every figure against its source.

If you are reviewing the work

Drafts, open requirements, failures

The specifications, what they do not yet cover, and the register that tracks it in public.