# Processing correctness Run from the repository with the normal Python runtime, reference profiles, Rust/Cargo, Node, and Playwright installed: ```sh .venv/bin/python scripts/check-processing.py --film-gpu ``` This builds the current Rust source in `build/processing-cargo`, generates synthetic fixtures, executes independent references, and checks the production engines, exports, browser application, and actual Metal display frames. It uses an isolated photo library and runtime; it does not edit a personal catalog or replace the installed application. The command exits nonzero for a numerical mismatch, missing runtime, unavailable GPU/display, timeout, or incomplete suite. It never converts a missing native display into a passing skip. All requested suites must finish before the report can say PASS. On macOS, run outside restrictive process sandboxes in an unlocked graphical login with Metal available. ## Evidence `build/processing-correctness/` contains: - `report.md`: suite outcomes and failures. - `report.json`: source hashes, revision, dirty-tree state, individual metrics, tolerances, runtime identity, and artifact paths. - `junit.xml`: CI results; blocked suites are failures. - Per-case reference, actual, and amplified difference PNGs for RGB comparisons. - Floating-point stage arrays, requests, logs, grain statistics, and submitted native drawable records. Film arrays retain their physical units. - Full browser-app screenshots and separate pixel-scored canvas captures. Never bless a failing output by updating a screenshot golden or increasing a tolerance until it passes. Locate the first divergent stage. If a reference contract changes intentionally, explain the model change and derive a new tolerance before comparing results. ## What the pixel gates cannot establish Every suite above compares two implementations of the same model. Three kinds of defect survive that by construction, and three further gates exist for them. All of them run in the ordinary unit suite. | Gate | Question it answers | Command | | --- | --- | --- | | `tests/test_control_coverage.py` | Is every control exercised by anything at all, and does any slider travel past the clamp the server applies? | `python -m unittest tests.test_control_coverage` | | `tests/test_control_semantics.py`, `tests/test_film_semantics.py` | Does each control do what its name says, across its whole travel? | `python -m unittest tests.test_control_semantics tests.test_film_semantics` | | `tests/test_processing_goldens.py` | Did the rendered result of a complete recipe change? | `python -m unittest tests.test_processing_goldens` | | `tests/test_raw_capture.py` | Do the capture-stage controls work, on a real mosaic? | `python -m unittest tests.test_raw_capture` | **Coverage.** The case lists are written by hand, so a control added to the schema or to the markup is checked by nothing until someone remembers it. `tests/control_inventory.py` derives the full set of controls by probing the production cleaning functions and parsing `web/index.html`, then the coverage test fails on any control no case list or audit reaches. A control a pixel gate genuinely cannot own needs an entry in `DELEGATED` naming the suite that does own it; an unexplained entry fails. **Semantics.** A parity gate passes happily when both implementations are wrong in the same way, and says nothing about a control that was never wired up. The audits sweep each control across its full travel and check that a named, physically meaningful measurement moves the way the label promises: Exposure raises mean luminance, Texture raises one-pixel local contrast, noise reduction lowers flat-field variance, a mask type selects part of the frame rather than all of it, `subtract` shrinks a selection. Every control earns a verdict, and `inert`, `leaky` and `contradicted` fail the build. `build/control-audit/audit.md` lists them all after a run. Known defects live in `TRACKED_DEFECTS` with a written explanation of what the control does today and why that is wrong. They do not fail the build, because they are already recorded. A tracked defect that starts behaving correctly *does* fail, so the table cannot rot into a list of things nobody believes any more. **Frozen renders.** If the grade and the shader are edited together, or a profile's measured data is replaced, both sides agree and every parity gate stays green while the picture people already saved quietly changes. Fourteen complete recipes are stored in `tests/goldens/` at 160x112. Their film interpretations are explicit, so changing the new-photo default does not change a frozen recipe. Stored reference bytes are checked against SHA-256 digests. Live CPU renders allow at most one 8-bit code value in 0.01% of channels for cross-platform rounding at half-code boundaries; larger changes or systematic offsets fail. Re-blessing is deliberate: ```sh python tests/processing_goldens.py --bless ``` which refuses while `RENDER_CACHE_VERSION` is unchanged, because a changed render that keeps its cache key serves the old pixels from every warm cache in the field. Failures write amplified difference images to `build/golden-diffs/`. **Capture stage.** White balance, highlight recovery, demosaic profile and sensor denoise run in the RAW decoder, before the film model and before any grade, so no rendered TIFF fixture reaches them. `tests/test_raw_capture.py` decodes synthetic DNGs written by `tests/make_raw_fixtures.py`: a few kilobytes each, valid uncompressed mosaics carrying a known colour matrix, a known as-shot neutral, clipped highlights and a noisy flat field. They test the decoder, not a sensor; real camera formats stay with the RAW journey layers in docs/journey-testing.md. The audits and the goldens drive the resident `lighttable-engine` built from `rust-engine/`, the same implementation the film gate tests, never the prebuilt `engine/spektrafilm-rs`. Build it once: ```sh CARGO_TARGET_DIR=build/processing-cargo cargo build --manifest-path rust-engine/Cargo.toml --locked ``` Both suites skip themselves without it, so CI builds it before the unit step and then asserts it is present, rather than staying green while checking nothing. ## What each suite establishes | Suite | Implementation under test | Independent reference and coverage | | --- | --- | --- | | `film` | Freshly built Rust resident engine, CPU diagnostic taps | Python spectral runtime at film exposure, developed film density, print exposure, developed print density, and scanned RGB; separate NumPy exposure and measured characteristic-curve checks. Isolates exposure, development gamma, dye couplers, halation, diffusion, print exposure/preflash/filters, scanner blur/sharpen/glare, reversal, and B&W development times. | | `film --film-gpu` | Actual WGPU film output | Verified CPU output for deterministic recipes; backend identity must say WGPU. A CPU fallback fails. Intermediate GPU buffers are not exposed, so GPU stage-by-stage equivalence is not claimed. | | `export` | Actual `render_cli.py` subprocesses and current Rust resident export | Independent NumPy sRGB, linear-light exposure, rotation/crop and Lanczos equations; RGB16 precision and ICC profiles. Python grading/masks independently check Rust postprocessing from the same float film base. | | `webgl` | Production `web/gl.js` in Chromium | Python CPU grading for individual positive/negative controls, detail/noise/sharpening, all curve channels and HSL bands, Point Color, tonal color grading and combined operation order. Scores synchronous framebuffer output and composited canvas screenshots, with navigation, portrait geometry, and compare restoration. | | `app` | Complete server and browser application | Real UI commands and slider/save/render paths, Film on/off, print exposure, and portrait/landscape navigation. Requires the intended saved recipe and matching ready photo. Scores the actual frame against the CLI render endpoint with an explicit reference recipe, and the visible scaled canvas against its rendered pixels. The browser uploads the engine's raw RGBA8 preview surface, and `/api/render/file` reads the same lossless surface for its film base (never the quality-88 preview JPEG, whose chroma subsampling blends colour edges by up to ~25 code values), so both sides of the comparison are exact engine output. | | `native` | Production `NativePreview.swift` and `NativePreview.metal` in a real AppKit window | Python grading against the actual MTKView drawable submitted for display, including image edges. Exercises PNG/FLRA input, all shared grade cases, cached A/B/A navigation, compare sweeps/restoration, and portrait/landscape sizing. Requires GPU completion and drawable presentation. | The native helper compiles the production renderer and shader; it is not a replacement for the full native product journey. Run the existing `scripts/run-product-journey.sh pr` as well when changing native bridge wiring, window layout, or packaging. The complete browser application is exercised by the `app` suite; the Metal helper isolates native rendering correctness. ## Tone working space The tone stage of the grade (Exposure, Highlights, Shadows, Whites, Blacks, Contrast) is defined once in `grade._tone_stage` and reproduced by the Numba kernel, `web/gl.js` (`toneStage`), `app/NativePreview.metal` (`toneStage`), `rust-engine/src/grade_gpu.wgsl` (`tone_stage`) and `rust-engine/src/export.rs` (`tone_stage`). Every implementation takes the display-encoded base frame (the cached film render, or the developed RAW/positive) and: 1. decodes it to scene-linear light (sRGB curve; Display P3 shares it, and ProPhoto RGB decodes the ROMM curve on the wide-gamut export path); 2. applies Exposure as a linear gain of `2^EV`, then the luminance-masked Highlights (`1 + 0.85·h·m`, `m = clamp((Y − 0.35) / 0.65)^1.2`) and Shadows (`1 + 1.5·s·m`, `m = clamp((0.45 − Y) / 0.45)^1.2`) gains, with no clip; 3. re-encodes with the sRGB curve continued above 1.0 by the same power law, so the extended signal is still a monotone function of linear light; 4. applies Whites and Blacks as endpoint moves on that extended signal, `(t − b) / (w − b)` with `w = 1 − 0.35·whites`, `b = −0.25·blacks`, and Contrast as the smoothstep blend (`k > 0`) or the flatten toward mid grey (`k < 0`); above 1.0 the smoothstep continues as `min(1 + (t − 1)^2, t)`, which keeps its zero slope at white and rejoins the identity; 5. rolls everything above the knee off toward white with an exponential shoulder, `knee + (1 − knee)·(1 − exp(−(t − knee) / (1 − knee)))`, and bounds the result to [0, 1]. The knee is a function of the sliders alone (`grade.tone_knee`): the encoded value an untouched white reaches after Exposure, Highlights, Whites and Blacks, mapped through `1 − 0.15·(1 − exp(−2·excess))`. A recipe that cannot push anything past white has `knee = 1` and the stage is the exact hard clip it was before, so every value inside [0, 1] renders exactly as it did; the shoulder only opens with the headroom a recipe creates, and opens continuously (a 0.01 nudge of Exposure moves nothing by more than 0.02). Dehaze, Temp/Tint, Saturation, Vibrance, HSL, Point Color, Color Grading, the curves and the vignette stay display-referred after the stage, as before. Local masks run the same stage through `grade.apply` (their Whites and Blacks are always zero). Film renders feed the print-scan output through it exactly like a Develop frame; the film goldens did not move because their recipes carry no headroom. Measured against the previous clipped stage on the three real-photo fixtures and the golden source (luma in 8-bit code values): Exposure −1 and Highlights −100 are bit-identical; Exposure +1 moves mid-tones (0.35–0.65) by at most 0.3 and the highlight mean (> 0.75) by 0.5–2.6 (a pushed white lands on 248 instead of 255, and the share of clipped pixels drops from 7.5–66% to 0–51%); Whites +50 moves mid-tones by at most 0.2 and highlights by 0.1–2.1. No saved edit needs a migration factor, so the edit schema stays at version 2. The `tone.linear.*` and `tone.headroom` findings in `tests/control_semantics.py` hold each control to its direction in linear light and prove that a ramp sent past white by +1 EV still rises after Whites −100. Because the stage depends on the transfer function alone, a Display P3 or ProPhoto export whose grade uses nothing but these six controls (`grade.is_tone_only`) keeps the wider gamut: `wide_develop_edits_supported` accepts it and `render_cli.py` grades the delivered encoding directly. Any other grade control, mask, retouch or watermark still renders through the sRGB preview path with the sRGB-limited warning. ## Reference contracts and tolerances Film stage errors use log10 exposure or optical density. Scanned RGB and preview tolerances are measured as 8-bit code values, even when input/output arrays are float. The exact limits are stored alongside every result; RGB preview defaults are mean <= 0.75, p95 <= 2, maximum <= 5. Maximum error matters: a few badly rendered border pixels must not disappear into an image-wide mean. RGB8 encoding has a half-code quantization bound. Lossless RGB16 exports use limits near one 16-bit code. GPU and curve texture rounding need separate documented limits from float CPU stages. Comparisons reject NaN, infinity, empty output, wrong dimensions, wrong channels, and missing artifacts. Grain is stochastic. Different samplers cannot be expected to generate the same individual particles. The suite instead requires exact repeatability within an engine and tests flat-field means and variance against analytical Poisson/binomial expectations. With GPU enabled it also checks CPU/GPU scan noise moments. Statistical regions exclude the known blur support between flat fields; ordinary preview image comparisons retain image edges. Browser element screenshots include fractional CSS boundary coverage. The app check independently scores the framebuffer against the CLI, then compares the visible app canvas with a separate reference canvas containing those readback bytes at identical DOM bounds. Both use the browser's compositor, including its fractional scaling, rather than approximating it with an integer-sized resize. The reference page is labeled as a reference artifact, never an app screenshot. No image registration or border trimming is used. These are software-equivalence tests for the declared model and fixtures, not proof that a simulation matches a particular physical negative, chemical bath, scanner, printer, or monitor. The two spectral implementations share measured profile data, so agreement cannot establish that the profile data are correct. Real-camera RAW demosaicing/WB, all stock combinations, operating-system display calibration, and stochastic perceptual quality need additional fixtures or real-device comparisons. Regression coverage includes Vision3 200T/500T with automatic metering on/off, flat-field Texture/Clarity after global exposure or an earlier mask, real-photo full and partial masks, Heal/Remove, sequential Clone, sharpening after retouch, and manual optical boundaries. Native edit cases exercise the server's bake decision before submitting its surface and edit payload to Metal. Browser app cases additionally cover an exposure slider while a detail mask is active and clearing the last detail mask. Ordinary grade-only controls retain their live GPU path. Retouching uses the exact ordered CPU base correction before GPU grading. Masks with Texture or Clarity bake the complete global-grade/local-mask stack, because a single fragment pass cannot sample the image produced by earlier adjustments. These edits can increase preview latency; the expensive film base is reused. Corrected surfaces are RGBA8 for Metal and lossless PNG for browser presentation. Soft proof, compare, and reference overlays remain display steps. Active optical diffusion now uses the exact CPU PSF convolution on both render backends. Other supported stages remain on the selected GPU backend, but the diffusion recipe renders the full frame before viewport trimming. This removes the previous preview/export halo mismatch and can increase latency when the effect is enabled. The default diffusion strength is zero. ## CI and focused runs Pull requests and main run the film CPU, export, WebGL, and complete browser app gates in `.github/workflows/python-tests.yml`, and upload failure artifacts. Hosted macOS runners do not stand in for a verified physical Metal/display session. The full local command additionally requires native and film GPU proof. ```sh # Hosted/CPU-compatible subset (explicitly omits native and film GPU proof). python scripts/check-processing.py --suites film,export,webgl,app # Focus on a native shader change. python scripts/check-processing.py --suites native --output build/native-pixels # Scorer and reference self-tests; real engine comparisons are separate gates. python -m unittest discover -s tests -p 'test_processing_*.py' -v ``` Application GPU finishing and merge stages have additional device tests. Build the current resident engine into `engine/lighttable-engine`, then run these in an unlocked graphical session with GPU access: ```sh LIGHTTABLE_TEST_GPU=1 python -m unittest discover -s tests -p 'test_gpu_grade.py' -v LIGHTTABLE_TEST_GPU=1 python -m unittest discover -s tests -p 'test_merge_acceleration.py' -v ``` These compare float32 output against the existing CPU algorithms, including odd dimensions, image boundaries, tiled filter halos, local-mask composition, and direct resident export ordering. Eligible device cases fail if they silently fall back. Ordinary runs still exercise the explicit unavailable-device fallback. Export finishing retains CPU chromatic-aberration geometry and precision-sensitive Point Color luminance-uniformity recipes; basic-only grades use the existing Numba path when available because worker transport can cost more than it saves. `bench/gpu_finish_benchmark.py` measures full-resolution finishing including worker transport, and can compare encoded 16-bit TIFFs with `--export-check`. `bench/merge_benchmark.py` measures complete merges and their phases with fixed registration randomness. Its `--baseline-source` option accepts the original `merge_workflow.py` for a baseline/new-CPU/GPU comparison. Use fresh processes to measure cold startup separately, and report source dimensions, recipe, repetition count, output error and device identity alongside speed. Install a reproducible browser test runtime if one is not already configured: ```sh npm install --prefix build/processing-browser --no-save playwright@1.62.1 build/processing-browser/node_modules/.bin/playwright install chromium export LIGHTTABLE_PLAYWRIGHT_MODULE="$PWD/build/processing-browser/node_modules/playwright/index.mjs" ``` `LIGHTTABLE_BROWSER_EXECUTABLE` optionally selects the test Chromium executable; otherwise Playwright's installed browser is used. Pinned film runtime/data preparation is shared with the existing Python CI setup. The untracked legacy `engine/spektrafilm-rs` binary is deliberately not used as the film gate's implementation under test.