--- name: performance-profiling description: Diagnose or improve LunCoSim FPS, physics time, periodic stalls, Builder versus View differences, terrain/render cost, or Tracy captures. Use when performance must improve architecturally without changing BigSpace, substeps, shadows, terrain quality, or image fidelity. --- # Performance profiling: measure the owner, then remove avoidable work Read the current handover in [`docs/reviews/open-400fps-performance-handover.md`](../../docs/reviews/open-400fps-performance-handover.md) before changing code. A status-bar FPS number is a symptom, not an attribution. ## Required separation - Run one production session that you own for the product FPS number, launched from the task checkout on a verified free API port. Do this even if another user's or agent's session is active; never control, stop, restart, or change the scene in a pre-existing session. Report concurrent GPU/CPU workloads and classify affected numbers as contention-affected, not clean acceptance. Do not use Tracy's instrumented number as acceptance evidence. - Run a separate Tracy build/capture using the adjacent `../tracy` checkout; start `tracy-capture` before the production binary and inspect the settled window, not only startup. Check the profiler listener and connection against the owned PID with `ss -ltnp`/`ss -tnp`. A concurrent Tracy client can occupy 8086 and move the owned app to the next port; pass that port with `tracy-capture -p`. A capture from another process is not evidence for the task scene. - Compare Builder and View with the same scene, camera, rendering-quality settings, physics substeps, shadow settings, and terrain assets. A Builder- only cost usually means an editor observer, rebuild, projection, or UI path, not that physics needs a different global timestep. - Attribute CPU, GPU, physics, terrain, and UI separately. Never disable shadows, lower authored terrain quality, change BigSpace, or change the standard substep count to make a graph look better. ## Architecture checks Search the owning systems for unconditional writes, full-set topology scans, repeated observer registration, per-frame allocations, polling, and work that should be gated by a revision/change event. Structural edits should invalidate structural caches; transform propagation and telemetry output are not by themselves topology changes. Check both the Builder and View registration paths before fixing only one. For a system that queues many compatible ECS component changes, inspect its `system_commands` flush separately from the system body. Collect changed values and use the owning crate's existing batch command path where it preserves the same commit boundary; retain change filtering so a quiet fixed cycle queues nothing. In panel paint code, use `PanelCtx::resource` for reads. Its `resource_scope` temporarily removes and reinserts the resource, which marks it changed even when the closure only reads it. Scope a mutable borrow only around a real state transition; otherwise a steady Builder repaint can wake change-detection systems and full-scene reconciliation on the next Update. USD authored/composed text painting borrows `UsdViewportState` and its cached source strings. Check `usd_preview_render_layer_propagation` when diagnosing preview invalidation; report traversal cost separately from frame time. Opening several previews must not add bodies, colliders, or joints to the mounted scene. For SysML requirements rebuilds, inspect `sysml_requirements_rebuild` input flags. Domain notification publishers check `has_pending_events` before mutably draining their registry; an empty queue must not invalidate readers. The shell's steady `WorkbenchSnapshot` check compares borrowed dock and perspective iterators before materializing owned vectors; preserve that allocation-free stable path when changing layout publication. The top-level menu's unique label projection is built only when the menu registry changes. Width measurement borrows that projection and streams labels without collecting another vector; compute perspective-tab rows once per UI pass and reuse them for both width measurement and painting. Do not rebuild the same responsive menu inputs during steady repaint. For contributed top-level menus, defer collecting and sorting scripted contributions until the popup callback runs; a closed menu should only paint its row and publish its anchor. Use `workbench_custom_menus_render` to inspect that row path separately from the workbench aggregate. When a measured UI snapshot builds several indexes from the same entity population, combine compatible marker reads into the existing query and avoid another full-population traversal for each marker set. Keep queries over different populations separate unless measurements show that a broader scan costs less. For large Builder hierarchies, derive a lightweight row index from the cached tree and current expansion state, retain that index across repaints, then use `ScrollArea::show_rows` to create widgets only for rows in the viewport. Rebuild the index only after a source revision, active filter/scope change, or a branch disclosure change reported by the shared tree helper. Keep foldout state keyed by stable entity/group identity and preserve selection, drag, and tooltip behavior on painted rows. Immediate-mode widgets are expected to be repainted; the domain tree and entry records should be borrowed or shared, not cloned into panel-owned snapshots. Borrow selected-entity state during paint and retain an owned selection snapshot only when it changes; compare egui temporary values through borrowed type-map access instead of cloning a cached vector every frame. Gate panel-owned view models with `WorkbenchSnapshot::is_panel_visible`, and order their systems after `WorkbenchSnapshotPublishSet`. Hidden dock tabs have no reader and should not rebuild view data each frame; keep separate cleanup work transition-driven when a panel closes. In the Builder Spawn palette, use the catalog owner's revisioned category-to-entry index. Borrow category labels and spawn entries, and retain formatted display labels only until that revision changes. Do not rescan all entries for each open category or build cloned entry groups before egui determines which categories have a reader. For Builder Ports, retain matching and expanded/collapsed entity indexes by topology revision, filter, and expansion state; update the sampling request only when the expanded entity set changes. In the Inspector's material part selector, keep the entity index separate from its display text. Format the active label only for the selected part and format the other labels only while the dropdown is open. Retain the material-bearing entity index for projected USD roots by selected root and `UsdStageRevision`; rebuild it only when either changes. Recompute for non-USD roots or when the revision resource is unavailable. When deriving a chosen part's material controls, filter the selected root's existing material-bearing entity list instead of walking that part's child tree again. For live line plots, do not copy and decimate a full history in every UI frame. Keep one bounded build per binding, snapshot changed histories at a limited presentation cadence, transform immutable sample snapshots on the async-compute pool, and keep painting the last completed point buffer while the next build runs. `ScalarHistory::snapshot()` shares completed chunks and copies only its bounded open tail on the caller; flatten and decimate that snapshot on the worker. In Tracy, compare `line_plot_history_snapshot` (app-thread copy time) with `line_plot_series_build_worker` (worker time and input/output point counts); several plots can rebuild concurrently even when each panel's paint span is short. When live curves share an experiment plot, also inspect `line_plot_scalar_history_snapshot`, `line_plot_scalar_history_build_worker`, and `multi_series_plot_points_build_worker` to account for both history flattening and plot-point preparation. Auxiliary Graphs overlays should use the same snapshot cache and keep painting their last completed buffer while a new one is built. Key transformed point buffers by source identity, transform mode, and pixel width; apply line style while painting. For experiment and live-overlay plots, retain transformed `PlotPoint` buffers by immutable source identity, log-Y mode, and display width; build them on the async-compute pool and borrow them while drawing. Min-max decimate time-sorted series to about one sample per logical display point on that worker so line geometry does not traverse every retained history sample each repaint, preserving narrow peaks. Cache each buffer's full-data bounds so egui auto-fit does not scan all samples on every repaint. Keep experiment variable groups and positivity summaries in the change-gated view model instead of regrouping names or scanning all sample values during graph painting. Time-series hover may use monotone-X interval pruning; keep the dependency's general segment search for phase-space curves where X can reverse. For the entity tree, derive parent and grid facts through indexed lookups along named candidates' deduplicated ancestor closure instead of copying every scene entity's `ChildOf` and `Grid` membership into the snapshot. Keep topology invalidation separate from the query-heavy snapshot producer: mark the revision dirty when scene facts change, then gate snapshot work until no prior worker is active. This lets a revision change reject an in-flight tree result without re-entering the full ECS query system just to discard it. For Builder telemetry, inspect `telemetry_catalog_patch_capture`, `telemetry_catalog_patch_worker`, and `telemetry_catalog_patch_commit` separately. Each patch handles at most 64 changed descriptors and relevant ancestor facts. There is one initial identity enumeration; later notifications update the persistent index, affected alias groups, and tree paths. Selection, samples, and unchanged hierarchy writes must cause zero descriptor preparation. Use `InspectTelemetryCatalog` counters and owner timings to verify this in an owned production session, and `scripts/api/test_telemetry_catalog.py` for repeated metadata edits with authored Rhai verdicts. Separate initialization from settled frame costs and worker time from app-thread capture/commit time. Use `workbench_panel_render` child zones inside `render_workbench` to attribute painting and visible-row index costs independently of descriptor work. For startup asset graphs, separate asynchronous source reads from discovery, composition, and UI/physics admission. Read all known dependencies in each breadth-first frontier through bounded batches, then merge results in stable authored order; this avoids serializing independent branches behind one parent at a time. Preserve the source-change reload graph when using child load contexts. The USD layer-closure owner and its reload receipts are documented in [`21-domain-usd`](../../docs/architecture/21-domain-usd.md#composition-closure-and-partial-scene-loading). Keep CPU-heavy source compilation off asset I/O workers too: register literal dependencies first, then prepare source ASTs on the existing async-compute pool. Startup discovery must keep filesystem enumeration, manifest reads and parsing, and installed-artifact checks off the app thread. Merge the resulting typed registry snapshot and publish readiness on the owning app schedule. Open Twin manifest scans can overlap, but commit in discovery order and discard results when their owning Twin closes. Reuse authored Rhai source classifications for the same manifest and policy revision; asset-content changes do not require a second classification pass. Application session metadata such as recents follows the same boundary: load, normalize, and persist it on workers, merge typed results on the app schedule, serialize writes, and finish the final write during shutdown. For Modelica startup, separate Rumoca compile time, prepared-solve cache lookup, `lower_for_live`, and ordered result commit. When equivalent requests share the same structural solve key, coalesce in-flight lowering instead of occupying a second worker; preserve one result position per participant in the existing stable commit queue. Keep the combined compiler/root/lowering admission bounded to the solve-pool worker count plus two staged operations, so the serialized Rumoca owner can compile later programs while solve workers are busy without building an unbounded DAE backlog. Application policy startup has separate Tracy spans for `application_policy_source_prepare_offthread`, `application_policy_compile_offthread`, `application_policy_prestartup_wait`, and `application_policy_activation`. The activation span separates `application_policy_clear_registry`, `application_policy_validate_order`, `application_policy_validate_prepared_hooks`, `application_policy_install_prepared_hooks`, and `application_policy_publish_registry`. Preparation starts during runtime plugin construction; the compile span includes authored installer-order selection. Owner filtering and per-hook contract/arity validation run in `application_policy_filter_manifest_offthread` and `application_policy_validate_prepared_hook_offthread`. `PreStartup` validates the selected order and its manifest records, then registers the prepared callables before Startup consumers. Keep source text and callables in the Rust preparation bundle; the application selector receives hook identities only. Compare the activation children before moving additional registry work across the lifecycle boundary; preserve authored install order and visible failure reporting. Read-only USD projectors use `CanonicalStages::reader_for` or `reader_for_entity`: generation-zero reads consume the worker-prepared plan, and later authored generations consume the live canonical stage. Do not call `get_or_build` just to read startup facts. Use the prepared schema/path indexes to find initial candidates. Change batches should advance unrelated edits and invalidate only owners whose consumed paths changed. If a live generation has no exact worker-prepared plan, snapshot the canonical stage recipe on its owner thread and submit replacement facts through shared bounded `AsyncWorkAdmission`; cap this owner's pending stages, then commit only when asset-plan identity and canonical generation still match. Hold startup progress only until the first topology index commits. For later generations retain the last committed facts so current simulation continues, while new or reprojected prim admission stays queued until replacement facts commit. A required preparation failure, including a host without worker transport, must be visible through the scene fault owner. Do not run a whole-stage topology traversal synchronously in `Update`. For route-edit latency, record four separate spans: the bounded Rhai input hook, the durable route `ApplyUsdOps` and its one incremental projection, reference-marker admission, and `UpdateUsdCurveView`'s stroke preparation and terrain annotation-image publication. Fragment lookup uses bounded spatial bins; verify idle camera/terrain changes do not rebuild the stroke index. Ribbon presentation preparation is presentation-only and must not issue another `ApplyUsdTransientOps`, change document generation, or run inside the fixed scenario event that observes a projected route. Compare the UI hook and fixed tick against frame/physics budgets; a fast worker result does not excuse a slow synchronous query or document edit in the input path. For general `SpawnEntity`/`DeleteEntity`, measure the command, USD add/remove, live structural reconciliation, and referenced asset admission separately. Keep edit-to-projection wall time separate from frame time. A new referenced entity may still be preparing its source plan across frames; each live instance retains that prepared asset so later matching spawns are warm. Measure the first spawn, a warm spawn, and deletion separately, while checking render-frame and fixed-step cadence during each. Use the production spans `scene_spawn_entity_command`, `usd_apply_ops_document`, `usd_document_apply_change_set`, `usd_reference_layer_closure_merge`, `usd_reference_instance_plan_remap`, `usd_reference_root_author`, `usd_reference_instance_promotion`, `usd_live_structural_reconcile`, `usd_visual_projection_batch`, `scene_runtime_spawn_commit`, `usd_live_subtree_despawn`, `scene_delete_entity_command`, `scene_delete_entity_persist`, and `usd_apply_one_op_document` to separate command, document, reference, projection, and removal costs. Warm instances of one immutable asset recipe should bypass layer-byte cloning and comparison; asset reloads use a new recipe identity and still merge changed bytes. Path reconciliation uses the lifecycle-maintained stage/path index, not a full population scan per changed prim. Raw-file spawn identity collision checks use the lifecycle-maintained API identity registry and a hash set of queued roots, not a world scan or a walk of all pending inputs. Document projection waits only for references named by that edit's typed AddPrim operations; an unrelated reference must not extend its completion time or document-projection hold. An authoritative pending reference can retain its separate `SceneReferences` simulation hold until its own instance is committed. A long completion wait is not itself a frame stall, but an unready authoritative reference can hold simulation progress. A normal document-backed spawn projects from the prepared instance plan and authors only a lightweight live root; `usd_reference_root_author` must not include reference composition. Measure `usd_reference_instance_promotion` separately because the first later edit that needs full composed instance facts performs that one-time promotion. Deleting an unpromoted root must not promote it. `scene_runtime_spawn_commit` measures each changed prim's live ECS structure insertion, including deferred reference descendants. `usd_visual_projection_batch` measures the bounded per-update component and visual projection pass; keep its frame cost separate from total time until the last descendant is admitted. The `usd_reference_instance_plan_remap` span must remain independent of asset prim count: an instance shares the immutable prepared snapshot and carries only its namespace and root overrides. The stage test verifies snapshot sharing with `Arc::ptr_eq` and verifies that an instance view exposes only the source asset's `defaultPrim` subtree. For Twin-open stalls, profile the active Twin policy loader separately from policy activation. Native manifest and Rhai source reads should run through bounded `AsyncWorkAdmission`; `twin_policy_source_prepare_offthread` measures that preparation and `twin_policy_activate` marks the lifecycle-bound commit. Verify that stale completions are discarded after a Twin switch and that the active `assets_mounted` plan and first authoritative tick wait for activation, with the app/UI schedule continuing. Browser builds still use synchronous WebStorage reads because they do not have a worker transport. One-shot `RunRhai` and tool callbacks share the prepared `ScenarioDriver` engine. When profiling a callback stall, separate one-time startup/prelude or tool-generation refresh from request execution; do not rebuild the engine for each callback. For Bevy visibility costs, distinguish the active scene camera from auxiliary shadow-map subviews. If adapter capability supports GPU culling, put `NoCpuCulling` on the scene camera so camera frustum work can move to GPU preprocessing while per-mesh light and shadow visibility remains CPU-owned. Do not put it on shadow-casting `Mesh3d` entities: Bevy excludes those from its CPU-built per-light visibility lists. Unsupported adapters retain CPU camera culling. Measure GPU headroom and compare a clean FPS run plus a separate Tracy capture; do not trade away shadows or authored quality to reduce CPU time. Local-light shadow views are filtered at the render boundary: compare a point light's finite range sphere or a conservative bound of each extracted spotlight frustum with every active extracted `Camera3d` frustum and compatible `RenderLayers` before Bevy prepares shadow views. Do not mutate authored lights, omit offscreen cameras, or infer relevance from the light origin alone. Uncertain bounds and boundary cases keep the map. This saves only maps provably irrelevant to all outputs; Bevy's main-world per-light caster visibility pass is still a separate cost to measure. Bevy's camera driver also executes `Core3d` for point/spot shadow roots. Admit camera-only Core3d stage sets only for camera roots; keep the shared shadow passes and the GPU preprocessing needed by the depth maps active on light roots. Verify the same retained shadow pass inventory before and after so this scheduling optimization cannot silently remove shadows. Treat Bevy `Changed`/`Added` filters as population filters, not free events: a no-match query can still inspect candidate entities, and separate `is_empty()` queries can repeat that work. Combine compatible invalidation sources into one `Or` query when they drive the same decision. If this remains a hot path, audit every writer before adding a source-owned event/revision/dirty set; once all writers are accounted for, prefer that signal over an always-evaluated `Added` population query in a run condition. Keep readiness checks separate from reconciliation requests. An unresolved async participant may require a cheap readiness check on later frames, but a shared revision that wakes full topology or causal-graph work should advance only when topology or endpoint lifecycle facts actually change—not merely because readiness is still pending. Guard `ResMut` revisions by comparing the owner's state before calling a mutating method: Bevy marks the resource changed on mutable dereference even when an idempotent method leaves its value alone. `SimComponent` input/output shape is tracked by `lunco-port-core::ScalarPortMap`. Its identity key changes at insert/remove/clear boundaries, while numeric sample writes leave it stable; use borrowed-name `set` in hot publishers and strict `set_existing` for already-declared inputs so stable keys do not allocate. A strict write separates whether the name exists from whether its value changed. Callers that bypass Bevy change detection should mark the component changed only when the sample differs; generic `InputPorts` follows the same rule. `PortMap` keeps authored names at its boundary and resolves them to dense, process-local slots for the fixed propagation data plane. Map-backed ports, static scene-property inputs, shader inputs, Avian ports, and catalogued link-class outputs use owner slots. Shader live values are predeclared for driven parameters, and authored shape fingerprints update only at structural setters. Every owner that can change a declared surface publishes the shared `PortTopologyRevision`; compiled wire handles are rebuilt on that revision or on connection changes, and a stale slot is never retried through a name lookup. Numeric samples do not invalidate the compiled fabric. Do not re-hash every map when `SimComponent` changes for ordinary physics values. `PortHolds` advances its own revision only for effective intent changes; rebuild its target-index projection on that revision or wiring changes and reuse the aligned value buffer on steady ticks. Do not probe every target against the hold table or clone its presentation snapshot per tick. Readable `inputs:*` connection sources use a distinct input-side slot reader; write-only inputs can be targets but never become fabricated readable sources. Success diagnostics use entity-indexed borrowed-name lookups and retained tick scratch. Compiled targets cache surface presence until topology invalidation; pending/broken snapshots share compiled names as `Arc`, and fault warning keys are formatted only on first failure rather than every fixed step. For deferred USD projectors, keep one entity-work set per owner and feed it from the complete lifecycle boundary: identity arrival, projection readiness, invalidation, removal, and scene teardown. The same applies to deferred adapter steps such as wrapping a Modelica model into its shared port surface. Keep a single bootstrap discovery for entities predating plugin installation, and retry only work whose authoritative stage/readiness input is still pending. When a projector caps per-update work, select the bounded prefix in stable owner order and retain the remainder; avoid sorting the entire pending batch on the UI thread before applying that cap. For render-free USD discovery, wake candidates from changed prim identity, canonical stage generation, or stage-asset events, and keep only transient runtime prerequisites in the per-frame retry set. Dormant `BasisCurves` without the owner's relationship must be revisited when their stage changes, not queried from USD every update. The celestial admission projector uses the shared `PendingEntityWork` queue for one bootstrap and later lifecycle arrivals; its idle run condition reads that owner instead of scanning all USD prims. Keep `CelestialProjected` on the scene root after its source classification because static-light resolution consumes that boundary. Ordinary non-root prims do not need a completion command. Preserve authored-field candidates with a `lunco:` property, the scene root, `DistantLight` parent-body fills, and non-root `LunCoEpochAPI` diagnostics. When initial candidate processing spans many owners, queue the bootstrap IDs once and drain a fixed-size batch in stable entity order per app update; keep the remainder queued so a large scene cannot monopolize one UI frame. For USD visual projection, read both the system body and its Tracy `system_commands` flush: `frame_budget` bounds prim binding but cannot bound the deferred command batch applied after the system returns. Cap prim work items per update as well as elapsed time. Direct-child admission has its own `child_spawn_budget` and `max_child_spawns_per_update`; the parent keeps its awaiting/projecting markers until its direct children enter the ECS queue, so scene readiness must remain held while batches drain. For large procedural terrain fields, profile both scatter execution and its deferred entity commands. Admit generated body bundles and visual components in stable bounded batches while physics continues. Keep the terrain's applied marker pending until bodies and required visuals are committed. Cancel queued entries on refresh and teardown. Gate sparse lifecycle work on an existing pending marker or owner queue; an idle Update should not scan lifecycle state for a request that did not arrive. For joint admission, use one combined query over the existing `PendingUsdJoint` and `PendingJointAdmission` markers to gate the readiness scan. Reconcile candidates before the admission commit so the stable batch boundary is preserved. When a worker prepares facts for a composed plan, result commit must use the same prepared-plan identity for any derived cache. If the canonical owner records that exact plan as the source of the current live-stage generation, prepared facts remain valid at that nonzero generation until a live edit advances it. Later generations require complete change history or generation-keyed main-thread extraction. Otherwise a cache-key mismatch can repeat a full-stage scan during result publication. For dependent-stage refresh, pair `usd_canonical_stage_asset_sync` with `usd_sim_prepared_topology_cache` and `usd_sim_joint_topology_scan` to verify that publishing the new asset plan neither reopens the canonical stage nor repeats its prepared topology extraction. When initial preparation already builds composed type or API-schema indexes, run one-time topology and vehicle-output extraction against the prepared reader on a worker and use its indexed candidate query. Combine overlapping schema queries and property-prefix checks into one candidate traversal, so live canonical readers do not walk the stage once per consumer. Return each candidate's type, API-schema, and property-prefix facts with its path so later classification does not repeat native schema reads. Build per-prim candidate sets and authored vehicle-port lists from the prepared snapshot; skip extraction for unrelated entities while preserving their readiness marker. Do not materialize every prim path on the UI thread to find a small set of joints, attachments, vehicle roots, or policy prims. Live canonical edits keep using their owning-thread reader. Invalidate cached topology from `UsdSceneChangeBatch`: resynced paths that match an indexed source or currently carry a relevant schema, and info changes on indexed source prims, require refresh; unrelated paths only advance the cached generation. Transform-only info changes on topology source paths also advance the generation without a topology rebuild; structural resyncs or mixed edits still refresh. Check intervening changes before accepting prepared worker output. If another live read or the change observer has already made the cache current at that exact generation, discard the late prepared result instead of scanning the stage again. `usd_sim_prepared_topology_cache` records whether this check found a current-generation cache; pair it with `usd_sim_joint_topology_scan` to confirm that a repeated reconciliation did not scan the stage. For whole-index projectors such as USD telemetry, use one initial bootstrap, then coalesce relevant insert/remove observers into an invalidation flag. Keep stage-generation and asset-store invalidation as scalar checks. Do not repeat the same `Added`/`Changed` population filters in both the run condition and the projector, and do not reproject until the index is invalidated. When structural projection arrives in bounded batches, keep the index dirty and coalesce further invalidations until the owning queue settles; clear derived outputs once for that dirty interval. Keep processed entity identities in the owner's index so prim progress does not enqueue a deferred ECS marker command for every prim; clear the set with the stale outputs on invalidation. This still skips ordinary non-declaration prims on later projector passes without imposing marker insert and removal flushes on the app schedule. When a USD reader already exposes `has_authored_attribute`, use it to test one known property instead of enumerating every attribute name. If one enumeration feeds multiple derived port sets, derive them together from that single result. Temporary lookup indexes over immutable ECS queries should borrow path and port surface data instead of cloning those maps for a one-pass reconciliation. Build compatible per-entity indexes in one query traversal rather than running separate full-population passes for each index. For Modelica telemetry metadata, build `ModelicaIndex::component_name_lookup` once per dirty metadata batch and resolve all variables through that borrowed lookup. Do not call the linear `find_component_by_leaf` scan once per variable; steady sample batches should use the session's cached metadata without walking the document index. Cache each runtime producer's stable `SignalRef` identity for the session and borrow it when recording an existing channel. Avoid reconstructing and cloning signal paths for every due sample. For burst channel publication, keep retention depth as a logical limit and let the history buffer grow with recorded samples. Do not reserve every channel's full retention window when most new channels contain only their initial sample. When multiple runtime consumers inspect the same composed owner, acquire its canonical reader and child list once, derive owned typed facts for each consumer, then release the stage borrow before mutating ECS. Preserve the same typed resolution and commit helpers across initial and live projection paths; do not introduce a second fact cache. The presentation `project_env_settings` owner retains authored exposure and bloom facts with the exact stage-plan `Arc` and the matching environment prim paths. Reuse them across transform-only batches and info edits outside those paths; environment-path edits, resyncs, generation gaps, or plan replacement require a fresh composed read. Compare the stage signal without allocating a replacement vector on steady updates. When multiple USD consumers need the same composed-stage fact, put its generation/instance-keyed cache at the shared fact owner and reuse that cache; clear it at the scene teardown boundary instead of keeping consumer-local copies. For derived marker sets, compare current membership with the desired set and apply only additions/removals; unrelated rebuilds must not emit lifecycle churn. Removal invalidation should be qualified by the entity's authored USD identity and relevant endpoint capability, not by a generic component removal alone. Extract a USD program's declared interface once at admission and reuse it for validation, diagnostics, and publication instead of re-enumerating attributes. Derive a domain root's synthesizer once alongside member-role discovery and reuse the root-scoped selection during projection; do not traverse a large component collection again to repeat its role-schema queries. In Tracy, compare `domain_member_role_discovery` and `domain_synthesizer_selection_cache` against the enclosing `project_domain_islands` interval to attribute any remaining app-thread tail. Keep invalidation domains distinct: a wiring/topology latch may be raised by endpoint arrivals and must not automatically trigger domain discovery. Live canonical edits publish `UsdSceneChangeBatch` with stage generation and resynced/info paths; route those paths through the owning stage/root/member index. Transform-only info paths advance the generation without re-reading a Modelica network; structural and other member/root changes still invalidate affected roots. Stage-asset changes and missed generation batches requeue only roots on the affected stage. Reserve the all-prim discovery for initial admission, and keep entity arrivals on their queued-entity path. Generated-source document sync should share the `PendingEntityWork` contract, and its metadata publisher should consume the owner-published dirty flag rather than adding a parallel change query. When gating a dependency's multi-system transform schedule, distinguish its per-update output flags from authoritative input changes. Preserve any required settle pass after a floating-origin change, and test the change → settle → idle sequence so a gate neither stays open forever nor closes before propagation. Keep the settle marker between admitted passes so the idle admission check does not replace expensive propagation with a full-grid scan. For physics, distinguish persistent environmental state from transient commands: use the engine's persistent acceleration or passive non-waking force contract for gravity and contact support, and reserve waking force/torque writes for authored drive, braking, or actuator commands. Do not emulate this with a timer, sleep threshold, or a second cache. Use the existing cache/revision owner and preserve the USD projection boundary. Do not add a second cache, a timer, a compatibility path, or a quality fallback. If the hot path is not established, stop after source inspection and capture a bounded profile rather than guessing. ## Evidence For whole fixed-step timing, use the `PhysicsPerformance` query's `step_time_samples_ms` history in every host that installs the shared USD physics runtime. It returns retained `PhysicsTotalDiagnostics.step_time` samples, one per completed physics step; compute p50/p95/p99/max from that array after the measurement window. The query also returns the current `step_time_ms` and `step_number`, and the editor publishes the same current sample through `engine-health.physics_step_ms`. Query once per window because `PhysicsPerformance` also counts live topology and is not a per-step sampler. Record the clean FPS window, physics and render timings, Tracy capture path, scene/settings, and whether the result is startup or settled. Rebuild the production binary after a code change and repeat one clean A/B plus one Tracy diagnostic capture. Link the changed owner and state any platform/GPU evidence that was not available. For fixed-step versus UI diagnosis, query `SimulationTimingProfile` during the settled owned run. It summarizes the latest 240 fixed ticks and 240 app-update fixed-loop bursts: p50/p95/p99/max service time, per-tick realtime service-budget exceedances, fixed steps per app update, fractional overstep, and simulation-time demand clipped by `Time::max_delta`. Pair this with `QueryTelemetryHistory` for `engine.frame_time`, which covers the full app frame. The profile is read on demand and is not a replay or UI-isolation verdict; a nonzero clipped-time total means wall-clock demand was already excluded by the current policy, not that recoverable ticks remain queued. For CPU outliers, report p50/p95/p99/max for the app-thread frame and the authoritative fixed-tick transaction, plus fixed steps per app update and simulation deadline/backlog. Do not infer these from mean Avian time alone or add overlapping Bevy schedule spans. Bevy drains accumulated fixed schedules synchronously before `Update`: a fixed `Time` delta is not proof of a constant wall-clock physics rate, and a catch-up burst delays UI/input. Never improve UI timing by silently discarding authoritative overstep. If both wall-clock physics cadence and UI isolation are required, measure the whole simulation-owner boundary and consume immutable snapshots from the UI/render side; separate cycle labels alone do not provide thread isolation. When `tick_rhai_scenarios` or `tick_rhai_scenario_visualization` has an outlier, compare the nested `rhai_user_scenario_hook`, `rhai_user_event_hook`, `rhai_native_task_tick`, `rhai_mission_event_projection`, `rhai_prelude_scenario_driver`, and `rhai_user_visualization_hook` spans. Their fields identify the scripted entity, hook/driver, and queued-event count. If the outer scenario system is slow while these spans stay short, the time is in driver traversal or elsewhere in the exclusive fixed pass; do not attribute it to Rhai evaluation or move authoritative World access to a worker on that basis. For exclusive systems that borrow `&mut World`, inspect individual Tracy event timestamps and durations, not just aggregate averages. Align long calls with Twin-open and readiness milestones: one startup outlier can stall UI even when the same system is nearly free on settled frames. If the outer system is hot, attribute time to its internal owner operations before choosing an async boundary or cache. When `system_commands` dominates after a producer, inspect the number and shape of its deferred writes. For a large set of entities receiving the same bundle, use Bevy's existing batch command when its fallible/overwrite semantics match the owner, and preserve the owner's iteration order when it determines stable IDs. Relationship bundles such as `ChildOf` can use the same fallible batch: Bevy runs their relationship hooks per entity, in batch order, while continuing past stale entities. For `sync_local_gravity_to_avian`, keep the comparison against the existing `ConstantLinearAcceleration` component and batch only changed values in query order. Measure its `system_commands` flush separately; removals remain driven by the existing `RemovedComponents` streams. For `reconcile_frozen_subtrees`, profile readiness freeze and release flushes separately from the system body. Batch repeated joint, rigid-body, collider, and ownership-record inserts in held-root traversal order, with joint disables queued before endpoint body disables; retain the chained joint-release boundary after body restoration. For `pose_to_position`, treat `PhysicsPoseSeeded` as one-time readiness metadata. Insert it only on the first valid physics-pose write; later external pose updates still write position/rotation and update the bridge shadow, while sleeping bodies wake through Avian's `Sleeping` removal hook. For `apply_pending_forces`, leave already-zero accumulators untouched so idle bodies do not publish a `PendingForces` change every fixed tick. Clear every nonzero command after application or while physics is held, and retain the terminal non-finite fault path. For `process_queued_usd_visuals`, include queue preparation in the frame-budget review. Reuse its system-local child-key scratch set across updates, and keep duplicate-child identity checks scoped to the parent being admitted after the budget check; preparing keys for every queued parent can turn a small projection slice into a full-queue UI stall. For live composed-stage discovery of a known USD type, use the typed `UsdReadObject::prim_paths_matching` query. It preserves live traversal semantics while avoiding a full materialized path list followed by a separate type lookup for every prim. For fixed-step telemetry, keep static channel-presentation facts borrowed in the per-sample path and allocate their owned signal metadata only when that channel's metadata is first created or changes. Reuse one `HashMap::entry` for each sample's cached identity lookup and mutation, and update its global-owner association only when the owner identity changes. Add the shared `SignalSource` owner marker only when the entity first retains a sample; its removal observer owns history cleanup, so later batches do not need to queue marker writes. Dependent USD-stage refresh is owner work behind `sync_twin_overlays`. Snapshot the base/runtime revisions, serialize each persistent source snapshot once on bounded workers, share its bytes across dependent stages, and coalesce changed layers per target stage. Compare recipe bytes on workers and skip unchanged rebuilds. The main-thread owner validates source revisions, operation/revision, and target-plan identity before opening the live stage and committing the plan. Measure serialization and plan preparation separately from the live-stage build, reset owners, and visual projection; do not move the thread-affine stage across the worker boundary or let completion order select a commit. For a first-mounted document with a non-empty view layer, profile `view_ops_since_source_baseline` separately from the global operation journal. The initial recipe already contains base and runtime layers; only the bounded view suffix needs replay, while an expired view suffix still requires a full composed-source rebuild. For `project_usd_policies`, a cold cache can use the worker-prepared plan as its baseline only when every `UsdSceneChangeBatch` from generation zero to the live generation is present and proves that no policy prim changed. Existing cached facts use the same complete-generation and affected-path checks when promoted. A missing batch, generation gap, plan replacement, or policy-affecting change must use full live extraction. Compare startup and settled edit captures after changing this path; compile evidence alone does not establish a timing gain. For movement across terrain LOD bands, inspect `rebind_changed_shader_look` and source reflection before blaming picking or physics. `shader_source_validate` and `shader_source_schema` execute once per loaded shader revision; streamed tiles and material replacements reuse their facts. Mark the exact idle/movement windows in Tracy messages, keep the initial avatar pose fixed, and seed a pointer in SceneView when reproducing native picking. Distinguish shader reload events from material/look changes. Verify invalid-stage diagnostics and hot-reload invalidation as well as frame time.