# Feature parity: fabric-emulator vs. real Microsoft Fabric How the emulator's surface maps to real Fabric (as documented at [learn.microsoft.com/fabric](https://learn.microsoft.com/en-us/fabric/) / [MicrosoftDocs/fabric-docs](https://github.com/MicrosoftDocs/fabric-docs)), and β€” the point of this table β€” **whether real work happens or just the API shape**. The emulator's design bet is that the durable, testable surface is *contracts + storage + identity + orchestration*, and those are done for real (real signed JWTs, real Delta bytes on disk, real RBAC, a real pipeline interpreter, real cross-engine SQL, real Livy high-concurrency session packing). The heavyweight or proprietary **compute engines** run by default: `docker compose up` auto-loads an override that starts Sail (Spark Connect) and a SQL Server sidecar, so Livy sessions, notebook cells and the T-SQL warehouse do real work out of the box. What stays 🟠 is the narrower set that needs a *different* engine β€” the JVM overlay, or an opt-in profile (`rti`, `eventstream`) β€” and what cannot be done honestly at all is stubbed. ## What the denominator is **This table measures completeness against a self-selected surface, not against Fabric.** The rows below are the surfaces the emulator implements, chosen the way [07](07-control-plane-api.md) describes β€” an endpoint-frequency scan of `fabric-docs`, keeping what SDKs, `fabric-cicd`, git integration and deployment-pipeline automation actually call. That is a deliberate engineering choice and a good one; the ratio it produces is not a coverage figure. For scale: Fabric's REST reference publishes on the order of **880 operations across ~57 workload groups**. A green count here says the chosen surface is complete and proven, never that Fabric is covered. That figure is Fabric's published reference, and it is **not** the denominator `docs/surface-ledger.json` counts β€” that one reads the swagger vendored in this tree and totals 1002 operations, because it also counts Power BI's `/v1.0/myorg` surface and, for Fabric itself, can only count what is committed here. The two numbers are reconciled with their sources in [65](65-api-surface-coverage.md), which is also where the machine-written surface ledgers are explained. **Fabric-era workloads this map does not grade at all** β€” absent rather than red, so probing is the only way to discover them otherwise: external data shares, managed private endpoints, Digital Twin Builder, Data Agent, Org Apps, Graph models, Anomaly Detector, Ontology, Operations Agent, Datamart, paginated reports, Snowflake and Cosmos DB mirroring, Databricks catalog mirroring, and the Spark settings and Livy-monitoring endpoints. **Their ITEMS exist; their WORKLOADS do not.** Every one of Fabric's 50 documented `ItemType` values creates, reads, updates, deletes and round-trips its definition here, through the generic item surface and through the typed collection the reference prints for it. That is item CRUD and nothing more: a Data Agent does not answer questions, an Ontology models nothing, and `sqlEndpoints/{id}/refreshMetadata` is unimplemented. A client doing CI/CD over these items works; a client doing the workload's own job does not. The [Scope boundary](#scope-boundary-fabric-not-the-predecessor-azure-products) section below covers *predecessor* products (Synapse, U-SQL, ADLS Gen1). This paragraph is its Fabric-era equivalent. **"Real via our own wire-protocol implementation."** A row is 🟒 **Real** not only when an external engine/client does the work, but also when the emulator *itself* implements Fabric's wire protocol and the logic behind it β€” so a real, unmodified client gets byte- and behaviour-identical responses. Fabric's control plane, OneLake's ADLS/Blob surfaces, the Data Pipeline expression language + control flow, and the Livy **high-concurrency** session-packing layer are all in this category: no engine is being proxied, yet the observable contract matches real Fabric because we built the protocol, not a mock of it. Where a row's *execution* needs an engine that is not the default β€” JVM-only Spark surfaces, an opt-in KQL profile β€” that part is split out as 🟠 or πŸ”΄. ## Legend | | Meaning | |---|---| | 🟒 **Real** | Genuine work: real signed JWTs, real bytes on disk, a real engine/client computes, real logic enforced β€” no pretending. | | 🟑 **Emulated** | Faithful API contract + persisted state, but no engine β€” status is clock-derived / management-only. | | 🟠 **Non-default engine** | Real, but only on an engine that is *not* the default: the JVM Spark overlay (`docker-compose.spark-jvm.yml`) or an opt-in profile such as `--profile rti`. **Not** "bring your own": `docker compose up` already starts Sail and the SQL Server sidecar, so the everyday Spark and T-SQL surfaces are 🟒. | | πŸ”΄ **Not implemented** | Honest 501 or absent. | ## Platform / fundamentals | Fabric feature | Emulator | Type | |---|---|---| | Workspaces CRUD | Full. Display names are **unique tenant-wide** β€” duplicates 409 `WorkspaceNameAlreadyExists` (uniqueness per the REST reference; fabric-docs covers workspace naming portal-side only) | 🟒 Real | | Items CRUD + typed collections | Full. Display names are **unique per (workspace, type)** β€” duplicates 409 `ItemDisplayNameAlreadyInUse`; names stay reusable *across* types, which is why OneLake addresses items as `name.Type`. The `type` is validated against the documented **`ItemType` enumeration** (50 values, REST reference) β€” anything else is `InvalidItemType`, as real Fabric returns β€” and canonicalised case-insensitively so `notebook` and `Notebook` cannot become two types. 24 typed collections alias the generic surface; each collection segment is taken from its own reference page, since they are not derivable (`GraphQLApis` is capitalised, `variableLibraries` is not) | 🟒 Real | | Role assignments / workspace RBAC | Enforced from the validated bearer principal | 🟒 Real | | Folders | Full β€” list/create plus get, rename, delete (empty only), and move; a cycle is `InfiniteFolderHierarchyLoop` | 🟒 Real | | **Catalog Search** (`POST /v1/catalog/search`) | Cross-workspace item discovery over workspaces the caller can already see. `search` matches display name, workspace name and description; `filter` supports `Type eq`/`ne` with `or`/`and`/parens; `pageSize` 1–1000 (default 50). Dashboard and Dataflow are excluded as the reference requires. Returns `ItemCatalogEntry` (`catalogEntryType: FabricItem`). Does not grant data-plane access | 🟒 Real | | **Fabric Core MCP Server** (`POST /v1/mcp/core`) | Streamable HTTP JSON-RPC facade over the existing Core REST β€” same handlers, so RBAC, audit and LRO stay one path. Tool names match Microsoft's published Core MCP list. Does **not** run notebooks or write lakehouse tables (Microsoft's own limitation). GET is 405 (no SSE). Witnessed by the unmodified Python `mcp` SDK in CI. Distinct from local `Fabric.Mcp.Server` and from pbix-mcp | 🟒 Real | | **Fabric IQ MCP** (`POST /v1/mcp/fabriciq`) | Read-only MCP server over Power BI reports and semantic models, on the same transport as Core MCP. The six documented tools β€” `DiscoverArtifacts`, `ResolveFabricItem`, `GetReportMetadata`, `GetSemanticModelSchema`, `ValueSearch`, `ExecuteQuery` β€” run as the signed-in user: **delegated tokens only**, as Microsoft documents, so a service principal is refused before any tool runs; a Fabric or a Power BI audience token is accepted; `X-Variants: Fabric.Routing.FabricIQ.V1` is served and any other variant refused. Each tool needs Read on the item, not Build, and row- and object-level security apply through the loader `executeQueries` uses, so two users asking the same question get different rows. `ExecuteQuery` runs 1–4 DAX queries, 250 rows by default and at most 1,000, and refuses MDX, DMV and `INFO` functions; the DAX is the bounded subset, now with `ORDER BY`. `GetReportMetadata` reads PBIR and PBIR-Legacy definitions: pages, visuals and the fields they bind, filters at report, page and visual level (simple conditions restated, the rest kept raw), report-level measures, and the bound model. `queries` takes JMESPath, with `regex_match`. **Inferred, not published:** Microsoft does not document the tools' input or output schemas, so the argument names follow its Fabric IQ skill and the response documents follow the paths that skill queries. **Not modelled:** verified answers and AI instructions (present and empty), workspace-app URLs, the embedded CSV resource for large results, OAuth discovery. Witnessed by the unmodified Python `mcp` SDK in CI, as two users whose answers differ by their RLS role ([07](07-control-plane-api.md#fabric-iq-mcp)) | 🟒 Real (tool schemas inferred from Microsoft's skill) | | **Fabric Data Warehouse MCP** (`POST /v1/mcp/dataPlane/sqlEndpoint`) | Microsoft's MCP server for T-SQL, at both documented endpoints: the global one and one scoped to a single item. Its one tool, `execute_query(workspaceId, itemId, query)` (also answering to the Learn page's `executeSQL`), runs one batch on a Warehouse, or on a lakehouse's SQL analytics endpoint named by its endpoint id, **as the signed-in user**. It takes the TDS wire's own path: route and Read check, the read-only rules (a Viewer's session, the endpoint's data), Fabric's dialect and time travel, a login as the caller so that grants and row-, column- and object-level security are SQL Server's, and lineage and versioning for an accepted write. Returns the last result set as an embedded `text/csv` resource then `Query returned N rows.`, at most 10,000 rows, with a 300-second timeout; a SQL error is `Error -32002: `. **From a live capture, not Microsoft's docs:** the tool name `execute_query` (Learn says `executeSQL`), `serverInfo` `microsoft.fabric.sqlEndpoint`, the title, annotations, result and error shapes, as recorded by a third party (`iemejia/fabio`) and matching Microsoft's skills; the row and time limits are what the skills call observed defaults. **Not modelled:** the observed 20-requests-a-minute limit, OAuth discovery. Witnessed by the unmodified Python `mcp` SDK in CI on a real SQL Server, as two users whose rows differ by a row-level security policy ([07](07-control-plane-api.md#fabric-data-warehouse-mcp)) | 🟒 Real (contract from a third-party live capture and Microsoft's skills) | | **Eventhouse MCP** (`POST /v1/mcp/dataPlane/kqlEndpoint`) | Fabric's remote MCP server over one KQL database, at both endpoints: global (each call names `workspaceId` and `itemId`) and scoped to a single KQL database. The four tools the live server lists, `executeQuery`, `getSchema`, `getGeneralKQLExamples` and `getSpecificKQLExamples`, run as the signed-in user, who needs Read on the database (share-aware). `executeQuery` runs real KQL on the attached engine and answers a Kusto document whose rows are in `PrimaryResult`, capped at 1,000 silently; a failed query carries the cluster, the database and the engine's error. `getSchema` reads the real schema with row counts and sample rows. The grounding tools refuse an empty database (`Database is empty`); `clusterUrl` and `databaseName` redirect to another of the emulator's own databases. **From live captures, not Microsoft's docs:** Microsoft names no tools; the names, `serverInfo` `KustoMCP`, `executeQuery`'s schema and the result and error shapes are from two third parties' captures. **Ours:** the grounding documents and their ranking, which Fabric produces with Copilot. **Not modelled:** Azure Data Explorer clusters, learned examples. Witnessed by the unmodified Python `mcp` SDK in CI against kustainer, as two users ([07](07-control-plane-api.md#eventhouse-mcp)) | 🟒 Real (contract from two third-party live captures; grounding is ours) | | Capacities (list, assign / unassign) | Full state, no billing/SKU enforcement | 🟒 Real state | | **ARM-created capacities** (`FABRIC_ARM_URL`) | Opt-in poll of arm-emulator's `GET /_family/capacities`. ARM-created `Microsoft.Fabric/capacities` appear on Fabric `GET /v1/capacities` under the Fabric REST GUID ARM assigned at create. The seeded default stays. Empty `FABRIC_ARM_URL` is standalone (seed only). No job metering β€” SKU is a label | 🟒 Real (feed consume) | | Long-running operations (202 β†’ poll) | Clock-derived | 🟑 Emulated | | **Asynchronous outcomes forced** (`FABRIC_FORCE_LRO`) | The reference documents **two** outcomes for this API β€” 200 with the body, and 202 + `Location`/`x-ms-operation-id`/`Retry-After` whose `/result` carries the definition β€” and **a real tenant answered 202** (measured 2026-08-11 on a VariableLibrary). The emulator only ever produced the 200, so nothing here exercised the half a tenant actually takes, and the failure it hides is the quiet kind: a client that reads the 202's body gets `null` and reports an **empty definition** rather than an error. Exactly that happened when this repo first called the real API. **`createItem` is the sharper case and the reason this became one switch rather than two**: Create Warehouse documents 201 and 202, a real tenant answered 202, and it "does not support create a warehouse with definition" β€” while the emulator went async only FOR a definition-bearing create, so the one type measured asynchronous was the one type guaranteed synchronous here. A client indexing the 201 body got `None["id"]` against a tenant. So the async outcome is now available as a **testing lever**, the same idea as `FABRIC_LIST_PAGE_SIZE`: set it and every dual-outcome API forces the operation loop on the spot. **Three surfaces audited and covered**: `getDefinition` (200/202), `createItem` (201/202) and `git/initializeConnection` (200/202) β€” the last carries a real result body, unlike `commitToGit`/`updateFromGit` whose LROs have none. **Scope stated rather than implied**, with the full audit in [49-async-outcome-audit.md](49-async-outcome-audit.md): these are the surfaces checked against the reference, not proof no fourth exists. The durable output of that audit is a searchable rule rather than a list β€” **every surface that answers 202 says so in its description, in the words "This API supports long running operations (LRO)"**, and every clean one does not, so the next sweep is a grep rather than a reading exercise. Six further surfaces were checked and found clean (`git/connect`, `git/disconnect`, bulk `admin/domains` assignment, and mirroring, whose LRO routes to the already-async `updateDefinition`), which says the three fixed were anomalies rather than the front of a queue; the already-async ones (updateDefinition, commit/update from git, capacity assign, identity provision, deploy, job instances) were confirmed and needed nothing, and two documented LROs are *not implemented at all* (Load Table, `sqlEndpoints/{id}/refreshMetadata`) so they are gaps rather than wrong shapes. **Off by default** β€” 200 is equally legal and is what most calls see; the point is that the other path is now reachable rather than hypothetical, and both are asserted to return the same bytes. **A real client now drives the forced half in CI** (`ci:fabric-cicd-lro`): the same fabric-cicd publish as the synchronous job, against an emulator started with the toggle, so the poll loop that follows `Location`, honours `Retry-After` and reads the operation's error body is Microsoft's own and knows nothing about this emulator. Four arms, because a green publish alone would not tell them apart β€” the toggle demonstrably reached THIS server (a raw definition-less `createItem` answers 202, and its `/result` carries the item rather than the operation); an operation that is genuinely `Running` rather than one that succeeds on the first poll, with `Location` staying on the state url until it moves to `/result`; the same delay traversed by fabric-cicd itself; and an **injected failure**, which is the arm that proves the client parses the emulator's own `errorCode` instead of treating any 202 as success. **Mutation-tested rather than asserted**: making a succeeded operation keep pointing at itself is caught by fabric-cicd on the *synchronous* job too (a definition-bearing create was always an LRO here), while making the toggle skip `createItem` is caught ONLY by the new leg β€” which is what makes the new leg evidence rather than repetition. Nothing diverged: the emulator's 202 satisfied a client that had never seen it | 🟒 Real (both documented outcomes, one witnessed by Microsoft's own client) | | **Deleted display names held** (`FABRIC_NAME_RESERVATION`) | A tenant does not free a deleted item's display name at once: recreating a Notebook under the name of one just deleted answers **`409 ItemDisplayNameNotAvailableYet`**, *"is not available yet and is expected to become available in the upcoming minutes"*, with **`isRetriable: true`** (measured 2026-08-11). The emulator freed the name instantly, so a provision / teardown / re-provision loop passed here forever and failed on a real deploy β€” the permissive direction, which is the one that ships. **The distinguishing field is `isRetriable`, not the status**: a name held after a delete and a name taken by a live item are BOTH 409, and a client that treats them alike gives up on the one that would have succeeded seconds later. **A duration rather than a flag, because the tenant disagreed with itself** β€” the message says "the upcoming minutes" and the name was free again on a retry **20 seconds** later, so nothing observed licenses a constant and the window is the operator's to choose. **Off by default**, for `FABRIC_FORCE_LRO`'s reason: instant reuse is what a local create/delete loop wants, and the point is that the other behaviour is reachable at all before a tenant is the thing that finds out | 🟒 Real (shape and retriability; the window is policy) | | Item **job execution** (`jobs/instances`) | **Five item types really execute, and their status follows the work rather than the clock.** Each parks `CompleteAt` beyond the clock's reach at submit β€” so no poll can observe a clock-completed job whose work is still running β€” and finalises when the work reports: **DataPipeline** runs the interpreter, **CopyJob** moves the bytes, **Notebook** and **SparkJobDefinition** are executed by the Spark agent when one is configured (without an agent they stay open for the documented external callback, which is the honest contract rather than a hang), and **Apache Airflow Job** is decided by a real Airflow sidecar. An item type with **no execution engine** β€” a Lakehouse, Warehouse, SemanticModel, Reflex and most of the ~50-type enumeration β€” still accepts the POST, because Fabric's job surface is generic across item types, and its status is then derived from the virtual clock (`NotStarted` β†’ `InProgress` β†’ `Completed` as `CompleteAt` passes) rather than from work. The lifecycle is faithfully shaped; nothing ran. That is the honest grade for a type with nothing to execute β€” inventing a work-derived status there would be the lie the notebook reconciliation was built to kill. Dataflow's refusal is the one outcome still reached at submit, because a 501 really is instantaneous. **Capacity admission is real state, not a broker:** a per-capacity concurrent-job ceiling (default 999) throttles Manual submits with `430 CapacityNotAvailable` + `Retry-After`, and queues scheduled / event-triggered jobs as `Queued` until a slot frees, FIFO, on the same clock and list levers that fire schedules. Same-item jobs both occupy slots when capacity allows β€” the Delta collision Fabric allows, not a mutex. [36-capacity-job-queueing.md](36-capacity-job-queueing.md) | 🟒 Real (types with an engine) / 🟑 clock-derived where there is none | | **List Item Job Instances** (`GET …/jobs/instances`) | An item's runs, paged, newest first, with the same clock-derived status the single-instance read returns. Listing is itself an evaluation point for the scheduler and for capacity-queue drain, so a run that has come due β€” or a queued job whose slot has freed β€” appears without touching the control surface | 🟒 Real | | **Item Job Scheduler** (`…/items/{id}/jobs/{jobType}/schedules`) | Fabric's *own* per-item scheduler β€” distinct from the `ApacheAirflowJob` item, which delegates to a real Airflow sidecar. All four `ScheduleConfig` members: `Cron` (`interval` 1–5,270,400 min), `Daily`/`Weekly` (`times[]` ≀ 100, `weekdays[]`), `Monthly` (`DayOfMonth` or ordinal weekday, `recurrence` 1–12) β€” every documented bound enforced at write time, the 20-per-item ceiling as `ScheduleExceedsLimit`, and `localTimeZoneId` (Windows ids or IANA) honoured as real local wall time, so a daily 09:00 stays 09:00 across a DST change. Schedules **really fire**: each due occurrence starts a job instance down the same path a manual run takes, so a scheduled DataPipeline executes the interpreter and only `invokeType: "Scheduled"` tells them apart. Evaluation is driven by the **controllable clock** rather than a background worker β€” advancing time materialises exactly the occurrences that came due, and a past `startDateTime` triggers instantly with no special case. *Boundary:* catch-up is capped at 100 occurrences per evaluation (newest kept), because a controllable clock can be advanced a year against a one-minute Cron. [07-control-plane-api.md](07-control-plane-api.md) | 🟒 Real | ## Identity & security (`security/`, `admin/`) | Fabric feature | Emulator | Type | |---|---|---| | Entra OAuth2 tokens / JWKS / client-credentials | entra-emulator mints **real signed JWTs** | 🟒 Real | | Workspace managed identity handshake | Provisioned via entra admin API; the identity's own token passes RBAC | 🟒 Real | | Key Vault references in connections | Resolved against azure-keyvault-emulator | 🟒 Real | | **Governance domains** (`/v1/admin/domains`) | Full admin surface: domain/subdomain CRUD (the hierarchy stops at two levels, as documented), workspace assignment (single-valued β€” re-assigning moves a workspace), `nonEmptyOnly` listing, and bulk `Admins`/`Contributors` role assignment. Deleting a domain cascades to its subdomains, assignments and roles. **Tenant-admin gated**, and graded as Microsoft documents it rather than uniformly: a READ (`admin/workspaces`, `admin/items`) requires *"a Fabric administrator **or** a service principal"*, while a WRITE (`admin/domains/create-domain` and every other mutation) requires *"a Fabric administrator"* with **no service-principal escape** β€” so a service principal may read the tenant and may not change it. Administrators are declared by the operator (`FABRIC_TENANT_ADMINS`), never inferred from a token claim; with none declared every mutation is refused, because the pre-gate behaviour was that everyone was an admin. Refusals are `InsufficientPrivileges`/403, the documented code | 🟒 Real (mgmt + tenant-admin gate) | | **Activity log / audit** (`GET /v1.0/myorg/admin/activityevents`) | Real audit trail β€” the emulator **records events as operations happen**, using the documented audit vocabulary (`CreateWorkspace`, `CreateArtifact`/`UpdateArtifact`/`DeleteArtifact` from admin/operation-list.md; `InsertDataDomainAsAdmin` & co. with their `DataDomainObjectId`/`FoldersToSetCounter` properties from governance/domains-audit-schema.md). Enforces the documented request rules (single-quoted UTC bounds, same-day window) and pages with `continuationToken`/`continuationUri` until the token stops coming back. Nothing is synthesised at read time | 🟒 Real | | **Tenant settings** (`GET /v1/admin/tenantsettings`) | The documented `TenantSetting` object in full β€” `settingName`, `title`, `enabled`, `canSpecifySecurityGroups`, `tenantSettingGroup`, the three `delegateTo*` flags, enabled/excluded security groups (`graphId`/`name`), and typed `properties` validated against the documented `TenantSettingPropertyType` enum. Optional arrays are **omitted, not null**, as the reference's sample does; seeded with the setting names that sample uses. `POST /v1/admin/tenantsettings/{name}/update` is the **real** update API (`enabled` required; response wrapped as `{"tenantSettings":[…]}`) | 🟒 Real | | **Tenant-wide workspace admin** (`GET /v1/admin/workspaces`) | The documented `Workspace` shape β€” note it differs from the user-facing surface: the envelope key is `workspaces` (not `value`) and the field is `name` (not `displayName`). Filters `type`/`state`/`capacityId`/`name` are enforced, with undocumented enum values returning `BadRequest` as the reference specifies; `domainId` is reported from the real domain assignment. Emulator workspaces are always `Workspace`/`Active` (no soft delete), so `state=Deleted` is legitimately empty. **Tenant-admin gated on the read rule** β€” an administrator *or* a service principal, as this API's reference states | 🟒 Real (mgmt + tenant-admin gate) | | **Tenant-wide item admin** (`GET /v1/admin/items`) | The documented `Item` shape across every workspace, with `workspaceId`/`capacityId`/`type`/`state` filters and the reference's own error codes (`InvalidItemType`, `InvalidItemState`). Envelope key is **`itemEntities`** β€” a third spelling after the user-facing `value` and admin workspaces' `workspaces`, which is why each came from its own reference page. `Active` is the only documented state, so that is all the emulator reports. Fields it does not model (`creatorPrincipal`, `defaultIdentity`, `tags`) are omitted rather than faked | 🟒 Real (mgmt + tenant-admin gate) | | **Capacity tenant-setting overrides** (`GET /v1/admin/capacities/delegatedTenantSettingOverrides`, `POST …/{capacityId}/delegatedTenantSettingOverrides/{name}/update`) | The documented `CapacityTenantSetting` (a tenant setting plus `delegatedFrom` and `delegateToWorkspace`, and *without* `delegateToCapacity`/`delegateToDomain`). The update body is the documented one β€” `enabled` required, no `properties`, no `delegatedFrom` β€” and the response wraps as `{"overrides":[…]}`. An override may only be created for a setting whose `delegateToCapacity` is true. **Note there is no `/v1/admin/capacities` list API in Fabric**: capacities are listed on the Core surface at `/v1/capacities` | 🟒 Real (mgmt + tenant-admin gate) | | **Sensitivity labels** (`bulkSetLabels` / `bulkRemoveLabels`) | The documented admin bulk APIs, reporting per-item `successfulItems`/`failedItems` rather than failing whole calls. Every change writes the documented `SensitivityLabelEventData` to the audit log β€” `SensitivityLabelApplied`/`Changed`/`Removed` with `SensitivityLabelId`, `OldSensitivityLabelId`, `ActionSource` 3 (Manual), `ActionSourceDetail` 5 (PublicAPI), `ArtifactType` 12, and a `LabelEventType` genuinely computed from label order (upgraded/downgraded/same-order/removed). **The label taxonomy is emulator-provided** β€” real Fabric gets labels and their order from Purview, which cannot be attached offline; `GET /v1/admin/labels` exposes it. Gated as writes (administrator only), but note what the gate does **not** model: these two APIs are documented `User: Yes / Service principal: No`, so real Fabric refuses a service principal *even when it is an administrator*. The gate models the administrator requirement, not Microsoft's per-API Entra identity-support matrix | 🟒 Real (label APIs + audit) / 🟑 taxonomy | | **Sensitivity labels β†’ catalog** | Labels are Purview *Information Protection* objects, not Atlas entities, so the Atlas-API route a Purview β†’ OpenMetadata migration takes cannot carry them β€” assets, classifications, glossary and lineage cross; labels do not. The optional OpenMetadata profile exports them instead: the taxonomy becomes classification **`FabricSensitivity`** (`mutuallyExclusive`, since an item carries at most one label), each label a tag under it, and an item's `sensitivityLabel` is applied to that item's catalog entity. CI reads the tag back **through OpenMetadata's API**, on two items with different labels, and asserts that clearing the label in Fabric clears the tag in OM ([22-openmetadata.md](22-openmetadata.md)) | 🟒 Real (stored in OM) | | **Purview Data Map** (`{endpoint}/datamap/api/atlas/v2/…`) β€” *split out of the old single "Purview scanning / classification" row, because its three parts have different **ceilings**: this one is reachable and moving, scanning is reachable in-family, and the built-in classifiers can never move. One πŸ”΄ made them look equally stuck, and became actively wrong the moment this landed* | The Data Map **is Apache Atlas v2** β€” the Azure spec's own TypeSpec declares `@route("/atlas/v2/entity")`, `/types`, `/relationship`, `/glossary`, `/lineage` and annotates them *"This is Atlas API"* β€” so this is Atlas semantics, not a Purview-shaped faΓ§ade. Implemented: the **type system** (typedef CRUD, per-category reads, `typedefs/headers`) and **entities** (createOrUpdate, read/delete by GUID, `/entity/bulk`, lookup by unique attribute), on the documented `purview.azure.net` audience with the documented `AtlasErrorResponse` codes. What is enforced rather than shaped: an entity naming an unregistered type is refused; required attributes are checked **through the supertype chain**, which is how `qualifiedName` β€” declared once on `Referenceable` β€” is required of every type; `qualifiedName` is the per-type identity, so createOrUpdate genuinely updates instead of duplicating; Atlas's **negative-GUID placeholder protocol** is honoured and reported in `guidAssignments`; a batch validates fully before it writes anything; deletes are soft, as the spec's `EntityStatus` states. The four Atlas base types are seeded, as a real account has them. **Scope, explicitly: 96 routes in the spec, and this is the type system and entities only** β€” glossary, lineage, relationships, classifications, business metadata and search are NOT implemented. 🟑 rather than 🟒 for one reason: the row's evidence is 14 Go tests, and this repo's rule is that green needs a **real-client witness in CI**. `pyapacheatlas` β€” an off-the-shelf Apache Atlas v2 client, not a client shaped like this implementation β€” now drives these routes in CI (`e2e/purview-datamap`), which is what moved the grade. **The green covers the type system and entities ONLY**; the witness is deliberately scoped to them, because a suite extended to the unimplemented routes would fail honestly and a suite extended until it passed would not be evidence | 🟒 Real (type system + entities only; witnessed by `pyapacheatlas`) | | Purview **scanning** (`scan/`) | Not implemented. A scan is an **engine**, not a surface: it connects to a source, enumerates assets, infers schema and applies rules over real values. Against the emulator's own sources β€” OneLake Delta, the warehouse β€” that is genuinely computable, so this row has a **reachable ceiling of 🟒 in-family / πŸ”΄ external**, which is why it is not bundled with the row below | πŸ”΄ Not implemented | | Purview **system classifiers** (the ~200 built-ins) | Not implemented, and **this one can never move**. The built-in classifiers are proprietary detection patterns published in no OpenAPI spec and in no documentation; custom classification rules are implementable, the shipped set is not. Split out from scanning precisely because its ceiling is different: scanning is blocked by work, this is blocked by information nobody outside Microsoft has | πŸ”΄ Not implemented (no reachable ceiling) | | **Lineage** (catalog graph) | Via the optional OpenMetadata profile: OneLake **shortcut** edges and executed pipeline **Copy** sourceβ†’sink edges are persisted exactly and witnessed in OM's graph API. Notebook edges come from the engine's own report of what it read and wrote, or from the data plane observing it; Script/stored-procedure code is still not guessed. Catalog SSO can also be pointed at entra-emulator ([22-openmetadata.md](22-openmetadata.md)) | 🟒 Real (shortcuts + Copy) | ## OneLake (`onelake/`) | Fabric feature | Emulator | Type | |---|---|---| | ADLS Gen2 DFS surface (create β†’ append β†’ flush, ranged read, list) | Full, incl. the `x-ms-range` dialect | 🟒 Real (real bytes) | | Blob surface | Full | 🟒 Real | | Delta commits (put-if-absent atomicity) | Real; `-race`-tested concurrent-commit race | 🟒 Real | | **Item permissions** β€” sharing one item without the workspace | **Granted, validated and reported; not yet enforced.** Effective access is the workspace role's implied permissions unioned with a direct grant β€” so revoking a grant leaves what the role gives, as the product documents. Semantic models are shared through **Power BI's documented dataset-users API** (`GET/POST/PUT …/datasets/{id}/users`, and the in-group spellings), with its documented rules: Post needs ReadReshare, Get and Put need ReadWriteReshare, Write can be neither added nor removed, an App cannot be a target, `None` removes. Every other item is shared through an **authenticated emulator-native** `…/items/{iid}/_emulator/access`, because Fabric's Core REST has no item-permissions operation and the portal uses calls Microsoft does not document β€” not under the unauthenticated `/_emulator/` prefix, where anyone could grant themselves ReadAll. The documented **admin list** `…/admin/workspaces/{wid}/items/{iid}/users` reports effective access, which is **inferred**: no sample shows an inherited row. A grantor shares at most what they hold; a UPN is refused by name, since there is no directory to resolve it; a group is stored but matches nobody. **Enforced on OneLake**: a ReadAll grant admits a Viewer, or a principal with no workspace role, to that item's data on DFS and Blob β€” reading and listing inside it, never enumerating the workspace β€” and `principalAccess` answers for grantees too. All of it comes from one shared decision, `store.OneLakeReadAccess`, so no two surfaces can disagree about the same principal. When the item has OneLake security roles, the roles decide and ReadAll reads through `DefaultReader`. **Enforced on Direct Lake and `executeQueries`**: Build is required to query, and a grant on a Direct Lake source is read through the same decision. **Enforced at the SQL endpoint**, through the real relay against a real SQL Server: Read connects, ReadData reads, sharing never writes β€” and a revoke **stops working**, including by three-part name from another granted database, because no access means `REVOKE CONNECT` there, which also defeats an explicit T-SQL `GRANT` an owner once authored. Witnessed grant-and-revoke on every surface, and mutation-checked: with memberships only ever added, or with CONNECT left in place, the SQL witnesses fail. The admin list's inclusion of inherited access remains an **inference**. **A limit a real client found:** Microsoft's own `fab acl ls ` (1.2.0, the locked version) reads that admin list and then fails inside fab on `principal.displayName` β€” a field Microsoft's spec marks *optional* (`Principal` requires only `id` and `type`) and that fab evidently relies on real Fabric sending (unconfirmed without a tenant). This emulator has no directory to source a name from, and omits a field it cannot source rather than inventing one, so the response is spec-conformant and unusable by that client. Asserted as a failure in `e2e/fabric-cli`, so it cannot close unnoticed. [57](57-item-permissions.md) | 🟒 Real (the admin list's inherited rows inferred; `fab acl ls ` cannot read it) | | **OneLake security roles** (`dataAccessRoles`) | Real: roles are stored, evaluated deny-by-default, and a role WIDENS a Viewer without narrowing Contributor and above. `DefaultReader`'s virtual membership (`fabricItemMembers`) is modelled, because an evaluator with only explicit Entra members cannot express it β€” matched on the principal's **real item access**, and only when the member's `sourcePath` names this item: it was ignored until 2026-09-13, so a role written for holders of ReadAll on *another* item admitted ReadAll holders here. DENY rules are refused rather than silently ignored, as the product refuses them. **Row and column security are read from the documented `constraints` object, keyed by `tablePath`** β€” and the decision rule is decoded **strictly**: a restriction the parser cannot read with certainty (an unknown field, the old flat `rows`/`columns`, an empty column list, a non-`Read` column action, two constraints on one table) drops the rule rather than granting past it. Until 2026-09-13 flat fields were read instead, so a policy authored the way the REST reference documents it was enforced as **no policy** β€” and every enforcement witness wrote the flat dialect, so none could see it ([54](54-onelake-security.md#the-authoring-payload-was-the-wrong-shape-and-every-witness-spoke-it)). Dropping a rule is always safe: a constraint narrows only the rule it sits in. Witnessed with the reference's own sample payloads. **The item type is checked on write**: `Lakehouse`, `MirroredDatabase` and `MirroredAzureDatabricksCatalog` carry roles and everything else is `DataAccessRolesNotSupported` β€” a **Warehouse** most of all, since its data is secured by T-SQL and a role stored there would be one the emulator honours and a tenant ignores. Two enum members the supported-items table does not name by spelling (`MirroredWarehouse`, `MirroredCatalog`) are refused rather than admitted on inference: a wrong refusal comes back as a bug report naming the type, a wrong acceptance as a policy that did not hold in production. The GET is **not** gated, because the docs do not say what a read against an unsupported item returns; that asymmetry is pinned by a test rather than left to drift. **Unsupported role combinations are refused, in the one shared decision.** "OneLake security doesn't support the combination of two or more roles where one contains RLS rules and another contains CLS rules … users … receive query errors": a principal whose row filter and column restriction on a table come from different roles is marked on that table by `onelakesec.Effective`, where the union used to open both and grant everything. Direct Lake and direct DFS/Blob reads raise the error by name; `principalAccess`, whose response has no field for an error, gives an engine **both** restrictions, never more than either role grants; the SQL analytics endpoint shows the reader no rows. Judged per table β€” inferred from "tables that are part of an unsupported role combination". Both kinds in one role remain supported. | 🟒 Real | | **Row- and column-level security, applied by the engine** | Real for Spark SQL and the catalog-resolved DataFrame readers: two callers get different rows and different columns from one `SELECT`. **Constraints are per table**: one rule can grant `*` and filter only `Tables/sales`, and each consolidated entry carries its own union, so the most specific entry decides β€” the Go evaluator and the Spark agent now read it the same way, where they had quietly disagreed on overlapping grants. The filter replaces the RELATION rather than rewriting user SQL, and every qualified registration of a secured table is swept, so `default.sales` cannot walk around it ([measured](54-onelake-security.md)). **Refuses** where the engine's catalog is shared, because reshaping one there would take the table from everyone. Under the two-context split the same filtering is done by the privileged half instead, and witnessed separately. | 🟒 Real | | **Two-context execution** β€” the user context holds no service credential | Real. A statement against an item with OneLake security roles runs in a child process carrying a token minted for the CALLER, with the agent's storage bearer and `ENTRA_CLIENT_SECRET` scrubbed from its environment; the privileged half reads the Delta files, applies the row filter and column projection, and hands over only the result. A table no rule grants is never registered there at all. **Enabled per item, not globally** (`FABRIC_TWO_CONTEXT`): the split exists to enforce policy, and where there is no policy there is nothing to enforce. [docs/54](54-onelake-security.md) | 🟒 Real | | **Direct path access blocked for a narrowed principal** | Real, and enforced by OneLake rather than by the engine. RLS and CLS cannot be applied to bytes, so a principal who may see only part of a table cannot fetch its files β€” on the DFS and Blob surfaces alike, and **from a notebook cell**: a secured session runs on an engine of its own carrying that caller's token, so `spark.read.load("abfss://…")` arrives as the caller and is refused 403. That is Fabric's own mechanism, not a patched reader. Sharing one engine across callers is what used to defeat it, and Fabric shares one only within a [single-user boundary](54-onelake-security.md) too. | 🟒 Real | | Shortcuts (OneLake β†’ OneLake) | Symlinks with target-side RBAC (trusted-workspace-access) | 🟒 Real | | Shortcuts to external targets (S3 / ADLS Gen2 / Dataverse) | **Real Amazon S3 / S3-compatible read-through**: a Connection carrying the Access Key Id + Secret Access Key pair β€” as `Basic` credentials, because Fabric's S3 connector uses authentication kind "Access Key" while the REST reference's `CredentialType` enum has no `AccessKey` member and `Basic` is its only two-secret type β€” makes the emulator sign upstream requests with **AWS SigV4** (`internal/awssig`, verified against AWS's own published example signature). Witnessed by `e2e/s3` against a real **SeaweedFS** server started with an identity config β€” the suite proves an unsigned GET is 403 and that a *wrong* secret is refused, so the pass cannot be vacuous; the object itself is written by **boto3**. **ADLS Gen2 reads are witnessed too**: `e2e/azurite-shortcut` drives a real SAS through the shortcut against **Azurite**, Microsoft's own storage emulator, proving an unauthenticated GET is 403 and a *tampered* SAS is refused. Scope stated honestly β€” Azurite implements Blob/Queue/Table and **not** ADLS Gen2, so the endpoint is the Blob one; that witnesses the read path (a plain authenticated GET, identical on both endpoints of a real ADLS Gen2 account) but **not** DFS-specific behaviour, which stays unwitnessed offline. **Dataverse is now a real target type**, at a scope worth stating precisely: Microsoft documents the target's four fields (`connectionId`, `deltaLakeFolder`, `environmentDomain`, `tableName` β€” the only documented target with no `location`) and documents that the tables must already exist in the Dataverse Managed Lake. It does **not** document that lake's byte layout, so the emulator serves ordinary Delta from `environmentDomain/deltaLakeFolder/tableName` and **that composition is ours, not Microsoft's**. This is a target type in the existing shortcut machinery, *not* a Dataverse emulator: no Dataverse Web API, no OData metadata, no table discovery. The **read-only rule is real and enforced**: Fabric states twice that "Dataverse shortcuts are read-only. They don't support write operations regardless of the user's permissions", and both the flush and the delete paths refuse, with `e2e/azurite-shortcut` confirming via the Azure SDK that the refused write never reached the target and the refused delete left it intact. Auth diverges and is disclosed: Fabric's delegated model is organizational account (OAuth2) or service principal, while the emulator presents whatever the Connection carries β€” the witness uses a SAS, so what is proven is that the credential is really presented and really checked, not that OAuth2 delegation works. **S3 read-only is by design** β€” Fabric documents S3 shortcuts as read-only regardless of permissions, so a write there is refused locally and never reaches the target. **ADLS Gen2 write-through is real**: a file flushed under an ADLS shortcut is PUT to the storage account and the local buffer dropped (it would otherwise shadow later target changes), and deleting a file within a shortcut deletes it at the target, as documented. Witnessed by `e2e/azurite-shortcut`, which reads the result back **with the Azure SDK straight from Azurite** rather than trusting the emulator's own 200 | 🟒 Real (S3 SigV4 + ADLS reads) / 🟠 Dataverse (opt-in layout) | ## Data Engineering (`data-engineering/`) > Engine-level claims below are measured, not asserted: the generated > [Spark engine matrix](engine-matrix.md) probes each capability against both > Sail and the JVM overlay and is regenerated in CI. | Fabric feature | Emulator | Type | |---|---|---| | Lakehouse item + Tables/Files storage | Full (via OneLake) | 🟒 Real | | Notebook authoring / definition round-trip | Full | 🟒 Real | | `notebookutils` / `mssparkutils` β€” **the whole documented surface** | All **eight** namespaces Microsoft's overview table lists β€” `fs`, `notebook`, `credentials`, `lakehouse`, `runtime`, `session`, `udf`, `variableLibrary` β€” with **all 44 documented members** present, each carrying the documented parameter names in the documented order. The module list is Microsoft's, not ours, so it can name a module nobody here thought to look for. **A second Microsoft source now sits beside the transcription**: `third_party/notebookutils-stubs/` pins the `dummy-notebookutils` wheel Microsoft publishes, and `check_notebookutils_surface.py` holds both β€” the docs decide behaviour, the stub decides what exists, and where they disagree the divergence is COMPUTED rather than declared. Holding them together turned "we transcribed carefully" into a list of 15 gaps, all now closed: the whole `help()` family (Fabric's `fs` page opens by documenting `notebookutils.fs.help()` as the discovery mechanism, and it was on no module β€” prose on the page rather than a row in the method table, so transcription could not yield it), `getHelpString`, `runtime.getCurrentWorkspaceId`, `fs.nbResPath` and `fs.refreshMounts`, `lakehouse.getDefinition`/`updateDefinition`, and `udf.run`. **Two of those gaps had gone stale before anyone reread them** β€” `fs.nbResPath` was waived as "Axis C lists it as absent" after Axis C landed it, and `udf.run` as "the UDF item has no engine here" after `getFunctions` began running the item's own code, which is why the check fails on a waiver that is no longer a gap. Graded on every conformance run by contract 2 on both lakehouse backends. Getting there fixed **25 defects**: 17 members absent, and 8 that existed and would still have been declined because the parameter NAMES differed (`dst` for `dest`, `recursive` for `recurse`, `maxBytes` for `max_bytes`) β€” a framework introspects a signature and refuses to start without ever calling it. Plan and measurements in [56-notebook-capability-parity-plan.md](56-notebook-capability-parity-plan.md). **Shape is half the claim; behaviour is the other half.** All 44 members are also EXERCISED end to end β€” 42 by `ci:notebookutils`, and `session.stop`/`session.restartPython` by `ci:livy-native`, which is the only harness shaped right for them: a Livy session is a persistent REPL, so one statement can restart the session and the NEXT one can report what survived. Each confirms out of band β€” the call that acted is never the call that answers. Giving the last members an e2e found **four defects that shape could never have caught**: `fs.rm` deleted a non-empty directory without `recurse=True` (ADLS answers 409 `DirectoryNotEmpty`, so the emulator was more destructive than the thing it emulates); the LRO `Location` header hardcoded `https://`, so a client following it β€” the documented way to reach a result β€” could not reach an emulator started with `-disable-tls`; `session.stop`/`restartPython` raised *"running outside a notebook session"* **inside** one, because nothing in the tree ever set the environment variables they read, and in a secured session the identity never crossed into the user-context child; and `restartPython` left the session without `display`, `displayHTML` and `%run`, which are kernel builtins on Fabric and cannot be lost to a Python restart | 🟒 Real | | **`notebookutils.notebook` management** (`create`, `get`, `getDefinition`, `update`, `updateDefinition`, `delete`, `list`) | The seven documented CRUD methods, with the documented parameter names in the documented order. **Added because framework-conformance contract 2 measured all seven ABSENT** ([38-framework-conformance.md](38-framework-conformance.md#2-the-api-shape-is-the-contract-independent-of-behaviour)) β€” and absence is what a CI/CD framework reads before it will start, without ever calling one. `create` takes `.ipynb` as Fabric documents it, and the executable `notebook-content.py` is derived **server-side** so the created item can actually run: before that, a documented create returned 201 and its `RunNotebook` job then died on `notebook-content.py is missing`. The 200-or-202 outcome is followed rather than read as a body, which is what `FABRIC_FORCE_LRO` exists to force a client to prove | 🟒 Real | | `notebookutils.notebook.runMultiple` β€” DAG structure and ordering | Runs a DAG in dependency order, accepting a bare list or Fabric's `{activities, timeoutInSeconds, concurrency}` shape. A failed activity's dependents are reported skipped with the reason rather than run. A cycle, a dependency naming an activity outside the DAG, a duplicate activity name, or a malformed shape is refused. `validateDAG` exposes those same checks ahead of a run, and `workspace` accepts a name or an id | 🟒 Real | | `notebookutils.notebook.run` β€” exit values | Returns the child's exit value, the exact string passed to `notebookutils.notebook.exit(...)`, and `""` when the child never called it. Applies to `runMultiple`'s per-activity `exitVal` too. Until 0.18.0 this returned the terminal job status, so a parent branching on a child's exit value took one path here and the other on Fabric | 🟒 Real | | **`notebookutils.notebook.run` β€” exit values OVER REST** | **A shim limitation on the real target, NOT an emulator divergence β€” the emulator is the side matching Fabric here.** In Fabric you call `run()` inside a notebook kernel, which has an in-process channel to the child's exit value; that contract is the row above and the emulator implements it. This shim also offers `run()` from OUTSIDE, over REST, and there the value is read from `…/jobs/instances/{jid}/notebookRun` β€” the emulator's own endpoint, paired with the `notebookRunResult` engine callback. **Real Fabric answers it `404 EntityNotFound`** (measured 2026-08-11), so a REST-submitted run cannot carry an exit value at all. Reported as `None` β€” *could not be obtained* β€” rather than `""`, which claims the notebook exited with no value and is what this returned until 2026-08-11, so every run on a tenant looked like a deliberate empty exit. Both are falsy, so `or ""`/`or 0` are unaffected. The shim now WARNS on the real target naming the two paths that work: call `run()` inside a notebook, or have the child write a one-row Delta table. Fixing this by making the emulator report `None` would move it AWAY from Fabric's documented contract | 🟑 REST path only; notebook path 🟒 | | `runMultiple` β€” failure contract | Raises `RunMultipleFailedException` (importable from `notebookutils.common.exceptions`) when any activity did not complete, carrying every result on `.result`. Each value carries Fabric's `exitVal` and `exception` keys; `status`/`error` ride alongside as emulator extras for local debugging and nothing should depend on them | 🟒 Real | | **Materialized lake views** | **The refresh is real**: the view's own query runs on the engine the emulator already hosts and the result is written as a **Delta table under `Tables/`**, so a Delta reader, the SQL analytics endpoint, the flow stream and lineage all see a refresh exactly as they see any other write. A refresh is only reported successful if **a Delta commit actually landed** β€” checked against OneLake rather than inferred from the statement's exit code, because a cheerful statement that wrote nothing is precisely how a view whose rows do not exist would be reported `Materialized`. **Staleness is measured, not assumed**: the declared sources' Delta versions are recorded at refresh time (read **before** the query, so a write landing mid-refresh cannot be counted as data the view read β€” forced in the witness by advancing a source from inside the engine fake) and compared with what those tables are now; the answer names *which* source moved. A view that has never refreshed is `NeverRefreshed`, not stale β€” there is nothing there to be out of date. *Emulator-native definition surface, and deliberately:* Fabric defines these with Spark SQL DDL that no capture here has observed, and this repo does not invent a syntax β€” so the definition hangs off `…/lakehouses/{id}/materializedlakeviews`, on the same precedent as the Reflex trigger binding, and `dependsOn` is **declared** rather than parsed out of the SQL (a wrong parse would not fail, it would silently report a stale view as fresh). Dropping a definition leaves the materialised table: the rows are real data a reader may still be using. When a DDL capture arrives it becomes a second front door onto this model, not a rewrite | 🟒 Real (refresh + staleness) / 🟠 definition surface emulator-native | | Reference-run lakehouse rule | A referenced child notebook bound to a **different** default lakehouse than its parent is blocked, matching Fabric; a child that declares none inherits, a matching one runs, and `useRootDefaultLakehouse` in the arguments bypasses the check. The failure names the cause and the way out rather than "The job failed." | 🟒 Real | | `runMultiple` β€” retry and timeouts | Per-activity `retry`/`retryIntervalInSeconds` (last error reported, no trailing wait after a final attempt) and the DAG-level `timeoutInSeconds`, defaulting to Fabric's 12 hours. `timeoutPerCellInSeconds` is applied per **cell**, multiplied by the notebook's real cell count **where that count is readable** β€” it comes from `…/notebookRun`, which real Fabric does not serve, so on a tenant the shim falls back to a whole-notebook ceiling instead. It previously assumed **one** cell there, which handed a twelve-cell notebook a one-cell deadline and produced a spurious timeout for a run that was working | 🟒 Real | | `runMultiple` β€” concurrency | An explicit `concurrency` is honoured with a bounded pool per dependency level, `0` meaning unlimited; a level boundary stays a barrier and results keep their listed order regardless of completion order. **The default diverges deliberately**: Fabric defaults to 3Γ— the CPU count, the emulator to sequential, because a harness comparing two runs needs the same sequence more than it needs speed | 🟑 Real, sequential by default | | `runMultiple` β€” child Spark sessions | Fabric runs children on isolated REPL instances **within the parent's Spark session**, so they share its compute and its session-scoped state. The emulator gives each child its own session, so sibling activities cannot see one another's temp views. Deliberate: the benefit is narrow and the change is large. Scoped in [39-run-multiple-parity-plan.md](39-run-multiple-parity-plan.md) | 🟑 Isolated sessions | | Spark session / statement / batch via the **Livy API** | **Native termination** (`--spark-agent-url`): the emulator implements the Livy contract and drives a persistent statement-executor agent. The default agent is a PySpark Connect client of Sail, not Apache Spark. An external Livy backend remains configurable with `--spark-livy-url` | 🟒 Real (default engine) | | Notebook **cell execution** | The emulator parses and records the notebook run, resolves the attached lakehouse and **applies a bound Environment to the session before the first cell runs**, and Sail executes cells by default. The same fixture runs on Spark 3.5 JVM and proves unqualified `saveAsTable`/`spark.table` bind to OneLake `Tables/` | 🟒 Real (default engine; per-capability detail in [engine-matrix.md](engine-matrix.md)) | | Notebook **cell languages** (`%%sql`, `%%pyspark`, `%%configure`, `%%html`, `%%markdown`, `%%spark`) | A cell's language now decides what happens to it. Before this the parser RECOGNISED the magic and the run loop sent everything that was not `sql` to the **Python executor** β€” so correct Scala came back as a Python `SyntaxError` pointing at the user's own code, and a `%%configure` block of JSON failed the same way. Four dispositions: run, render, ignore, refuse. `%%html`/`%%markdown` are content and are rendered, never executed. **`%%configure` is accepted and IGNORED, out loud** β€” the cell records that it changed nothing and that the requested executors, memory and conf were not applied; refusing would be worse, since it must be the first cell on Fabric and the results it would produce are correct, just not on the requested hardware. Scala, R and C# are **named**, with what to do instead and explicitly "Real Fabric runs it", so the message cannot be read as "your cell is invalid". An unknown magic still falls back to Python, because Fabric adds magics and refusing every unheard-of one would break notebooks on upgrade. **Witnessed through a real RunNotebook job** (`ci:conformance-sail`), not only by unit tests on the classifier: the defect was never in `Disposition()` β€” the parser always classified the magic correctly and the RUN LOOP ignored the answer β€” so a test on the classifier passes on both sides of the bug. Every cell in the probe is chosen to be INVALID PYTHON (`%%configure`'s `false`, the `

` markup), so a regression cannot pass quietly; reintroducing the old behaviour was measured, and it fails with the original signature, `NameError: name 'false' is not defined`. The Scala leg asserts the refusal NAMES the language and that the run STOPPED β€” proven by the absence of a marker the next cell would have written, read over DFS rather than taken from the job record | 🟒 Real (Python and SQL execute; other languages refuse by name) | | **`%run`** | Splices another notebook's code into **this** session, so what it defines is usable on the next line β€” the whole difference from `notebookutils.notebook.run`, which starts a separate session and returns an exit value. A line magic inside an ordinary Python cell, so it is a source rewrite applied before `ast.parse`: a `%run` inside a string or a comment is left alone, and source with no `%run` comes back byte-identical. Parameters bind before the referenced code runs, nested `%run` expands, and a non-Python cell in the reference is skipped rather than spliced. **Proven through the real agent** β€” the only harness that can, since the notebookutils e2e runs its notebook as a plain script and `e2e/notebook-run`'s runner is itself the engine | 🟒 Real `ci:conformance-sail` | | **`display()` / `displayHTML()`** | Notebook BUILTINS, bound in the session namespace β€” a notebook writes `display(df)` with nothing above it. They were **absent entirely**, so one of the most common lines in any Fabric notebook raised `NameError`. `summary=True` is honoured rather than ignored, because Fabric documents it as column name, type, unique values and missing values β€” a data QUALITY read a notebook branches on, so there IS something to switch. **Under a kernel it publishes a MIME bundle** β€” `text/html` plus a `text/plain` alternative, both built from one description of the data so they cannot drift; without one (the agent's RunNotebook path) it still prints, which is what every stdout-reading suite asserts on. "Nothing here to render into" was true of the agent and false of the repository: the `jupyter` profile ships a real JupyterLab against this same stack, and its kernel now binds Fabric's `display` rather than IPython's β€” the two are not interchangeable, and `display(df, summary=True)` is a TypeError to IPython's. **Witnessed by `ci:notebook-display`**, which runs nbclient against the kernel from that same image and reads the notebook's own outputs: a stdout reader cannot tell a published bundle from a print, which is how this row stayed unevidenced. Removing the bundle was measured β€” the harness reports the `stream` output it fell back to. Column names and cell values are escaped, so a value holding `