--- name: fhirpath-test-designer description: > Design and generate comprehensive FHIRPath test suites using input domain partitioning and Pathling's DSL test framework. Use this skill whenever the user asks to write tests for a FHIRPath feature (function, operator, type system behavior, traversal pattern, etc.), review existing test coverage, identify missing test cases, or discuss what dimensions a feature needs testing across. Trigger on phrases like "write tests for", "test coverage for", "what tests do we need for", "review tests for", or any mention of testing a specific FHIRPath feature. Also trigger when the user mentions testing dimensions like singular/plural, empty propagation, cardinality, or HAPI resources in the context of FHIRPath tests. --- # FHIRPath Test Designer You design and generate FHIRPath test suites by combining specification research, input domain partitioning, and Pathling's fluent DSL. Your output is a test matrix reviewed by the user, followed by generated test code. The DSL reference below is authoritative — it was derived from `fhirpath/src/test/java/au/csiro/pathling/test/dsl/`, and `DslApiContractTest` in that package exercises every construct documented here. If a method you want does not appear below, read the package rather than assuming it exists. `references/DSL_Testing_Strategy.md` sets out the reasoning behind the partitioning approach — read it when deciding whether a dimension is worth testing, or when justifying a matrix to a reviewer. ## Workflow Three phases. Present findings to the user between phases. Accepts `--unattended` when dispatched by a caller that has no user to present to (e.g. `implement-pathling --unattended`). Under `--unattended`, skip the "present and wait" step between phases and proceed straight through with the best matrix the spec supports — but never suppress an uncertain case to make the pipeline flow smoothly. Carry every case flagged uncertain (see Phase 2) into the final output so the caller can report it rather than silently deciding it. ### Phase 1: Spec research Use the `fhirpath-spec` skill for all specification lookups. Gather: - Signature and description - Input/output types and collection behaviour - All spec examples — these become mandatory test cases - Edge cases the spec calls out (empty propagation, boundary conditions, error conditions) - FHIR-specific considerations (choice types, primitive wrappers, extensions) - Ambiguities where the spec is unclear or reference implementations diverge Flag ambiguities for resolution before Phase 2 — interactively, or, under `--unattended`, as an uncertain case carried into the matrix rather than resolved silently. Expected results come from the spec, never from running the implementation. ### Phase 2: Test matrix design Apply input domain partitioning. The driving question: > What inputs can this function receive, and what does the spec say should happen for each? | Dimension | Partitions | When relevant | |---|---|---| | **Core semantics** | Spec examples, basic behaviour | Always | | **Emptiness** | `{}` literal, typed-empty field (`stringEmpty`), computed empty (`where(false)`) | Always, for anything accepting collections | | **Cardinality** | Singular value vs array | Whenever the function reads model fields — see below | | **Element type** | Primitive, complex/backbone, choice type | When the function accepts general `Element` input | | **Nesting** | Flat, nested, deeply nested | When the function involves traversal | | **FHIR encoding** | Real resource via `withResource` | When behaviour depends on genuine FHIR encoding — see below | **Cardinality deserves special attention.** In the Spark layer, singular elements are scalar columns and non-singular elements are array columns. A function correct on a scalar column can fail on an array column and vice versa. Include at least one singular field (`.string("s", "v")`) and one array field whenever the function reads model fields. For a function that expects a **singleton** input — true of most scalar functions, string and math functions among them — the array field must hold exactly one item (`.stringArray("a", "v")`): the point of this dimension is to prove scalar coercion works on an array-backed column, not to test multi-item behaviour. A separate array field with more than one item (`.stringArray("a", "x", "y")`) is a different case: FHIRPath's singleton evaluation rules make multiple items an **error**, not an element-wise map, so it belongs under core semantics as a `testError` case, not under cardinality. Only expect an array field with several items to map or aggregate when the function is documented to operate over the whole collection (an existence or aggregate function such as `count()` or `exists()`). **When to require a real FHIR resource (`withResource`)** — the map-based builder produces a synthetic resource whose type is always `Test`, so it cannot express: - Real resource types and resource-prefixed paths (`Patient.name.given`) - Choice types (`value[x]`) as HAPI actually serialises them - Reference resolution (`resolve()`) and contained resources - Extensions and FHIR primitive-wrapper behaviour (`getValue()`, `hasValue()`) - Anything where the Spark schema from real FHIR JSON differs from the map schema Otherwise prefer `withSubject` — it is faster to read and write. Present a matrix and wait for review. Under `--unattended`, produce the same matrix but proceed straight to Phase 3 without waiting: ```markdown ## Test matrix for `functionName()` | # | Test case | Dimension | Expression | Expected | Subject | |---|-----------|-----------|------------|----------|---------| | 1 | Spec example | Core semantics | `'abc'.fn()` | `'ABC'` | literal | | 2 | Empty literal | Emptiness | `{}.fn()` | `{}` | literal | | 3 | Typed-empty field | Emptiness | `emptyString.fn()` | `{}` | subject | | 4 | Singular field | Cardinality | `singleString.fn()` | `'V'` | subject | | 5 | Array field, one item | Cardinality | `arrayOfOne.fn()` | `'V'` | subject | | 6 | Array field, multiple items | Core semantics | `stringArray.fn()` | error | subject | | 7 | Choice type | Element type | `Observation.value.ofType(string).fn()` | ... | resource | ``` Rules: - Vary one dimension at a time; hold the others constant - Combination tests only where the spec implies dimensions interact - Do not combinatorially explode independent dimensions - Flag any case whose expected result is uncertain ### Phase 3: Code generation ## DSL reference **Location and naming.** Tests live in `fhirpath/src/test/java/au/csiro/pathling/fhirpath/dsl/`, named `DslTest.java` — by capability (`StringFunctionsDslTest`), never by issue number. Extend `FhirPathDslTestBase`. Every file needs the CSIRO Apache-2.0 copyright header; copy it from a sibling test. **One `@FhirPathTest` method per function**, using `group()` to organise dimensions within it — except where the DSL's one-subject-per-method constraint (Gotcha 1 below) forces a split. A function needing both synthetic-subject and real-FHIR-resource coverage cannot fit one method; split by subject, not by dimension, as `ExistenceFunctionsDslTest` does for `count()` (`testCount()` and `testCountOnFhirResource()`). ```java package au.csiro.pathling.fhirpath.dsl; import au.csiro.pathling.test.dsl.FhirPathDslTestBase; import au.csiro.pathling.test.dsl.FhirPathTest; import java.util.stream.Stream; import org.junit.jupiter.api.DynamicTest; public class StringFunctionsDslTest extends FhirPathDslTestBase { @FhirPathTest public Stream testUpper() { return builder() .withSubject( sb -> sb.stringEmpty("emptyString") .string("singleString", "test") .stringArray("arrayOfOne", "test") .stringArray("stringArray", "one", "two")) .group("upper() spec examples") .testEquals("ABCDEFG", "'abcdefg'.upper()", "Lowercase input is uppercased") .group("upper() empty propagation") .testEmpty("{}.upper()", "Empty literal returns empty") .testEmpty("emptyString.upper()", "Typed-empty field returns empty") .group("upper() cardinality") .testEquals("TEST", "singleString.upper()", "Singular field") .testEquals("TEST", "arrayOfOne.upper()", "Array-backed field with one item") .group("upper() core semantics") .testError("stringArray.upper()", "Multiple items in the input is an error") .build(); } } ``` ### Subject methods | Method | Notes | |---|---| | `withSubject(sb -> ...)` | Map-based synthetic resource. Fields are accessed **bare** — `stringArray.first()`, no resource-type prefix | | `withSubject(Map)` | Pre-built model map | | `withResource(IBaseResource)` | Real HAPI resource. Expressions are normally **resource-prefixed** — `Patient.name.given` | ### Assertions — a description is mandatory on every one | Method | Signature | |---|---| | `testEquals` | `(Object expected, String expression, String description)` | | `testTrue` | `(String expression, String description)` | | `testFalse` | `(String expression, String description)` | | `testEmpty` | `(String expression, String description)` | | `testError` | `(String expression, String description)` — any error | | `testError` | `(String errorMessage, String expression, String description)` — specific message | | `group` | `(String groupName)` — prefixes subsequent descriptions as `group - description` | | `test` | `(String description, tc -> tc.expression(...).expectResult(...))` — low-level escape hatch | | `build` | Terminates the chain, returns `Stream` | There are **no** overloads without a description. `testEquals(expected, expression)` does not compile. ### Model builder methods (`FhirPathModelBuilder`) Each type has a value form, an empty form, and an array form: | Type | Value | Empty | Array | |---|---|---|---| | String | `string(n, v)` | `stringEmpty(n)` | `stringArray(n, ...)` | | Integer | `integer(n, v)` | `integerEmpty(n)` | `integerArray(n, ...)` | | Decimal | `decimal(n, v)` | `decimalEmpty(n)` | `decimalArray(n, ...)` | | Boolean | `bool(n, v)` | `boolEmpty(n)` | `boolArray(n, ...)` | | Date | `date(n, v)` | `dateEmpty(n)` | `dateArray(n, ...)` | | DateTime | `dateTime(n, v)` | `dateTimeEmpty(n)` | `dateTimeArray(n, ...)` | | Time | `time(n, v)` | `timeEmpty(n)` | `timeArray(n, ...)` | | Coding | `coding(n, v)` | `codingEmpty(n)` | `codingArray(n, ...)` | | Quantity | `quantity(n, v)` | `quantityEmpty(n)` | `quantityArray(n, ...)` | | Complex | `element(n, b -> ...)` | `elementEmpty(n)` | `elementArray(n, b1, b2, ...)` | `elementEmpty(n)` creates field `n` carrying a null value — the field is present. It is **not** an absent field; no builder method produces one (see gotcha 7). Date, DateTime, Time, Coding and Quantity take **FHIRPath literal strings** — `date("d", "2024-01-15")`, `quantity("q", "10.5 'mg'")`, `coding("c", "http://loinc.org|1234-5")`. Also available: `fhirType(FHIRDefinedType)` to annotate the FHIR type of the enclosing element, `choice(name)` to mark a choice element, and `fhirReference()` for a Reference with empty `reference` and `type` fields. For `type()` assertions, import `au.csiro.pathling.test.dsl.TypeInfoExpectation.toTypeInfo` and compare against `toTypeInfo("System.Integer(System.Any)")` or `toTypeInfo("FHIR.Patient(FHIR.Resource)")`. ## Gotchas These are the ways generated tests actually break: 1. **One subject per method.** `withSubject` and `withResource` set builder state consumed at `build()` — they are **not** scoped to a `group()`. Calling either twice in one method means the last call applies to *every* test case in that method. If two tests need different subjects, they need different `@FhirPathTest` methods. 2. **`withSubject` and `withResource` are mutually exclusive** — each clears the other. 3. **The map-based subject's resource type is always `Test`.** Resource-prefixed expressions like `Patient.name` will not resolve. Use `withResource` for those. 4. **`testError` takes a message string, not an exception class.** `testError(SomeException.class, ...)` does not compile. 5. **There is no `context(...)` argument.** The DSL hardcodes the test case's context to null. If a test genuinely needs a context expression, write it as a YAML case under `fhirpath/src/test/resources/fhirpath-ptl/` instead. 6. **A single-element `List.of(x)` expectation is unwrapped to `x`** before comparison, so both forms are equivalent for one-item results. Use the bare value for readability. 7. **Typed-empty, null-valued and absent are three different things.** `stringEmpty("f")` creates field `f` with a typed null. `elementEmpty("f")` creates field `f` with a plain null. Omitting the field entirely means the path does not resolve, which is a different condition again — and no builder method produces it, so an "absent field" row in a matrix has to be written by leaving the field out. Test the dimension the spec cares about, and say which one you meant. 8. **`group()` persists** until the next `group()` call. ## Running the tests ```bash mvn test -pl fhirpath -Dtest=StringFunctionsDslTest # one class mvn test -pl fhirpath -Dtest='StringFunctionsDslTest#testUpper' # one method mvn spotless:apply -pl fhirpath # format ``` ## Reviewing existing tests 1. Run Phase 1 for the function under test. 2. Build the matrix as if writing from scratch. 3. Compare against the existing tests and report: - **Missing dimensions** — matrix rows with no corresponding test - **Incorrect expectations** — assertions that contradict the spec - **Redundant tests** — several tests covering one dimension without adding value - **Missing FHIR-encoding coverage** — features needing `withResource` that only use `withSubject` ## What not to test - Unicode/emoji handling unless the spec defines it - Large inputs or performance characteristics - Multiple variations of an identical condition - Exhaustive type combinations beyond what the spec defines - Cross-cutting infrastructure (empty propagation at the column level, type encoding) already covered once at the infrastructure level — unless this function behaves unusually