--- name: test-design description: Procedure for designing and writing Matft's XCTest cases with high coverage — boundary values, dtypes, memory layouts, NaN/inf, empty arrays, broadcasting, platform differences, performance and agreement with NumPy — with expected values generated by numpy (python/gen_*.py) instead of written by hand. Use this skill whenever tests are written or extended in Matft: implementing a new function or fixing a bug test-first (TDD), adding coverage for an existing function, reviewing whether tests are sufficient, or turning a reported numpy mismatch into a regression test — e.g. "write tests for X", "add test cases", "improve coverage", "is this tested enough?", "check it matches numpy", or in Japanese「テスト書いて」「テストケース追加して」「網羅性を上げて」「境界値のテスト」「型ごとのテスト」「Numpy と一致するか確認して」「TDD で実装して」— even if the word "skill" is never mentioned. For benchmark runs and the docs performance table use the benchmark skill; for image visual checks use image-visual-check. --- # Designing Matft tests Matft promises "behaves like NumPy". Most of the bugs found so far were not in the main path but at the edges: integer wraparound, zero-length dimensions corrupting the heap, slice views reading past their data, complex imaginary parts copied from the real part, vDSP behaving differently on x86_64, NaN dropped by a kernel. A test that only checks `f([1, 2, 3])` in `.Float` catches none of these. This skill turns "write tests" into a systematic pass over the viewpoints where Matft actually breaks, with every expected value coming from numpy so the test cannot encode a misunderstanding. TDD is required in this repository (CLAUDE.md): write the tests, watch them fail, then implement. ## Workflow ### 1. Pin down the numpy specification Before writing anything, run the numpy counterpart in Python and read its docs for the target function: signature and defaults, the output dtype rule, the output shape for each `axis`/`keepdims`, what it does with NaN, inf, empty input, negative axes, out-of-range parameters. Also note where Matft deliberately differs (see "Matft conventions" below) so you don't report them as bugs, and check `Sources/` for the existing signature if the function exists. Use the venv from `python/README.md` (`.venv/bin/python`). If it does not exist, create it as the README says. ### 2. Make a test plan over the viewpoints Read `references/viewpoints.md` and go through every viewpoint for the target. For each one, decide **cover** (list the concrete cases) or **N/A** (with a one-line reason, e.g. "no axis parameter"). Write the plan as a short table in your reply before writing code. The table is what makes the coverage reviewable: a missing row is visible, an unexplained gap is not. | Viewpoint | Cases | Notes | |---|---|---| | Values / boundaries | 0, ±1, negatives, ties, max/min of each int type | | | dtype | Float, Double, Int, UInt8, Bool, Complex | output dtype == np.result_type | | ... | ... | | Size the plan to the change: a new public function gets the full pass; a one-line bug fix gets the failing case plus the neighbouring viewpoints that the same code path touches (e.g. a fix in a reduction kernel → axes, layouts, dtypes). ### 3. Generate the expected values with numpy Expected values are never computed by hand or copied from Matft's own output — that would make the test agree with the bug. Choose one of: - **Generated test file (default for anything with more than a handful of cases).** Add cases to an existing generator (`python/gen_numpy_gaps_coverage.py` for numpy-level functions, `gen_fft_audio_coverage.py`, `gen_image_coverage.py`), or create `python/gen__coverage.py` from `assets/gen_template.py`. Each case writes the Swift expression and the numpy expression side by side, so a reviewer can check them against each other. The output file starts with `// Generated by ... Do not edit by hand.` Register a new generator in the table of `python/README.md`. Details: `references/generator.md`. - **Hand-written XCTest** for things a generator expresses badly: in-place mutation and aliasing, thrown errors, view identity, or a single regression case. Still compute the numbers in Python and paste them with the numpy expression as a comment (`// numpy: np.diff(a, n=4, axis=1).shape -> (3, 0)`). Compare arrays with `XCTAssertClose` (= `np.testing.assert_allclose`, NaN matches NaN, inf must match exactly) and pass `checkType: true` — a wrong output dtype is a numpy mismatch too. For exact Int / Bool results use `rtol: 0, atol: 0`. Avoid `XCTAssertEqual(MfArray, MfArray)`: `==` ignores the mftype, treats floats within 1e-5 as equal, wraps integers into the type (UInt8 index 299 "equals" 43) and never matches NaN, so it hides exactly the bugs this skill is looking for. `XCTAssertEqual` is fine for shapes and Swift scalars. Tolerances: Float `rtol/atol ≈ 1e-5..1e-6`, Double `1e-10..1e-12`; loosen only with a comment explaining why (e.g. accumulated FFT error). ### 4. Red Run only the new tests and confirm they fail for the expected reason (not a compile error in the test itself): ```sh swift test --filter MatftTests. ``` If a case unexpectedly passes before the implementation, it is not testing the change — tighten it. If an existing function fails a new case, you found a bug: keep the test, fix the bug in the same PR (TDD), and mention it in the PR description. ### 5. Green, then the full suite Implement the minimum to pass, then run `swift test` (all). Then check the platform viewpoints that apply (`references/viewpoints.md` §Platform): x86_64 for new vDSP usage, WASI for code with a fallback path. ### 6. Performance (hot paths only) Add a performance case only when the function is a hot path: an elementwise kernel, reduction, sort/search, indexing/setter, conversion, linear algebra, FFT, or anything that scales with a large array. Skip it for small helpers, creation of small arrays, and error handling — and say so in the plan. Procedure (`references/viewpoints.md` §Performance): add a `measureWithWarmup` test in `Tests/PerformanceTests/PefTests.swift` using `PerfFixtures`, and register the same expression in `CASES` of `scripts/benchmark.py` with its numpy counterpart. Include a non-contiguous (transposed) input variant if the kernel has a separate strided path. Measuring and reporting the numbers is the benchmark skill's job. ### 7. Report End with the plan table updated to what was actually covered, the test files touched, the Red → Green results (counts of failures before and passes after), and any bugs found or cases left N/A. ## Matft conventions (not bugs) These differ from numpy on purpose. Write the expected value in Matft's convention and say so in a comment. - A reduction over all axes returns shape `[1]`, not a 0-d scalar (the generator helpers convert 0-d to `[1]`). - Every type except `.Double` is stored as Float32: integers above 2^24 lose precision unless `.Double`; Bool is stored as 0/1 Float. Integer results out of range wrap like numpy's fixed-width ints. - `MfType.result_type` follows numpy for array–array integer promotion; otherwise the higher `priority` wins (integers do not widen when combined with Float). - Invalid arguments usually hit `precondition` (a crash), which XCTest cannot catch. Only test errors of APIs that `throws`; for precondition paths, note the intended behavior in the plan instead of testing it. ## Helpers you should reuse (Tests/MatftTests/TestHelpers.swift) - `XCTAssertClose(actual, expected, rtol:, atol:, checkType:)` — assert_allclose with the worst index in the message. - `layoutVariants(a)` — the same logical array as row/column major, offset view, prefix view, strided view, reversed view and transposed view. Loop over it: `for (name, x) in layoutVariants(A) { ... name }`. - `rowValues(x)` — values in row-major order as `[Double]`.