# Checked kernel construction

Kernel construction now receives final GEMM matrix views and final shifted-cast chunks, resolves read storage, and calls the kernel-owned `KernelRun::bind`. That boundary interns the original access metadata and checks the current ABI, scalar arguments, specialization and physical view interpretation. Placement and final emission reuse the same relative addressing; no per-call scalar cache or new executable representation was added.

`low/call.rs` is removed. GEMM row/store interpretation and attention shape interpretation live with their families. The remaining family refactor must consolidate the centralized ABI/specialization rules and extend capabilities to mixed state, local copies and the other proposal requirements. This step is not completion of the proposal.

Two existing tests requested nonexistent implementations: a GeLU that changed AMP element order, and an AMP-left-to-AMP-left rearrangement kernel. They previously stopped before ABI validation. They now construct explicit supported conversions. A new regression checks that column slicing retains backing row strides, rejects fragmented dense-kernel operands before placement, accepts contiguous subregions, and rejects an unsupported GeLU format change.

Validation: `cargo test --workspace --release` passed 426 tests, with 0 failed and 7 ignored; `cargo check --workspace --all-targets` completed without warnings; `git diff --check` passed.

| Hardware fixture | Maximum absolute error | Minimum cosine | Modeled total cycles | Modeled exchange cycles |
| --- | ---: | ---: | ---: | ---: |
| Small FP16 MLP, Repeat 3 | 0.000183 | — | 32,613 | 18,087 |
| Small FP8 ViT, fused QKV, Repeat 3 | 0.189453 | 0.997459656 | 331,874 | 120,708 |

Both device fixtures passed with the same numerical results and modeled costs as the preceding tree. These are randomized small fixtures, not the final full-model accuracy check. The cycle figures are estimates, not measured device runtime. Commands, logs and binary profiles are adjacent.

| Expansion fixture | Variant | Expand ms | Recost ms | Footprint ms | Expanded RSS MiB |
| --- | --- | ---: | ---: | ---: | ---: |
| mlp-b2-repeat3 | before | 788.84 | 321.88 | 504.87 | 205.45 |
| mlp-b2-repeat3 | after | 1254.68 | 453.33 | 771.77 | 205.30 |
| attention-materialized-b1 | before | 459.47 | 92.96 | 330.74 | 120.52 |
| attention-materialized-b1 | after | 457.77 | 98.23 | 342.12 | 120.29 |

All generated-work counts, low cycle estimates, exchange row estimates and geometry counts match for both expansion fixtures. Attention expansion was similar. The MLP comparison was slower (789 to 1,255 ms), including recost and footprint work whose implementation did not change. Earlier controls of the same binary already ranged from approximately 805 to 1,248 ms, and single-core runs reversed that earlier apparent regression (see `../flat-operands/verification.md`). The incremental binding cost is not isolated by these whole-process measurements; no host speedup is claimed. That host timing variability remains an open performance investigation.

The MLP commands pin CPUs 8–15; attention commands pin CPUs 24–31. Each fixture uses the same affinity before and after. The control executable is `/tmp/ipu-family-bind-before` from commit `a740522`. The benchmark/device binary predates the final private-method rename from `append_in_place_cast` to `build_shifted_cast`; the complete workspace checks cover that rename.
