# Direct kernel input bindings

Kernel inputs are now direct `ShardView`s, matching indexed output bindings. The removed `KernelOperand` wrapper held a list, but every working constructor supplied one view and ABI validation rejected all other counts. A view can still have a strided/packed physical traversal; physical-view validation is retained.

Production source shrank by 104 lines (excluding test files and trailing test modules). Input/output requirements and placement constraints now walk the same direct binding structure.

Validation: `cargo test --workspace --release` passed 425 tests, with 7 ignored; the final `cargo check --workspace --all-targets` was clean. `git diff --check` passed. The two hardware fixtures retained their preceding numerical results and modeled costs:

| Fixture | Maximum absolute error | Cosine | Modeled cycles | Modeled exchange cycles |
| --- | ---: | ---: | ---: | ---: |
| MLP FP16, batch 1, 17 tokens, width 64, hidden 128, Repeat 3 | 0.000183 | — | 32,613 | 18,087 |
| Small ViT FP8, fused QKV, Repeat 3 | 0.189453 | 0.997459656 | 331,874 | 120,708 |

These are small randomized hardware fixtures, not a full trained-model accuracy check. Cycle fields above are modeled values, not measured device runtime. Adjacent command files preserve exact invocation details.

## Host expansion measurements

| Fixture / control | Expand ms | Recost ms | Footprint ms |
| --- | ---: | ---: | ---: |
| mlp-b2-repeat3-before | 1247.54 | 455.17 | 513.25 |
| mlp-b2-repeat3-control-before | 804.86 | 311.89 | 341.08 |
| mlp-b2-repeat3-after | 1218.79 | 449.46 | 746.58 |
| mlp-b2-repeat3-recheck-before | 816.00 | 335.34 | 339.51 |
| mlp-b2-repeat3-recheck-after | 1231.88 | 440.25 | 766.29 |
| mlp-b2-repeat3-perf-before | 1244.35 | 451.49 | 514.26 |
| mlp-b2-repeat3-perf-after | 1224.38 | 461.48 | 778.46 |
| mlp-b2-repeat3-cpu8-before | 1235.18 | 455.44 | 548.36 |
| mlp-b2-repeat3-cpu8-after | 725.62 | 261.89 | 449.01 |
| attention-materialized-b1-control-before | 462.75 | 102.36 | 197.79 |
| attention-materialized-b1-after | 455.14 | 93.21 | 339.81 |

The control binary predates both the neutral tensor-geometry move and this change. CPU 8–15 runs use one NUMA node; CPU 8 runs also set `RAYON_NUM_THREADS=1`. The same old binary varied from roughly 805 to 1,248 ms for MLP expansion; the new binary varied from roughly 726 to 1,232 ms. Single-core comparison reversed the earlier apparent slowdown. These measurements do not establish a speedup or a consistent regression. The source of the bimodal host variation remains unisolated. No device benchmark was repeated for this investigation.

MLP low cycles, shard/kernel/copy counts, transfer and recipient counts, row footprint and geometry-cache counts were identical across variants, as were those metrics for attention. RSS after expansion was approximately 205–206 MiB for MLP and 120–122 MiB for attention in the eight-CPU runs.

CPU sampling is retained in the adjacent `perf-*.perf.data` and hotspot files. Much of the sampled analysis time is geometry hashing and exchange-footprint hashing; changing those hashers is separate work.
