Copy/Compute consolidation verification, 2026-09-14.

Both fixtures use 64 selected compute tiles and randomized weights. They are
small semantic/device checks, not full-size trained-model performance results.
Command JSON files record the exact build and reference-run invocation.

- F16 MLP, [1,17,64] -> 128 -> 64, Repeat 3: numerical PASS, max absolute
  error 0.000183. Modeled final cycles 32613, exchange 18087, both identical to
  the pre-change direct-mid fixture.
- Small FP8 ViT with 3 layers and fused QKV: numerical PASS, cosine
  0.997459656, max absolute error 0.189453, both identical to the pre-change
  direct-mid fixture. Modeled final cycles 331874 (previously 335983), exchange
  120708 (previously 119296). These figures are model estimates, not hardware
  cycle measurements; binary cycle profiles were also captured.
- Codegen release tests: 313 passed, 5 ignored; one doctest passed. The
  coordinate interpreter checks composed and uncomposed copies with a mix of
  automatic, physical and logical traversal policies, ownership and padding.

The retained *-uncomposed-policy files show a rejected intermediate build:
refusing to compose a retile with a logical transformation retained canonical
one-tile staging. The final composition keeps logical traversal explicit while
eliminating that intermediate, restoring the MLP exchange footprint.
