# Broader cleanup and cross-review, 2026-09-12

Range: f571a62..5f34a28. Approximate production-line change: Rust -403, schemas -29. Excludes tests.rs, tests directories, the ipu-tests crate, and terminal cfg(test) test modules; counts comments/blank lines consistently. Tests and the legacy binary fixture are additional verification code.

## Representation and boundary changes

- Mid implementation resolution consumes owned bodies instead of cloning Repeat graphs twice. Distributed packing directly retains its chosen operations/values instead of simulating candidate rollback.
- One tiling projection routine handles pointwise and product operands.
- Removed unused mid alignment, access-tail and distinct-memory-element fields. Actual low KernelRequirements retain alignment/overread/bank constraints. MemoryOperand now lives with low calls.
- Repeat storage enumeration is shared across lifetime and padding analysis; ordered/mixed exchange assembly and kernel-format extraction also share implementations.
- Primitive point-to-point/multicast rows share sender emission and receiver-tail construction. Scheduler timing adapters and duplicate payload records are gone. Receivers carry a single vector of complete timing records, replacing three synchronized arrays.
- Production lowering and replay share materialized-phase finalization and release-mode horizon validation.
- Packages and measured profiles share Cap'n Proto profiling records and Rust codecs; profile records now have their own module. Existing generated Rust names remain aliases.
- Prepared and streaming host calls share the rendezvous sequence; input-slice copying is shared too.
- Removed a memory Pareto coordinate that was a monotone function of another coordinate.

## Correctness fixes

1. Repeat binds every local fragment for carried, invariant and iterated values. Previously the tile lookup bound only one fragment. Regression includes rotated linear ownership and verifies placement aliases and resident parameter sequences.
2. Local casts match canonical fragments and validate owners/logical bounds. A normal high-to-mid-to-low FP32/F16 regression demonstrated reuse of row zero for later output rows before the fix. Physical tail differences remain allowed.
3. Exchange captures reject paired sources borrowing an inactive transmitter in an odd-sized active topology. Capture parsing constructs topology once rather than once per paired transfer.
4. Host command/slice bounds reject arithmetic overflow; the old Option comparison accepted None as in-bounds. File-offset overflow is also rejected.
5. Identical linear ownership no longer fails storage-identity comparison or acquires spurious coarse exchange cost. Shape, ownership-offset and padding differences still reject identity.

## Validation

- Full codegen suite before the final linear-identity change: 261 passed, 5 ignored. Final rerun: 262 passed, 5 ignored (artifacts/cleanup-ultra-final-codegen-tests.log).
- ipu-exchange: 52 tests passed. Two pre/post checks covered 211,548 cases, with identical primitive rows and payload timings (including overflow-boundary offsets); artifacts/exchange-cleanup-20260913/.
- Package: 8 tests passed; driver: 3 passed. Package fixture was captured using the pre-change schema; current schema encoding of the same text is byte-identical. See crates/ipu-package/tests/fixtures/.
- Final workspace check passed: artifacts/cleanup-ultra-final-workspace.log.
- Actual BS2 package replay through the updated host driver: initialization plus two inferences PASS, embeddings byte-identical to saved output. artifacts/cleanup-ultra-host-replay.log. This replays the existing device package; it does not claim hardware validation of a newly compiled full model.
- Independent agent cross-reviews covered low kernel requirements after mid cleanup, Repeat/conversion fragment matching, schema compatibility, and driver handshake/fences.

## Limits and retained design

- Old search checkpoints may fail their graph/configuration compatibility check because recorded compiler metadata changed. Keep their pinned binaries for resumption; saved device packages remain readable.
- Current multi-output compute generators use explicit-axis layouts with one shard per tile. Remaining local_shard calls are limited to those paths. General multi-fragment multi-output kernels would need corresponding matching.
- Linear FP8 fusion still has a separate row-width eligibility restriction. Fixing identity costing does not implement that new kernel/layout pathway.
- Distinct exchange algorithms keep their own ready-queue/update logic. Their differences are substantive, so a generic scheduler framework would not be a cleanup.

This review is broader than the earlier random sample; it is not evidence that every remaining part of the repository is optimal.
