- Reintegrate exchange placement.
- Reintegrate optimistic planner.
- fillzero seems wrong level?
- ensure low is not doing other planning ("lol")
- Pull copy strategy out into planning/optimization (was this already done?).
- streamed reduction grouping (what?)
- change consumers to accept views etc
- fix MLP regression
- fix that exchange stress case

  3. Reduction seed selection was lost. Reduction construction (crates/ipu-codegen/src/planner/reduction.rs:103) always seeds from contributor zero. Before
     66a9ad4b, it preferred a contributor already on the destination tile. Copy elimination cannot recover that choice once planning selects the remote
     contributor. This remains a plausible contributor to the previously reported MLP slowdown; I haven’t measured its individual impact.

non-p2 splits/partitioning
swap multiplication
broader reduction ownership ?
axis orderings
cast before replication
"uneven grain-balanced shards"
tile placement ?
SDK-style "panel-local subdivision"
cast before replication
Is it doing cursed prefiltering again?

gelu padding striding?

  - Recognize explicitly zeroed weight tails when eliminating activation clears. The current pass mainly tracks zero-padding through parameter provenance. It
    doesn’t establish the useful fact here: “these weight-tail regions were explicitly cleared and remain zero when GEMM reads them.” With that proof, the
    existing finite-scratch contract could remove the corresponding activation K-tail clear—approximately 966 cycles on that tile. This needs region/lifetime
    tracking, not simply treating every parameter-derived buffer as zero-padded.

  - Deliver weight padding with the weights. Preserve padded physical panels through redistribution, so the destination receives zeros alongside useful
    weights. That could remove the weight clears themselves, at the cost of extra transfer bytes. It needs costing; shipping padding is not automatically
    cheaper.

graph calibration generator

"read" and fix planner
