# Reading reduction seeds in their producer layout

The FP16 reduction accepts a byte stride between seed panels. Its worker loop
and the compact layout of the other contributors are unchanged. Address binding
derives the stride from the selected view in its backing allocation; dense seed
buffers use the original block-size stride. The supervisor replaces one pointer
increment, so this does not add instructions to the reduction cycle estimate.

A low reduction pass removes seed staging when its copies exactly describe that
regular view. It checks storage lifetimes, other readers, aliases, and the actual
source/destination byte mapping. Mutable loop bindings and externally visible
staging remain independent. This pass runs before low costing and placement.

Randomized tests cover panel dimensions, contributor counts, contiguous and
strided copy descriptors, exposed seeds, and writes to the source after staging.
The small hardware MLP passes with maximum absolute error 0.000732. The full
729x1152x4304 FP16 MLP passes with maximum absolute error 0.001465.

The initial strided-seed-only build removed 810 copy descriptors (five panels on 162 tiles), but its
normalized runtime is 192,270 cycles versus 192,264 before. The remaining 1,296
tiles still gather their local contributor into the other-partials buffer.
Contributor zero is remote on those tiles. Earlier discussion incorrectly
described all these copies as seed staging: strided seed access alone does not
remove the copies of local non-seed contributors.

## Local contributor selection

The follow-up low pass in `low/passes/reduction.rs` also selects a local
contributor previously staged in the second operand. It redirects the old seed's
receives into that contributor's vacated slot, binds the local contributor as
the new seed, and removes its copies. Mid construction and copy mappings are
unchanged. This is a permutation of an existing reduction's contributors, before
low costing, exchange scheduling and placement. It may change FP16 addition order.

The pass requires private reduction operands, complete seed receive coverage,
and exact byte correspondence between the removed copies and the selected local
view. It rejects conflicting writes, overlapping receives and externally visible
or mutable loop-boundary staging. It handles both contiguous and regularly
strided local panels. The old seed-borrowing pass was moved out of general
movement handling into this pass, sharing its validation.

The full FP16 MLP now runs in **190,224 normalized cycles**, compared with
192,264 before either change (1.06% faster). Maximum absolute error remains
0.001465. All up-projection reduction copies around 78k are gone; the four
remaining down-projection `copy_u64` samples and unrelated GeLU-input copies
remain. The pass removes 7,605 copy descriptors in the selected full plan.

Randomized contributor-token simulation covers row-major and AMP layouts,
rotated ownership, sharded and replicated results, multicast destination
redirection, and exposed intermediates. Each output must receive every original
contributor exactly once. All 314 unit tests pass, with five opt-in tests ignored.
The small FP8 MLP also passes on hardware (maximum absolute error 0.005859).

Updated profile: `profiles/siglip-mlp-f16-b1-local-seeds.html`.
Captures and packages: `artifacts/reduction-selection-20260924/`.

Initial profile: `profiles/siglip-mlp-f16-b1-strided-seed.html`.
Initial packages, reference results and captures: `artifacts/reduction-seeds-20260924/`.
The attempted repeated fixture is rejected before lowering because the current
planner does not implement its parameter-format construction.
