# Saved BS1 fusion trace, 2026-09-13

Compiler main 2d4e69d; loaded migrated extended-search checkpoint at 63 attempts, zero additional optimization, no profiling, --inspect-exchanges (no hardware invocation). Debug targets mid::elementwise and mid::residual. See run.log.

- MLP bias add (19): row-major 729 row groups x 2 channel groups / 1458 tiles. GELU (20): flat 4-element grains / 1472 tiles. elementwise::fuse_region only looks up the immediate single-use producer, requires equal input/output tensor types, and does not cross this redistribution. local recipe proposals can open the boundary and re-enumerate consumers, but no explicit bias/GELU ownership proposal exists there; only four alternative plans per operator are offered per incumbent.
- Residual/statistics fusion before LN (17): two statistic parts, redistributed=true, before 8211 / after 8632 estimated cycles. Rejected by cost, not missing kernel.
- FP8 fusion after LN (17) sees the up-projection cast's ownership: 6 channel partitions, 9 row partitions, 27 replicas. Moving ordinary LN to these owners violates the complete-row requirement.
- FP8 fusion after GELU (20) sees down-projection input owners: 27 channel partitions, 6 row partitions, 9 replicas. Moving GELU there repeats work across replicas. Existing pass considers consumer-side producer relocation; it does not instead quantize at original producer owners and rewrite intervening copies to FP8. No exact fused GELU rejection cost was logged, so replication is an explanatory cost pressure, not a measured speed comparison.
- FP8 GELU/LN output kernels already exist; previous suggestion that they need introducing was too broad.

Concrete reduction opportunities: low/expand/reduce.rs creates initial, remote-partials and result scratch and copies the final result to its output intersection. Direct seed reads/final writes require physically compatible slices and bank constraints; do not assume every copy can disappear. Complete/Streamed/Batched staging all exist in lowering, but unconstrained GEMM generation offers only Complete and Streamed. Intermediate Batched candidates could trade a few exchange epochs for smaller receive buffers. This does not shrink already-computed GEMM partials or the separate staged weight panel. Smaller weight staging requires smaller GEMM tiles or explicitly chunking weight delivery/compute, with an associated performance tradeoff.
