# GELU search and fusion rejection trace

Replayed the saved 63-attempt BS1 extended-search checkpoint with one additional step and debug logging at the local recipe screen and elementwise fusion price check. Compiler changes are diagnostics only.

Opening ValueId(360), the MLP bias-add output, causes normal reselection of operation 20 to retain the bias's 1458-tile row/channel layout. The candidate is NOT visited and IS generated. It is rejected by the strict coarse-cycle screen: incumbent 18,688,103; candidate 18,969,740 (+281,637). The prior forced checkpoint builds measured placed-program estimates of 11,763,240 and 11,618,412 respectively. Thus increasing search steps alone cannot recover this candidate under the current coarse comparison.

For the resulting compatible pair, BiasGelu is recognized but estimated at 15,726 cycles against 14,034 for separate Add and Gelu. This is a separate cost decision, not a missed match. Hardware speed of the fused variant was not measured.

Reduction-copy change: use view_byte_traversal(..., Physical).contiguous_span() on the destination intersection, validate byte length/alignment, and bind the last reduction run directly to that ShardView. Allocate ping-pong scratch only for intermediate stages, retaining existing staging for fragmented destinations. A local seed can similarly be read through its original contiguous view, but cannot become a writable intermediate accumulator without a separate last-use/donation proof. Current reduction kernel uses ordinary loads/stores, no distinct-element constraint; standard contiguous input/output ABI suffices. Test complete/streamed/batched reduction, both stage parities, partial destination offsets, fragmented fallback, multiple contributors/output groups, and Repeat liveness before hardware verification.
