# Cause of the GELU cost inversion

Instrumentation commit bdc6c63; benchmark recipes include reduction-output bypass. Debug logs and analyze.py / breakdown.json record mid operation prices and low exchange geometry. Scheduled comparisons use reduction-bypass-20260913/run.log for BS1 and gelu-ownership-20260913/open.log (before bypass) for the preserved layout. The compared exchange IDs map to the same operation provenance; reduction-output bypass does not change their exchange geometry.

Mid, per repeated layer:
- GELU op20: remove a 4141-cycle conversion; kernel changes 11370 -> 11382. Saving 4129.
- Down-projection op21 preparation: 30248 -> 44808, entirely exchange (25080 -> 39640), penalty 14560.
- Net +10431 per layer x27 = +281637; this is the entire coarse rejection margin.

Low down-projection input exchange:
| layout | max endpoint payload | max fragments | predicted cycles | scheduled event horizon |
|---|---:|---:|---:|---:|
| flat GELU | 39040 B | 134 | 21440 | 10169 |
| preserves bias owners | 39040 B | 244 | 39040 | 10165 |

Low uses max(payload/4 + 600, fragments*160). Both variants are dominated by the fragment term despite their actual schedules having effectively equal durations. Preserved ownership also removes the preceding GELU redistribution: estimated 2752 cycles, scheduled event horizon 3578. Scheduled horizons exclude the estimator's 600-cycle global-phase allowance.

An unchanged bias-preparation exchange illustrates the absolute overestimate: 10496 B, 406 fragments -> 64960 estimated cycles, but 2731 scheduled event cycles (3331 with the modeled barrier).

The 160-cycle coefficient dates to c130f98 (2026-08-16, attention conversion/exchange fusion). It is an empirical fragmented-conversion price, not an ISA restriction or universal per-fragment cost. Logical spans do not imply independent serialized endpoint/address setup. Applying this coefficient to all geometry, then hard-rejecting nonimprovements, is the main demonstrated issue. Low costing retains the same heuristic, so more exact span expansion alone cannot correct it.

Secondary estimator approximations: mid sums isolated operation maxima (low composes tile-local timelines); low endpoint costing combines outgoing tile pairs via source.tile/2 even though ordinary transmit lanes are independent and only paired mode borrows its partner. These are separate from the demonstrated fragment penalty and were not changed here.
