# Reduction output copy bypass

Commit e3fa887. Current reduction ABI writes contiguous physical destination slices directly in the last stage; fragmented/alias destinations retain the deferred packed-result copy. Single-stage direct outputs also avoid the result scratch allocation. Seed copies are unchanged.

Tests cover offset destination guards, differing contributor counts, complete/streamed/batched staging, odd/even stage counts and fragmented fallback. Compiler library run: 281 passed, one old assertion requiring final copies failed, five ignored. Updated that assertion to accept direct reduction writes; the randomized test then passed. No functional failures remain from that suite.

Full 27-layer BS1 rebuilt from the original saved extended-search recipe (no additional search). Two resident inferences pass against FP32 with cosine 0.994264800 both times. Saved output.bin is byte-identical to artifacts/search-state-20260912/bs1/extended/resident/output.bin.

Placed cost: 11,697,188 cycles, down from 11,763,240 for the same recipe before bypass (-66,052, 0.56%). Resident replay: 7.792905 ms/image, 128.322 images/s; hostExchangeReplay PASS. Host timing includes image upload and embedding download and is subject to host variation.

See run.log, busy.log, model.ipuexe, resident/ and memory/placement-537633.html.
