# Bound local-copy calls

Base: 125375bb. Date: 2026-09-15. The overall compiler-structure refactor remains unfinished.

LocalCopy remains a raw byte/stride recipe. Executable low work now retains CopyRun, whose family checks the selected storage ranges and chooses the helper before append. Coalescing explicitly rebinds its changed descriptor. Placement, detailed costing, runtime retention and emission consume the binding; the late helper selector, screening walk and unconditional eight-byte requirement are removed. Arithmetic kernels and copies use the same raw StorageAccess requirement, without fake tensor formats for byte copies. Halfword copies reserve their possible read/modify/write tail.

Validation:

- Release workspace: 444 passed, 0 failed, 8 ignored. SDK-backed saved-search replay: 1 passed. All-target workspace check, formatting and diff checks passed.
- New binding checks cover out-of-bounds contiguous/strided access, invalid backing and cross-tile copies. Existing ABI/worker byte-coverage tests moved from tile/geometry to the copy family. Existing alias, Repeat and randomized complete-graph tests passed.
- 82 distinct hardware copy cases passed full source, output and guard-byte verification. Contiguous/strided 32/64-bit helper commands and fixtures are retained here; each distinct program ran once. Existing halfword hardware results from the previous commit remain applicable.
- Measured contiguous 64-bit cost: 246 + 12 * ceil(words/6). Strided 32/64-bit cost: 294 + 6 * ceil(rows/6) * (2 * words_per_row + 6). Both exactly match the measured cases. Scalar 32-bit: 104 + 23 * words, within three cycles. These fixtures do not establish arbitrary bank-conflict costs or different pointer setup costs.
- MLP Repeat3: package byte-identical, build 0.329 s, peak RSS 50,832 KiB.
- Small FP8 ViT Repeat3: package differs on four tiles, with corresponding weight/profile binding addresses; support sizes and profile metadata are unchanged. It passed a hardware/reference run with minimum cosine 0.996164615 and maximum absolute error 0.275879. Build 0.964 s, peak RSS 208,764 KiB. Detailed estimated cycles changed from 339,623 to 336,152; this is a cost correction, not a measured speedup.
- Saved full 27-layer SigLIP: package byte-identical, build 88.766 s, peak RSS 4,354,204 KiB. No duplicate full-model hardware run. Build timings are observations, not controlled speedup claims.

Implementation count includes the new source file and excludes tests/comments under the existing counter: 50,830 -> 50,966, +136 lines. The baseline remains 48,402. This slice removes duplicate selection but adds explicit binding/range/access checks; overall code-size reduction is still unmet.

Remaining: logical useful-work coverage for local copies (physical padding is not automatically useful), movement policy ownership, geometry-cache consolidation/measurement and final comprehensibility/extension-path and size audit. No new optimizer/retry layer was added.
