# Local-copy binding — in progress

Base 7955237e, previous turn was progress and completed shared low storage binding. Full goal still unfinished.

Trace: tile::local_copy_call selects helper/ABI; place forces alignment 8; estimate/program uses the pre-optimization six-instruction strided loop; package::runtime_retained_symbols redoes helper selection; profile_work excludes local copies. Plan is a bound CopyRun retaining its byte descriptor plus CopyKernel family choice, used by all these consumers, with explicit rebinding during copy coalescing. Do not fabricate tensor formats for byte copies or keep optional bound metadata / an extra IR.

Found correctness issue before implementing family binding: static_copy_u16 takes an inline table at the return address, but local_copy_call describes a normal destination/source/count callable. lib.rs converts only absolute one-halfword calls into that table. Multi-halfword and Repeat-relative calls fall through to an incompatible ABI. The existing randomized tile test only checks argument values, so misses this.

Completed ABI correction replaces static_copy_u16 with the ordinary contiguous halfword ABI and remove the inline-table special path in lib.rs. This deletes a special backend protocol and supports all counts/Repeat pointers. Hardware, package comparisons and workspace validation are complete; see verification.md. All three model packages are byte-identical. The largest measured microbenchmark penalty is 3.10%, while longer copies and Repeat pointers now work. Overall implementation size is 50,830 lines, still above the proposal baseline. No process or device job remains active.

Main family work remains to implement. Cost model should describe current two-instruction strided loops, actual helper width and six-worker row distribution; supervisor 16/32-bit copies need their own work model. Preserve separate high-level approximations where exact helper is not yet bound. Alignment/tails must account for u16 read/modify/write of aligned 32-bit words.


Profiling detail for the next family work: package/profile_work.rs currently returns None for every local copy because byte descriptors may include padding. Do not simply count every physical byte as useful. Recover or retain the logical byte coverage from the selected source/destination geometry; shared geometry-cache work may help. Distinguish payload issue work from supervisor/worker setup in the cost and usefulness models.

The new current copy_check accepts --word-bytes 2/4/8, --contiguous, --repeats and --runtime-source; JSON cases optionally include repeat_stride. Default strided-u64 mode remains compatible with earlier affine packing fixtures. Baseline worktree contains only the updated hardware tester over old production code. Current main CLI was rebuilt after the baseline build; do not run an old compiler binary against the new halfword runtime.
