# Halfword-copy ABI simplification

Base: 7955237e. Date: 2026-09-15. This is a prerequisite correction within the unfinished local-copy family refactor.

The runtime halfword helper now takes normal destination/source/count/return arguments, as its caller already advertised. The general emitter no longer converts selected calls into an inline address-pair table. The old conversion only covered one-halfword calls with absolute addresses; larger counts and Repeat pointers fell through to the incompatible ordinary ABI. The obsolete table batching, decoding helpers and manual emitter cursor are removed.

Hardware validation uses the production runtime and an extended copy_check fixture. It verifies the entire source and destination buffers, including unmodified bytes:

- 24 batches of 1, 2, 6, 32, 128 and 512 single-halfword calls, across all four source/destination alignments modulo four.
- 40 contiguous cases, 1 to 1,023 halfwords, across those alignments.
- The same 40 cases with three iterations and changing Repeat pointers.
- All cases passed. Each distinct package ran once.

The valid old single-halfword batching path was measured with codegen/runtime from isolated worktree /tmp/ipu-halfword-baseline-7955237e, using the same updated hardware checker. For aligned source/destination, 1/2/6/32/128/512 calls took 246/390/960/4680/18408/73320 cycles before and 240/384/972/4794/18906/75354 after. Across all alignments the largest measured penalty was 3.10%; one isolated copy is slightly faster. Commands, input fixtures, baseline/current binaries and per-case results are retained here. The old invalid larger-count/Repeat paths were not executed on hardware.

Representative model packages remain byte-identical:

- F16 MLP Repeat3: 0.334 s build, 49,728 KiB peak RSS.
- Small FP8 ViT Repeat3: 0.940 s, 207,820 KiB.
- Saved full 27-layer SigLIP: 95.895 s, 4,474,000 KiB.

These exact executable plans have no runtime change. Their existing real-weight numerical validation remains applicable; no identical full-model hardware benchmark was repeated. Build durations are observations, not controlled speedup claims.

Final release workspace run: 443 passed, 0 failed, 8 ignored (/tmp/ipu-halfword-workspace-final.log). All-target workspace check, SDK-backed saved-search replay, formatting and diff checks passed. An earlier doctest run encountered a removed rlib during concurrent rebuilding; the final run was performed after builds finished and passed.

Non-test implementation count fell by 63 lines, from 50,893 to 50,830. The proposal baseline is 48,402, so the overall size-reduction criterion remains unmet.
