# Simultaneous copy loads and stores

`copy_u64` and `copy_strided_u64` now prime one 64-bit value, repeat
`ldst64pace` for the remaining words, and drain the last store with `st64pace`.
They never prefetch beyond the selected row. Single-word rows execute only the
prime and drain. The 32-bit strided helper retains separate accesses because
its pointers/row widths need not support 64-bit operations.

The copy binding declares a source/destination memory-element conflict for its
64-bit implementation. Placement collects it alongside ordinary kernel operand
conflicts, including copies nested in Repeat, and resolves it through the normal
alias groups. Identity reductions also declare the conflict because they use
the same copy kernel. No retry or alternative allocator was added. Randomized
copy-chain tests check the actual placed source/destination element sets as well
as interpreting the resulting byte movements.

This is required by the ISA: simultaneous accesses must use distinct memory
elements (Tile Vertex ISA IPU21, sections 2.9.12 and 3.7.5.1.1).

## Measurements

Hardware fixtures check complete destination buffers, including gaps and guards.
The updated copy cost formulas match the fixture timings exactly:

- Contiguous 64-bit: `264 + ceil(words/6) * 6`.
- Strided 64-bit: `288 + ceil(rows/6) * 6 * (words_per_row + 10)`.
- Small supervisor 32-bit: round `102 + 21 * words` up to six cycles.

The first two paths approach one memory issue per copied 64-bit word, rather
than two. For 3584 bytes, contiguous copying falls from 1146 to 714 cycles;
three strided rows of that width fall from 5706 to 3036 cycles. Short rows pay
more setup: a one-word strided row takes 354 rather than 342 cycles.

The supervisor 32-bit helper also uses post-increment loads/stores instead of
separate pointer additions. A 256-byte copy falls from 1578 to 1446 cycles.
Unlike the combined instruction, those MRF forms are supported in supervisor
mode and impose no additional memory-element separation.

Artifacts: `artifacts/copy-ldst-20260926/`. `kernel-audit` contains the new
measurements; `baseline-copy-*` reruns the same fixtures with the previous
runtime assembly. All four copy fixtures pass (172 cases total).

Full batch-one SigLIP-sized MLPs both fit and pass hardware/reference validation:

| Precision | Previous profile | New profile | Change |
| --- | ---: | ---: | ---: |
| FP8, scale -2 | 129384 | 128670 | -714 cycles (0.55%) |
| FP16 | 189084 | 188754 | -330 cycles (0.17%) |

Maximum absolute errors remain 0.005432 and 0.001465 respectively. The FP16
comparison uses the last recorded full profile, before the intervening shared-row
fill upgrade's 30-cycle setup increase. These are end-to-end comparisons, including
placement/scheduling effects. Profiles are `profiles/siglip-mlp-f8-b1-ldst.html`
and `profiles/siglip-mlp-f16-b1-ldst.html`. All 314 codegen tests pass (five ignored).

## Other paths inspected

The FP16 AMP-left packer's tail copies now use post-increment loads/stores too.
Fully padded panels use four 64-bit zero stores instead of eight 32-bit stores
and eight pointer additions. Its cost model was updated. The 12x50 packing
fixture falls from 1350 to 1242 cycles; full-panel cases are unchanged. The
extended 80-case hardware packing fixture includes completely padded rows and
panels and passes bitwise.

GEMM and the main FP8 cast paths already use simultaneous load/store operations.
The FP16 coefficient packer's 32-bit output stores address separated positions,
so adjacent register numbers do not make them mergeable into wider stores.
GeLU's vector loop still has separate input/output accesses; pipelining those
across iterations is a further candidate, but needs an arithmetic-pipeline
rewrite and a decision about retaining in-place execution. It is not changed
here. Sparse exchange patching also still has explicit pointer increments;
its scattered 32-bit writes cannot directly use `ldst64pace`.
