# Critical-path MLP padding

The batch-one FP16 729×1152×4304 MLP still initializes exchange destinations
after GeLU. On the last-arriving tile, an activation clear takes 966 cycles,
followed by six 512-byte weight clears taking 300 cycles each. The weight
regions have a 4608-byte pitch. These are physical K-tail regions, not useful
matrix elements.

Copy lowering now combines equal-width, regularly spaced clears when one
strided fill is cheaper than separate contiguous fills. It uses the existing
storage `StridedSpan` representation. Six workers traverse independent rows;
large individual ranges retain the contiguous kernel's within-row parallelism.
The kernel family owns the cost used for selection and low costing.

The complete hardware MLP takes 189,084 normalized cycles, down from 190,224
with local reduction seeds (0.60%). The six weight clears become one 660-cycle
call; post-GeLU critical-path padding falls from 2766 to 1626 cycles (41.2%).
The maximum absolute error remains 0.001465. The rendered profile is
`profiles/siglip-mlp-f16-b1-padding.html`.

The hardware fixture checks 35 combinations of row count, width and stride,
including preservation of every gap and guard byte. All pass bitwise. Its
timings match `246 + ceil(rows/6) * 6 * (words_per_row + 5)` exactly for
1–73 rows and 1–192 words per row. Reproduce with `scripts/kernel-cost-audit.py
--cases fill-strided` and the script's required SDK/output arguments.

Packing already initializes its complete output, including padding. These
remaining destinations instead receive logical values through exchanges.
Eliminating their activation padding using zero weight tails would require
proving that the explicit destination clears survive subsequent writes. The
current padding pass proves that property for resident parameters and complete
physical copies, but not partial/rearranged writes. This change keeps that
conservative behavior. It also avoids introducing a finite-scratch requirement,
which would restrict allocation reuse.

Artifacts: `artifacts/padding-20260924/`, including the hardware fixture's
`kernel-audit/fill-strided/cycles.json`.

Validation also includes all 314 codegen unit tests (five ignored), randomized
zero-range binding checks, and the existing randomized byte-level interpreter
checking logical values and physical padding through copy chains.

## FP8

The same batch-one workload with `--fp8-scale=-2` passes hardware/reference
checking (maximum absolute error 0.005432), taking 129,702 normalized cycles.
Profile: `profiles/siglip-mlp-f8-b1-padding.html`; artifacts: `f8.*` in the same
directory.

No strided fills are selected. After GeLU the last tile still executes a
27,648-byte clear (3690 cycles) and three 3584-byte clears (684 cycles each):
5742 cycles total. Three large rows would leave half the new kernel's workers
idle, costing 2964 cycles versus 2052 for the three contiguous calls. Improving
that case needs within-row parallelism while amortizing the launches. It does
not address the larger contiguous activation clear. The initial critical-path
weight clear also remains 810 cycles.

The older prepacking FP8 profile takes 129,132 cycles; that comparison includes
the intervening local reduction seed changes. Padding has the same 368 calls
and 612,336 summed tile cycles. The exchange before GeLU takes 3462 cycles after
the last arrival, versus 2322 previously; reduction changes recover part of this
increase. Thus the FP16 padding speedup does not carry over to this FP8 plan,
and the combined changes currently regress total FP8 runtime by 570 cycles
(0.44%).

## Multiple workers per row

The strided fill now has two worker arrangements under the same kernel entry:
workers take independent rows, or all six workers split each row into contiguous
pieces. Both use `rpt` for stores and a branch between rows; the ISA cannot nest
repeat loops. The kernel family chooses the cheaper arrangement, and lowering
compares that cost with separate contiguous launches.

Both arrangements pass 77 hardware cases each, checking the entire destination
including untouched gaps/guards. Measured cycle models are exact over these
cases (1–73 rows, 1–448 words per row, limited to 64 KiB fixture buffers):

- Independent rows: `276 + ceil(rows/6) * 6 * (words + 5)`.
- Shared rows: `300 + rows * 6 * (ceil(words/6) + 5)`.

The three FP8 weight clears take 1740 cycles with shared rows instead of 2052
with separate launches. The extra selection/setup adds 30 cycles to independent
rows versus the original strided helper (the FP16 six-row example is now 690
cycles). Single contiguous clears still use the unchanged kernel.

Fixture artifacts: `artifacts/padding-shared-20260924/kernel-audit-v2/`.
Reproduce with `scripts/kernel-cost-audit.py --cases fill-strided fill-shared`
and the required SDK/output arguments.

The full FP8 MLP passes with unchanged maximum absolute error 0.005432. It takes
129,384 cycles versus 129,702 before this variant (318 cycles, 0.25%). Post-GeLU
critical-path clearing falls from 5742 to 5430 cycles. Sixty tiles use the new
shared-row call. The large activation clear remains unchanged.
Profile: `profiles/siglip-mlp-f8-b1-shared-padding.html`; full-run artifacts:
`artifacts/padding-shared-20260924/f8.*`. All 314 codegen unit tests pass, with
five ignored.
