# Kernel cost audit, 2026-09-22

The cost functions are issue-time estimates, not bounds. This audit compares
them with current IPU21 binaries, including call overhead, before changing the
estimates. No exchange/search policy or numerical implementation was changed.

## Reproduction

Build `cargo build --release -p ipu-tests --bins`, then run:

```sh
python3 scripts/kernel-cost-audit.py --sdk "$POPLAR_SDK_ENABLED" --output artifacts/kernel-audit
```

The script retains commands, fixture output and failures. `--cases` selects
individual fixtures. Two host jobs compile concurrently; device access uses
the same lock. Fixtures test numerical or bitwise results as well as timing.
Packing now supports production-style constant specialization and FP16;
copy checks also support zero fill. Softmax can compare output memory classes.

Measurements from this audit are in:

- `artifacts/kernel-cost-audit-20260922`: elementwise, residual statistics,
  FP8 output variants, casts (including shifted overlap), packing, unpacking,
  softmax, GeLU and reductions.
- `artifacts/kernel-cost-copy-20260922`: dense and strided copies. Two fixture
  input-generation errors (host-transfer alignment and a buffer exceeding the
  fixture's 64 KiB limit) are retained here; corrected cases passed under
  `artifacts/kernel-cost-copy-extra-20260922`.
- `artifacts/kernel-cost-fill-20260922`: zero fill.
- `artifacts/kernel-cost-mlp-20260922`: full MLP tests after the initial cost
  corrections; final FP16-packing/fill changes are checked separately below.

`comparison.log` in the first directory contains the diagnostic comparison
against elementwise/softmax formulas. Its temporary Rust reporting tests were
removed; retained regression tests compare against independent device timings.

## Corrections

| Family | Problem | Updated model |
| --- | --- | --- |
| FP8 C++ packing | Flat ten cycles per element despite word loads and compile-time unrolling | Worker row groups, word-loop costs, alignment, and zero columns |
| FP8 coefficient assembly | Charged data transposes for wholly zero panels | Separate transpose/zero-panel work |
| FP16 AMP-left packing | Flat element count omitted setup and panel geometry | Row assignment, wide versus narrow loads, tails and zero panels |
| FP16 coefficient assembly | Selected the one-panel price even with extra physical columns | Require both logical and physical width 16; otherwise price all physical panels |
| Transposed FP16 unpack | Ignored guards executing throughout tail-specialized loops | Logical/physical extents and guarded versus zero row pairs |
| FP16-to-FP8 cast | Old four-bundle loop and outdated panel setup | Current two-bundle pipelining, short vectors, panel scheduling and stride eligibility |
| FP8 GeLU/BiasGeLU | Counted source padding as arithmetic; omitted output padding | One shared model for logical panels, quad tails, and zero stores |
| Packed reduction | Charged full call setup for each block | Retained call frame; worker setup/rendezvous per block |
| Small/unaligned Add | Priced scalar gather/wrap as ordinary vector arithmetic | Distinguish quad, pair and halfword paths |
| Zero fill | Omitted setup and the six-worker issue multiplier | 234 cycles plus six cycles per worker wave |

Examples from the independent measurements:

| Kernel/shape | Previously estimated | Measured |
| --- | ---: | ---: |
| FP8 AMP-left pack, 53×192 | 101,771 | 29,280 |
| FP8 AMP-left pack, 73×32 | 23,371 | 2,706 |
| FP16-to-FP8 dense cast, 17,664 values, separated banks | 9,162 | 5,130 |
| Transposed FP16 unpack, 2×31×65, padded to 32×66 | 18,720 | 33,348 |
| Fill 192 u64s | 43 | 426 |

The new FP8 GeLU model matches the aligned measured linear-output cases exactly;
aligned packed outputs differ by at most six cycles. The reduction measurements cover
120 combinations including multiple retained blocks. The old reduction model
overcharged each additional block by approximately 90 cycles.

## Models retained and limits

- Dense/strided copies still match their instruction models closely. Large
  vector Add, FP16 GeLU, BiasGeLU, layernorm, residual statistics and application
  estimates remain close to hardware; they do not need global retuning.
- FP8 layernorm is less exact for packed output (up to about 8% on large cases).
- Whole-row softmax matches the one-active-worker cases closely. Multiple
  active workers can be about 20% slower than the estimate; changing output
  from standard to interleaved memory did not remove that difference. Split
  softmax differs by up to about 17% in this sweep. No unexplained multiplier
  was added. The precise contention source remains unisolated.
- Current MLP profiles provide GEMM comparisons. Interleaved-weight estimates
  are generally close for the large kernels; the standard-weight 73×32×1088
  kernel is estimated at 38,221 versus 34,998 measured. The standard-weight
  model remains approximate. This audit did not build a new isolated GEMM or
  attention-merge timing fixture.
- Cast prices assume favorable bank separation; actual placement can force
  slower loads/stores. C++ packing prices remain approximate, especially for
  byte-aligned tails and compiler-dependent unrolling. Rare FP16 C++ fallback
  and attention-merge estimates were not newly fitted.

These limits matter when two plans have very similar estimates. Recalibration
does not make kernel costs a substitute for placement or device measurement.

## Planner results

The initial corrections alone recovered both plans previously found only
under tighter memory budgets, using the default budget and unchanged search:

| FP8 SigLIP-sized MLP | Before | After | Reduction |
| --- | ---: | ---: | ---: |
| Batch 1 | 222,180 | 200,520 | 9.7% |
| Batch 4 | 608,454 | 571,104 | 6.1% |

These are renderer-normalized device cycles, excluding its leading interval.
Both passed numerical validation. This demonstrates that stale kernel costs
were enough to lose these plans; it does not establish that all catalogue
pruning or exchange-estimation problems are resolved.

The final sweep, including the FP16 packing and fill corrections, is under
`artifacts/kernel-cost-mlp-final-20260922`. All three packages passed hardware
validation; each directory also contains a rendered `profile.html`.

| Batch | Previous cycles | Final cycles | Change |
| --- | ---: | ---: | ---: |
| 1 | 222,180 | 200,514 | −9.7% |
| 2 | 333,738 | 350,886 | +5.1% |
| 4 | 608,454 | 571,104 | −6.1% |

Batch 2 regressed. Its up-projection changed from K288/C336 with 52–53 local
rows to K288/C320 with 56–57 rows; down-projection changed from K160/C384 with
81 rows to K256/C192 with 104–105 rows. The profile's aggregate exchange phase
spans rose from 119,154 to 173,736 cycles, and packing/casting also grew.
These phase spans are not additive critical-path accounting.

The new up-projection packing price is 48,300 versus 48,330 measured; its
down-projection FP16 packing prices are 20,928/23,844 versus 20,970/23,994
measured (unpadded/padded columns). Thus these kernels' corrected estimates
are close even though the selected whole program is worse. Determining where
the previous plan is lost or misranked requires a separate search trace; this
audit does not establish that cause or change search policy to conceal it.

Validation: 300 codegen tests and the doctest passed, alongside the direct
hardware fixtures and these three end-to-end runs.
