# GEMM preparation order and SDK padding

The GEMM solver now offers an unreplicated distributed packed intermediate for
each operand with a supported packing kernel. It compares that with packing
after distribution to the compute grid. These are ordinary mid copies;
construction and the local cost model walk the same preparation chain. Existing
copy composition preserves pack-once staging before replication. No packing-order
override or benchmark-specific grid was added.

This is not necessarily packing on the original input owners: a linear input
must first reach rectangular owners supported by the packing kernel. The extra
exchange and temporary storage are charged. Packing is then performed once per
distributed panel, rather than once per compute replica. FP8 casts remain at
their existing point after preparation.

The randomized exhaustive solver comparison includes both preparation orders.
Distributed numerical interpretation and ordinary lowering tests pass, as does
the full suite: 312 tests and the doctest, five opt-in tests ignored.

## Hardware result

Same batch-one 729×1152×4304 standalone MLP, 1472 execution tiles; normalized
profile intervals as in the previous comparison:

| Precision | Previous cycles | New cycles | Maximum absolute error |
| --- | ---: | ---: | ---: |
| FP16 | 213,732 | 192,264 | 0.001465 |
| FP8 F143, scale -2 | 155,826 | 129,132 | 0.005493 |

Both full numerical reference checks pass. FP16 is 10.0% faster than the previous
plan and 5.9% faster than the recorded 204,243-cycle SDK result. FP8 improves
17.1%; no equal-precision SDK comparison was measured. The SDK comparison uses
an outer cycle counter whereas ipu-stack has detailed per-step instrumentation.

The FP16 activation packing now runs on 1464 tiles, with 822-cycle ordinary
shards and 972-cycle tails, instead of 14,724 cycles on each compute replica.
The selected up GEMM also changes from M/N/K splits 6/30/8 to 3/54/9; the
down GEMM retains 4/12/30. The performance improvement therefore includes a
geometry change, not just replacement of a kernel on an otherwise identical
plan. There are six exchange phases instead of four.

Planning through low reranking took approximately 271 seconds for FP16 and
233 seconds for FP8, with both builds overlapping. The extra preparation choices
increase search time. Rendered profiles:
`profiles/siglip-mlp-f16-b1-prepacking.html` and
`profiles/siglip-mlp-f8-b1-prepacking.html`.

## SDK evidence

Re-ran the existing standalone SDK MLP fixture with execution instrumentation:

```sh
POPLAR_ENGINE_OPTIONS='{"autoReport.all":"true","autoReport.directory":"artifacts/mlp-prepacking-20260924/sdk","debug.instrument":"true"}' \
  flock /tmp/ipu-stack-device.lock artifacts/sdk-mlp-20260922/benchmark \
  artifacts/mlp-prepacking-20260924/sdk-result.json
```

The timed SDK body begins with
`Copy_{ValuePadder/padding,input}_to_up/Conv_1/weightsRearranged`:

| Part | Cycles | Active tiles |
| --- | ---: | ---: |
| Exchange | 571 | 1469 |
| Local copy | 924 | 24 |

This combines input rearrangement and padding; the entire cost cannot be
attributed to padding alone. There is no standalone zero-fill compute set in
the timed body. `poplibs/lib/popops/Padder.hpp::ValuePadder` creates constant
padding tensors, which are included in subsequent copies. They are not newly
zeroed by a kernel each invocation.

The SDK up-projection is swapped. Its channel grain pads the original 729-token
dimension to 736; the original 4304 output channels form the field dimension,
whose split grain is one. Our previous ordinary up-projection instead had 30
144-column panels and a 16-column tail. Its initial clears wrote 4608 bytes on
48 tiles, taking at most 810 cycles per tile. These are different padded axes,
not identical GEMM layouts with a padding pass inexplicably absent in the SDK.

The instrumented SDK run took 205,066 cycles and passed its sampled reference
check (maximum absolute error 0.00246625). The prior uninstrumented reference
was 204,243 cycles. Raw execution data, packages and hardware logs are under
`artifacts/mlp-prepacking-20260924/`.
