# Standalone SDK comparison, 2026-09-22

Both programs execute one batch-one FP16 MLP, 729×1152 → 4304 → 1152,
without biases, on 1472 tiles. Input ownership is compact/linear; weights are
resident in compiler-selected layouts. The timed region includes both GEMMs,
tanh-approximate GeLU, preparation, reductions and exchanges. Host transfers,
compilation and reference computation are excluded. Each executable ran once.

| Compiler | Device cycles |
| --- | ---: |
| Poplar/Poplibs SDK 3.4 | 204,243 |
| ipu-stack `3ab3f84a` cost models | 254,136 |

ipu-stack takes **24.4% more cycles**. SDK uses its default matmul planning
options except FP16 partials and disabled AMP stochastic rounding. There is
no SDK plan tuning or prepacked host activation. The returned output mapping
is unconstrained in both fixtures.

SDK timing uses `poplar::cycleCount` around the whole sequence, with internal
barriers. ipu-stack uses the renderer-normalized span, excluding its leading
profiling interval. ipu-stack has detailed instrumentation; SDK has just the
outer counter. These are cycle comparisons, not clock-frequency-normalized
wall-time measurements or perfectly identical instrumentation.

The SDK fixture checks every output for finiteness and 36 sampled outputs
against a dense FP32/FP64 host calculation using uploaded FP16-rounded random
inputs/weights. Maximum absolute error was 0.00246625. ipu-stack's existing full
reference check passed at maximum absolute error 0.001221. Seeds and weight
distributions are not bit-identical between fixtures.

Artifacts: `artifacts/sdk-mlp-20260922/`, including the SDK executable, build/run
logs and result JSON, plus the fresh ipu-stack package and profile under
`ipu-stack/fp16-b1-n1-default/`.

With the SDK enabled:

```sh
c++ -O2 scripts/sdk-mlp-benchmark.cpp -o /tmp/sdk-mlp-benchmark \
  -lpoplar -lpoplin -lpopops -lpopnn -lpoputil
flock /tmp/ipu-stack-device.lock /tmp/sdk-mlp-benchmark /tmp/sdk-mlp.json
python3 scripts/planner-memory-sweep.py --sdk "$POPLAR_SDK_ENABLED" \
  --output artifacts/new-fp16-control --precision fp16 --batches 1 --budgets 0 \
  --jobs 1 --threads 24
```

No standalone SDK FP8 result was measured in this comparison. The previous
246,186-cycle ipu-stack FP16 result preceded kernel-cost recalibration; the fresh
selected plan is 3.2% slower. This is another plan-ranking regression, rather
than evidence that the unchanged execution kernels became slower.

## Plan and boundary comparison, September 23

Reconstructed the SDK graph with the benchmark's exact shapes and options,
including GeLU, and queried both cached matmul plans and tensor mappings. This
is a graph-construction diagnostic, not a new hardware timing. Sources and raw
results are in `artifacts/sdk-mlp-20260922/{layouts.cpp,plans.txt,mappings.csv}`.

Split counts below use the original logical M/N/K axes, even for swapped
multiplication (`Bᵀ × Aᵀ`). SDK's convolution `fieldSplit` becomes logical N
after swapping; `outChanSplit` becomes logical M.

| Operation | SDK | Selected ipu-stack |
| --- | --- | --- |
| Up projection | swapped; M=12, N=15, K=8; 1440 tiles | ordinary; M=12, N=15, K=8; 1440 tiles |
| Down projection | swapped; M=16, N=6, K=15; 1440 tiles | ordinary; M=9, N=6, K=27; 1458 tiles |

Both SDK grids occur in ipu-stack's `gemm::choices`. Each has 60 variants
across tile orders, reduction layouts and weight-memory classes; all 60 pass
the result word-alignment checks. Enumeration is not the blocker.
The completed catalogue contains 62 up-projection and 124 down-projection
fragments, with **zero swapped GEMMs** in either. Thus the outer DP never sees
the SDK-like orientation in this run.

### Preparation makes the swapped choices unusable

`planner/construction.rs::copy` unconditionally assembles a whole tensor on
one tile when converting a different element order to `Amp::TransposedLeft`,
before distributing it. This was introduced by `22513639` as a workaround for
sub-word transfers. Both MLP weight matrices contain 4,958,208 FP16 elements:
the single-tile temporary is at least 9,916,416 bytes. The current resident
parameter default is ordinary 16×16 block-major storage, so both swapped SDK
geometries hit this path.

The SDK creates distributed resident weights using the selected GEMM plan.
Its up weights occupy 1440 tiles with 3424–3456 elements each; down weights
occupy 1440 tiles with 3264–3456 elements each. Neither matrix is replicated
in this resident mapping. Runtime operand broadcast is a separate matter.

For a controlled comparison using the current compact parameter home and a
linear activation boundary, the cheapest coarse variants were:

| Geometry | Coarse cycles | Estimated peak bytes/tile |
| --- | ---: | ---: |
| SDK-like up | 6,293,684 | 10,023,032 |
| Current up | 137,720 | 188,416 |
| SDK-like down | 6,315,903 | 10,079,400 |
| Current down | 118,006 | 134,144 |

These are fragment estimates, not measured runtime or the complete DP path's
cost. They expose a real oversized temporary constructed in mid, rather than
merely an inaccurate exchange price. Increasing DP width cannot repair it.

### Boundaries also differ

SDK's `Convolution.cpp::remapOutputTensor` explicitly copies the result to a
balanced layout grouped by 16 output channels. It then runs GeLU in place.
The hidden tensor uses 1464 tiles with 2128–2144 elements each. Tile 0 owns
columns 0–15 of rows 0–133; subsequent tiles continue the grouped traversal,
potentially crossing a column-group boundary. The final output uses 1458
tiles with 576 elements each. This is neither our whole-row boundary nor our
logical row-major linear boundary.

The selected ipu-stack hidden boundary is row-major linear ownership across
1472 tiles. Its catalogue does not offer SDK's grouped-channel boundary. SDK
also redistributes between its hidden output and the down GEMM's preferred
input: it is not avoiding all conversions or jointly optimizing both GEMMs.

There is an additional detailed-cost defect: a staged transposed-AMP unpack
to a linear boundary is charged `u64::MAX`. `OperandWindow::local_extents`
represents a linear shard as a one-dimensional vector; the rearrangement
estimate passes that destination geometry to a matrix kernel, whose selection
requires at least two axes. This is not evidence that the real unpack costs
that much. The coarse estimator does not encounter this kernel-selection path.

The immediate obstacle to SDK-like plans is therefore preparation construction
and parameter-layout choice, with a separate output-costing defect and a
narrower boundary vocabulary. This comparison does not establish how much of
the measured 49,893-cycle deficit each issue explains; that requires executable
corrected plans and a hardware comparison.

## After removing the fallbacks

Removed whole-tensor single-tile packing and the single-owner default for
sub-word rows. Unsupported distributed row-tail conversions now return a
planning error instead. GEMM candidates offer their own operand layout as a
resident parameter input, with replication removed and tile numbering compacted;
the previous baseline remains an alternative. First-use parameter selection and
whole-execution resident accounting use the existing search mechanism.

Unpack estimates now use source matrices and source scratch ownership. Kernel
selection failures no longer become `u64::MAX`: mid estimation reports an
unavailable estimate, and low costing propagates the kernel error. Source unpack
eligibility is shared with lowering. The original failure combined destination
geometry for both directions of conversion with the one-dimensional linear-owner
estimate; the matrix kernel correctly rejected that geometry.

Both full-size batch-one standalone MLPs passed hardware reference checks:

| Precision | Previous cycles | New cycles | Maximum absolute error |
| --- | ---: | ---: | ---: |
| FP16 | 254,136 | 275,562 | 0.000977 |
| FP8 | 200,514 | 198,426 | 0.005554 |

These use renderer-normalized cycles. Both new plans select swapped
up-projections, proving that this orientation now reaches execution. The FP16
regression remains a plan-selection problem to investigate, not a claimed
speedup from these changes. No additional ranking rules were introduced to
force a preferred benchmark result.

Artifacts: `artifacts/native-resident-20260923/{fp16,fp8}/`, including packages,
profiles, memory reports and run logs. Validation also included 301 passing unit
tests (five explicit opt-in tests ignored), a passing doctest, and randomized
checks comparing unpack estimates with kernels emitted by lowering. The final
planner-only rerun covers the narrower rejection of unsupported GEMM inputs.

## FP16 selection replay

Reconstructed the previous ordinary GEMM fragments from the current geometry
generator. Neither exact fragment survives in the current catalogue. Injecting
them into the catalogue makes the existing DP select the new up-projection and
old down-projection, without changing search or ranking rules:

| Available candidates | Selected mid estimate | Pre-scheduling low estimate |
| --- | ---: | ---: |
| Current catalogue | 376,366 | 284,320 |
| Current plus previous GEMMs | 372,694 | 260,542 |
| Previous GEMMs only | 376,694 | 251,234 |

These are model estimates, not new hardware measurements. The hybrid has not
been run on hardware. This establishes loss of a useful candidate before DP,
as well as a mid-cost ranking error: the current complete plan wins by only
328 estimated cycles although the previous plan ran 21,426 cycles faster.
Low costing distinguishes them before placement or exchange scheduling.

The replay also exposed a separate low-cost cache defect: equal padded extents
can have different logical tails and different unpack kernel costs. Cache
matching now includes logical extents. The randomized unpack test exercises
these tails and checks cached costs against individual kernel selection. This
cache is downstream of DP and did not cause its selection.

Replay logs and the temporary diagnostic patch are under
`artifacts/fp16-ranking-20260923/`. The diagnostic instrumentation was removed;
no search-policy changes were made in this investigation.

### Exact rejection and cost discrepancy

A second trace matches the complete candidate values and operations against
the reconstructed old fragments, rather than matching only GEMM geometry.
Both are evicted by `planner/search.rs::shortlist` when its frontier exceeds
32 entries. This happens before detailed costing and boundary-compatible
dominance pruning, not because either candidate fails memory screening.

The cap preserves cycle/memory extrema and otherwise deletes the interior
entry whose neighbors have the smallest cycle difference. It ignores distance
in memory dimensions. The old up-projection (coarse price 137,720) sits between
137,139 and another 137,720 entry. That equal-cycle neighbor has 51,480 standard
and 231,424 interleaved peak bytes, versus the old candidate's 58,648 and
148,480. The old down-projection (118,006) sits between 116,414 and 119,547.
Both old candidates are tradeoff points, not dominated points. Adding native
parameter alternatives increases crowding in this capped frontier.

The main detailed-mid error is `estimate/mid.rs::operation_cost` charging
`2 * maximum_destination_bytes / local_copy_bytes_per_cycle` plus
`maximum_intersections * local_copy_call_cycles` for every non-DirectRetile
copy. These intersections include remote arrivals. Lowering receives them
directly into the staging buffer; it does not launch a local copy for each.
Mid then adds the actual packing kernel price on top of that generic charge.

For the old up-projection's activation preparation, this charges 86 local
calls and 35,136 bytes (29,160 cycles), though traffic analysis reports only
one local intersection and 288 local bytes. Packing itself is priced at 7,824
cycles. The old down-projection similarly charges 89 calls and 51,840 bytes
(32,112 cycles), versus one local intersection and 320 local bytes; packing
is 10,800 cycles. These are errors in which work is charged, not just poorly
calibrated bandwidth constants.

| Activation preparation | Detailed mid | Incremental low cost |
| --- | ---: | ---: |
| Old up | 41,976 | 13,150 |
| New up | 39,096 | 36,006 |
| Old down | 49,992 | 18,966 |
| New down | 31,476 | 15,738 |

The low column is the cost difference between successive lowered graph
prefixes, including exchange consolidation and low passes, before scheduling
or placement. It is not a hardware measurement. This category alone
overpenalizes the old plan relative to the new one by 41,024 cycles.

Other discrepancies partly offset that bias: mid prices separate reduction
copies independently while lowering batches their exchanges; generic copy
charges also affect weight preparation and output conversions. GEMM and
reduction compute costs match exactly between these mid and low replays.
Across the whole graph, mid overestimation is 125,460 cycles for the old plan
and 92,046 for the new one, a 33,414-cycle relative error that reverses the
ranking. No exchange scheduler execution is needed to expose it.

Evidence: `exact-trace.txt` records exact-candidate eviction;
`deep-trace.txt` records per-operation prefix costs; `copy-components.txt`
records generic copy charges and actual local intersection counts. All are
under `artifacts/fp16-ranking-20260923/`, with the reproducible temporary
instrumentation in `exact-diagnostic.patch`. Production code remains unchanged.

### Copy-estimator correction

Detailed copy costing now charges local bytes/intersections only, plus the
separate packing or unpacking kernel. LocalKernel does not get an extra generic
copy charge. Generic destination transforms without a specialized packing
kernel retain a separate approximate transform charge. Removed the unused
all-destination/all-intersection counters from traffic analysis.

The old up/down activation preparations now estimate 13,140 / 18,208 cycles,
versus 13,150 / 18,966 from lowering. Their previous mid estimates were
41,976 / 49,992. The complete old plan now estimates 264,986, versus 251,234
from low costing. Exchange consolidation remains outside this mid model.

Randomized tests compare remote FP16/FP8 packing and purely local packing with
lowered kernel costs. Corrected an older test that expected a staging surcharge
for row-major-to-row-major movement, where no transform occurs.

The broader GEMM tests exposed a previously avoided invalid preparation: a
transposed operand required a 162-byte gathered remote payload, but source
gathering cannot send a partial final word. Traffic analysis now records these
intersections; costing rejects them for the destination-staging path that
clips physical padding. This does not reject physical retiles that can transfer
padding, or tile-local halfword copies.

The coarse shortlist is unchanged: 32 survivors per chunk of 4,096 GEMM
assignments, then 32 after merging chunks, before detailed refinement and DP.
The correction does not establish globally correct rankings: the first replay
selects a different catalogue plan at 237,432 mid cycles versus 275,574 low
cycles, even when the old fragments are supplied. These remain estimates; no
new hardware timing is claimed. Logs are `estimator-fix-suite.txt` and
`estimator-fix-replay-final.txt` under the same artifact directory.
