# Planner memory-budget sweep, 22 September 2026

Compiler: `cadf775a`, after native packed reduction staging. The sweep records
the executable SHA256, exact commands, complete logs, packages, numerical checks,
runtime profiles and estimated/exact placement profiles in
`artifacts/planner-memory-20260922/`.

## Method and limits

One SigLIP-sized MLP: 729 tokens per image, 1152 input/output channels, 4304
hidden channels, no bias, 1472 execution tiles. FP8 inputs and weights use scale
-2; arithmetic results/reductions are FP16. Inputs have compact linear ownership;
the planner chooses parameter storage. All parameters remain resident.

Five planner budgets per batch: default (586,960 bytes), 384, 256, 128 and 64 KiB
per tile. The default 48 KiB support reservation is included in the bound.
Four builds run concurrently, each with 12 Rayon threads; device execution is
serialized. Each build has a ten-minute timeout. Reported cycles use the profile
renderer’s start cutoff, not the leading shared-clock interval. Each package is
executed once with the fixture’s randomized-reference validation.

**This is a planner-budget test, not a physical SRAM reservation test.** The
allocator still sees the full device. Successful placement therefore establishes
that the chosen graph fits alone, not that arbitrary additional resident tensors
will fit alongside it. Exchange-row estimates affect ranking but are not included
in the hard tensor-memory bound; final package placement checks actual rows.

Estimated peaks come from the MidGraph memory report. Placed peaks are unions of
simultaneously live allocation address ranges, maximized over tiles and lifetime
steps. Aliases and sequential reuse are not double-counted. The inaccessible
SRAM tail is excluded. Support includes the runtime/code/exchange/profiling
reservations. Independent tensor/support maxima must not be added together.

## Findings

All nine FP8 cases at 256 KiB or above pass placement and hardware validation.
Maximum absolute reference errors range from 0.005493 to 0.007385. The six
smaller-budget cases reject during search; there are no timeouts or lowering,
placement, or numerical failures in this sweep.

| Planner budget per tile | Batch 1 cycles | Batch 2 cycles | Batch 4 cycles |
|---|---:|---:|---:|
| Default | 222,180 | 333,738 | 608,454 |
| 384 KiB | 222,180 | 333,738 | 571,104 |
| 256 KiB | 200,520 | 333,738 | 820,914 |
| 128 KiB | Rejected | Rejected | Rejected |
| 64 KiB | Rejected | Rejected | Rejected |

| Batch / budget | Estimated tensors + support, bytes | Placed tensors, bytes | Placed all, bytes |
|---|---:|---:|---:|
| 1 / default | 388,352 | 334,208 | 377,316 |
| 1 / 256 KiB | 154,624 | 93,640 | 136,304 |
| 2 / default, 384, 256 KiB | 235,008 | 154,824 | 213,868 |
| 4 / default | 320,512 | 249,192 | 309,600 |
| 4 / 384 KiB | 320,512 | 249,192 | 309,520 |
| 4 / 256 KiB | 262,048 | 203,144 | 263,804 |

The batch-4 256 KiB case's estimate is only 96 bytes below the budget. Its
263,804-byte total placed occupancy includes fixed runtime storage already
excluded when defining the planner's usable-memory budget. An earlier version
of this report incorrectly called that a 1,660-byte budget overrun: the two
totals have different scopes. At the placed peak (tile 60, step 30), tensors
occupy 203,144 bytes and support 60,660 bytes, including 1,760 bytes of fixed
runtime state and code below the planned-data region. This case costs 43.7%
more cycles than the 384 KiB choice.

Tightening the budget exposes faster choices that default planning misses. At
batch 1, 256 KiB selects 200,520 cycles rather than 222,180 (9.7% faster), while
placed tensor peak falls from 334,208 to 93,640 bytes. The up-projection changes
from K=32 panels to K=192 panels; the profile's reduction timeline contribution
falls from 38,766 to 21,366 cycles. At batch 4, 384 KiB selects 571,104 rather
than 608,454 cycles (6.1% faster). These are measured ranking deficiencies;
the follow-up trace below separates candidate availability from ranking.

Rendered runtime profiles are saved for batch-1 default/256 KiB, batch-2 default,
and batch-4 384/256 KiB under each case's `profile.html`. Every successful case
also has an exact interactive placement profile under `memory/placement-*.html`.

All batches reject at 128 and 64 KiB, at high boundary 3 before lowering.
Rejection times are respectively 47.9/73.7/67.6 seconds and 21.5/25.3/22.7
seconds for batches 1/2/4. A search rejection is not proof of infeasibility.

The default build takes 119.5/267.3/411.9 seconds for batches 1/2/4. Of that,
MidGraph construction/search takes 71.6/175/291 seconds, provisional exchange
scheduling 31.3/65.2/60.8 seconds, and final exchange lowering 11.7/21.7/53.0
seconds. Provisional/final tensor placement itself takes under 0.1 second per
call. These are concurrent-build wall timings, not isolated CPU benchmarks.

## Ranking trace

For batch 1, generate the catalogue at both default and 256 KiB budgets, then
search each unchanged catalogue under both budgets. Default generation retains
60/2/78 candidates for up/GELU/down; 256 KiB generation retains 58/2/70. In both
catalogues, default search selects the same larger plan and 256 KiB search the
same smaller plan. **The faster plan survives default shortlisting.**

| Batch-1 selected plan | Mid estimate | Estimated exchange | Measured cycles |
|---|---:|---:|---:|
| Default-budget winner | 378,269 | 126,472 | 222,180 |
| 256 KiB winner | 378,342 | 63,840 | 200,520 |

The model favors the wrong plan by just 73 cycles. Only the up-projection differs:
its estimated total is 229,274 versus 229,347 cycles; GELU and down-projection
costs are identical. The larger memory allowance admits the wrongly favored
plan; tightening the budget removes it.

One concrete error is the row-major-to-AMP FP8 packing fallback in
`kernel/rearrange.rs`: it uses 10 cycles per padded element. For the smaller
plan's 53-by-192 panels that is 101,771 cycles versus 29,280 measured. For the
default plan's 73-by-32 panels it is 23,371 versus 2,706 measured. The estimated
packing penalty for the larger panel exceeds the measured penalty by 51,826
cycles, easily overwhelming the 73-cycle ranking margin. Copy-call and exchange
estimates also differ; this is not a claim that changing this one coefficient
fully calibrates the model.

Batch 4 has a different failure. Cross the default and 384 KiB generation/search
budgets while preserving the true `[4,729,1152]` input shape. Default generation
retains 52/2/124 candidates; 384 KiB generation retains 52/2/112. Both search
budgets select the default plan from the default catalogue, and the faster plan
from the 384 KiB catalogue:

| Batch-4 catalogue | Search budget | Selected mid estimate | Corresponding hardware cycles |
|---|---|---:|---:|
| Default | Default or 384 KiB | 974,606 | 608,454 |
| 384 KiB | Default or 384 KiB | 966,798 | 571,104 |

Thus the better implementation is lost during catalogue pruning, before the DP
search. The detailed estimator prefers it once it is available. Budget-dependent
operand-size filtering occurs before partitioning assignments into groups of
4096. Each group uses the coarse model and a frontier capped at 32 alternatives;
the merged frontiers are shortlisted again before detailed pruning. Changing
the budget changes that early selection, not just the final feasibility test.
This trace localizes the loss to catalogue construction; it does not identify
the individual dominance/crowding deletion which removes the winning geometry.

Diagnostic logs and the temporary tracing-test patch are saved alongside the
sweep as `budget-ranking-b1.log`, `budget-ranking-b4.log`, and
`budget-ranking-diagnostic.patch`. The tracing test was removed after use; no
production planning behavior was changed.

## Historical coverage

The old FP16 batch-1, three-block Repeat workload completed in 979,482 renderer
cycles (`artifacts/repeat-sweep-carried/mlp-b1-n3/`). Its compact host input was
1,679,616 bytes and resident weights 59,609,088 bytes. The same current workload
fails before candidate construction:
`parameter format for this consumer construction is not implemented`.
The FP8 batch-2 three-block case also fails there; a historical FP8 repeated
batch-2 profile exists at `artifacts/strided-copy/mlp-b2-n3/` (856,230 renderer
cycles). Repeat support is absent from the new planner; these are coverage
failures, not memory infeasibility findings.

The evidence supports budget-sensitive alternatives for standalone MLPs, but
does not establish an improvement over the historical planner under full-model
resident-memory pressure. Restoring Repeat coverage and testing actual resident
occupancy are still necessary for that conclusion.

The current standalone FP16 batch-1 control passes at 246,186 renderer cycles,
with 80.8 seconds of planning and 118.2 seconds total wall time (eight threads).
The historical standalone FP16 result in `artifacts/useful-work/mlp/` is 175,578
cycles, but loads 154,736,640 bytes of prepacked/replicated input versus the
current compact 1,679,616 bytes. This is **not** a controlled planner speed
comparison. Historical comparisons under identical standalone input contracts
remain unestablished by this sweep.

## Reproduction

`scripts/planner-memory-sweep.py --output <new-directory> --sdk <sdk-directory>`
runs the FP8 sweep; use `--precision fp16`, `--blocks`, `--batches`, and `--budgets`
for controls. `scripts/planner-memory-report.py <sweep-directory> ...` summarizes
completed cases and computes placed peaks from saved allocation reports.
