# Boundary layouts and private GEMM grids

The planner currently constructs GEMM, Add and GeLU implementations. Their
boundary layout vocabulary is established before implementation construction:

- Activations offer unreplicated row-major storage with whole-row ownership,
  and dense contiguous intervals aligned to eight bytes. Whole rows suit
  row-wise consumers; contiguous intervals balance pointwise work and memory.
- Both distribute storage up to the configured tile count, avoiding empty
  owners. Whole-row ownership spans batch dimensions as well as matrix rows.
- Explicit layout constraints remain constraints. Imported formats remain
  available. Unconfigured GEMM parameters retain the existing compact packed
  format; compute-grid replicas are preparation storage, not resident defaults.
- Compatible Add/GeLU chains propagate their layouts in both directions before
  the vocabulary is frozen. They cannot import arbitrary internal GEMM grids.

There is an existing copy-path limitation for packed matrices with sub-word row
tails. Their default is currently whole-matrix row-major storage, rather than a
split which lowering cannot realize. This is conservative and can prevent large
irregular matrices from fitting; improving that copy path is separate work.

For each input/output boundary combination, GEMM enumerates its internal grids,
tile orders and reduction ownership. Preparation, GEMM, reduction and the final
boundary conversion are one costed fragment. Boundary combinations and chunks
of internal grids are evaluated in parallel. A grid with even one operand/result
exceeding total tile capacity is rejected before construction.

`estimate/coarse.rs` ranks these MidGraphs using byte-volume copy prices and
the shared kernel estimates. It shares the allocation/liveness traversal but
uses maximum shard sizes instead of per-tile arrays. It does not estimate copy
intersections, exchange fragmentation, pairing or row storage. Its memory peak
is conservative about coincident owners; its cycles are heuristic, not bounds.

Each fixed-boundary chunk retains a bounded memory/cycle frontier (32 entries),
then the merged shortlist is bounded again before detailed analysis. Crowding
is measured in cycle order, preserving the fastest and minimum standard,
interleaved and total-memory entries. This deliberately sacrifices exhaustive
optimality, including contextual memory dominance. Detailed costing and the
resident/transient/retained input analyses run only on the survivors. Actual
search connections still receive detailed costing.

The approximation is the boundary vocabulary. A cheaper direct connection using
some other native layout is not represented unless that layout is offered. Costs
include the actual neutral-format conversions, not an assumed future fusion.
Exchange costs are still estimates before physical placement and scheduling.

On 2026-09-22, the 64-tile FP8 MLP smoke workload's mid construction fell from
35.6 seconds to 2.15 seconds. Hardware validation retained the same 0.002694
maximum absolute error. These are compiler timings, not device cycle counts.
The full-size SigLIP candidate enumeration still spends substantial time costing
copy/exchange geometry; the small-workload speedup does not establish full-size
planning scalability or device performance.

The full 1472-tile FP8 SigLIP MLP did not finish its first GEMM catalogue within
a five-minute bounded run using 56 Rayon threads. This includes interval indexing
of copy-traffic intersections; the cost model is unchanged. There is no new
full-size device timing from that run. This was the motivation for coarse
catalogue ranking.

With coarse ranking, the final 56-thread full-size run completed mid construction
in 63.5 seconds (21.9 seconds for the up catalogue, 37.3 for down). Lowering then
rejected the chosen plan with `view extents are not a valid subset of the shard`;
this is a compiler-throughput result, not a successful full-size device build.
The 64-tile FP8 smoke MLP planned in 185 ms with four Rayon threads and passed on
hardware with the same 0.002694 maximum error. Logs and a CPU sample are under
`artifacts/coarse-cost-20260922/` (the sample is from the intermediate revision,
before the axis-wise alignment check). It shows kernel invocation construction,
symbol formatting and layout geometry as remaining costs.

The storage-view failure was traced to FP8 row-major weight packing: without a
packing kernel, lowering expanded all 20 replicas into 99,164,160 single-byte
copies. Staged packing now reports a missing rearrangement kernel before that
expansion. Structured conversions realizable with ordinary word copies remain
supported; this check does not prune GEMM candidates or supply an FP8 packer.

## Pass integration audit (2026-09-22)

Copy composition runs before coarse shortlisting, detailed candidate pruning,
and connected-fragment costing (after live exports have been identified). Final
assembly composes across fragment boundaries and refreshes whole-graph costs.
The DP still ranks fragment-local costs; it does not credit transformations that
only become possible across the selected fragment boundaries. Explicit tile
mapping is applied to that final graph before costing and expansion. Memory
profiling observes this graph without privately optimizing it.

Other old-planner transformations are still disconnected: cast reordering,
distributed packing, reduction grouping, preparation redistribution, and the
producer-fusion implementation in the parked planner. These select algorithms,
ownership or numerical conversion order; they need candidate integration rather
than unconditional normalization. Add/GeLU candidate fusion remains active.
Low copy elimination, exchange grouping, copy merging, optional cast donation,
and padding elimination remain connected through `low::passes::run`.

The MLP fixture leaves resident parameter layouts unspecified. FP8 input
precision is now independent of layout overrides; native default FP8 parameter
panels use 32-element inner grains. Host activation layout remains explicitly
linear, so this still differs from historical fixtures that uploaded already
replicated/packed activations.

Validation: 19 planner tests passed (one catalogue-report test ignored), plus the
new randomized copy-chain integration test and all eight search tests after the
mapping restoration. The full FP8 batch-4 SigLIP MLP passes on hardware at 721,638
cycles (715,278 renderer cycles), versus 884,784 before composition and the
fixture repair. Maximum absolute error is 0.007080. Weight-packing kernels are
absent from this new profile, but activation packing and reduction staging remain.
A small FP16 MLP also passes (maximum absolute error 0.000488).
Artifacts: `artifacts/fp8-pack-check/mlp-composed.{ipuexe,capnp}`.

## Reduction staging

GEMM candidates offer both native packed and row-major reduction storage, with
gather and row/column scatter ownership. Only the row-major alternative unpacks
partial products on their producer tiles. The reduction kernel accepts a leading
contributor dimension in packed storage: `contiguous_axis_stride` derives each
contributor block from the storage representation, and the supervisor visits
the outer blocks without changing element order. This matters for AMP-left,
whose column panels enclose the contributor dimension. Transposed/block-major
layouts can instead place each entire contributor contiguously. The same
selected block size determines the kernel ABI and cycle estimate.

The device equivalence fixture compares randomized panel-major contributor data
against the old whole-tensor-stack kernel, including in-place outputs and guards.
All 120 reduction cases (206 cases including GeLU) are bitwise identical.

The full batch-4 FP8 MLP passes at 614,700 cycles (608,454 renderer cycles),
versus 721,638 before packed staging: 14.8% fewer cycles. Maximum absolute error
is 0.006470. Both GEMMs now proceed directly to their reduction exchange without
the large producer-side strided unpack. Smaller local copies remain; the largest
strided call is 6,222 cycles rather than 55,482. The expanded alternative set
increases mid construction from 76.6 to 130 seconds in these runs. Package,
rendered profile, and phase report are in `artifacts/packed-reduction/`.

Pairing audit: generation currently requires an entirely pairable transfer
(complete destination pairs, aligned endpoints, even word count). It does not
split off unpaired receivers, a common four-byte alignment prefix, or an odd
trailing word. A prefix only helps when the participating endpoints have the
same alignment residue; incompatible residues cannot be repaired that way.
Selection compares all eligible transfers paired against all ordinary; it does
not search selective subsets. A paired send reserves its neighbouring sender
lane, so selective pairing could avoid contention that the all-paired candidate
introduces. These are search/generation limitations, not additional ISA rules.
No exchange scheduling changes were made in the reduction-staging change.
