# Baseline-first local planning

Work branch: `baseline-local-planner`. Replaces the whole-graph beam and its
configuration/penalty retries; it must not become another production fallback.

The initial program uses canonical activation boundaries, compact persistent
parameter homes and explicit conversions. Operator implementations remain
whole-device mid selections generated by the existing catalogue. The baseline
must pass complete package construction before optimization begins.

Local proposals change operator selections or remove canonical boundaries in
small neighborhoods. Unaffected parameter ownership remains fixed. Every
accepted proposal has a complete jointly allocated, scheduled and linked
package. Physical persistent addresses may move; there is no partition of SRAM
into immutable persistent and scratch banks. Existing schedule replay may reuse
compatible phases, but successful local graph rewriting does not imply local
physical placement.

Validation must cover canonical boundaries, explicit Repeat argument/yield
formats, compact parameter sequences, rejection retaining the incumbent and
bounded optimization effort. Retire the old future-state beam bookkeeping,
configuration products, stronger-penalty retries and their obsolete knobs/tests
when switching the production entry point. Keep candidate-generation and kernel
correctness tests, adapting selection-specific assertions to the new contract.

## Implementation

The production entry point now builds one baseline and compares at most
`optimization_steps` ordered proposals (default eight). Parallel speculation can
also build later candidates before the first improvement is selected. It retains the actual
built incumbent, including its final allocation and schedule. Rejection never
starts a fresh global search. A failed baseline is reported as such.

Search runs start from the baseline; progress is not saved or resumed.

The default baseline uses contiguous row blocks targeting one 16 KiB SRAM
element per owner. Parameter sequences use compact native homes and a shared
rotation. Casting initially follows redistribution; early casting is a local
alternative. Operator selection uses implementation estimates and the conversion
inserter to price surrounding input movement.

`--capacity-baseline` (or `PipelineConfig::capacity_baseline`) instead distributes
whole rows over as many row owners as the shape and device permit, without
replication or tail-row padding. It compares both cast orders using the ordinary
conversion inserter and mid implementation/liveness analysis. Its memory score
includes source buffers, input conversions, compute scratch and conversion back
to the canonical output, retaining inputs that have later consumers. Automatic
parameters are costed in their compact homes, as in actual lowering.

Capacity selection prioritizes total live tensor memory, then maximum standard
allocation, exchange-row estimate and cycles. Its operator shortlist retains a
smallest-buffer candidate alongside the existing exchange-storage extreme. It
uses the same lowering and local optimizer, with no automatic retry between
baseline policies. It remains opt-in: PE batch two fits, but SigLIP encounters
exchange storage/alignment limitations. See the
[capacity-baseline results](CAPACITY_BASELINE_2026_09_12.md).

This is still a local memory estimate: it does not model every other live graph
value or prove joint placement. Compute operands may retain replicas when the
existing kernels need them; bounded panel broadcasting has not been added.

Proposals include operator alternatives, joint producer/store and consumer
changes, opening canonical boundaries, cast order, packing distribution and
independent reduction overlap. Physical tile mapping is also compared within
the validation budget. Unchanged operator recipes skip catalogue generation;
implementation, expansion and exchange-schedule caches are shared across trials.

Removed the whole-graph beam, future-state equivalence machinery, region-beam
cache/attachment adapter, configuration products, finalist admission layers,
penalty retries, and their CLI/configuration knobs. The remaining `mid/lowering`
module applies selections and inserts conversions. `--optimization-steps 0`
requests just the baseline; `--operator-candidate-limit` controls local catalogue
breadth. Expansion benchmarks and exchange capture inspect the baseline.

## Validation so far

- Codegen: 218 unit tests and one doctest pass; four manual tests ignored.
  Seven benchmark/diagnostic tests, workspace all-target checks and Clippy pass.
- Compact iterated and invariant Repeat bindings expand and place correctly.
- Rejected/slower local proposals retain the incumbent; validation budgets and
  immediate baseline failure are covered.
- Hardware exposed an existing Repeat bug: local copies used absolute weight
  addresses. They now use the iterated pointer, like kernels and exchanges.
- Small two-layer FP8 ViT with three accepted proposals: 8.27 seconds package
  planning, final modelled cost 300276 cycles, hardware maximum absolute error
  0.105469. See `artifacts/baseline-local-planner/local-small-r2/`.
- Current full-size one-layer baseline passes hardware/reference validation:
  139.62 seconds package planning, 62632 bytes maximum initial exchange tables,
  final modelled cost 1032274 cycles. The renderer-trimmed hardware span is
  988890 cycles (0.65926 ms); maximum absolute numerical error is 0.077148.
  Profile: `artifacts/baseline-local-planner/current-full1/model.html`.
  This is a baseline-only run, not a claim of recovered optimized performance.
- Current baseline, small FP8 ViT: two and three layers pass hardware/reference
  checks (maximum absolute errors 0.105469 and 0.189453). These exercise both
  table-based and arithmetic Repeat patches. Rendered three-layer profile:
  `artifacts/baseline-local-planner/hoisted-small-r3/model.html`.
- Full 27-layer baseline now clears exchange storage and generated-code placement.
  Exact arithmetic Repeat patches reduce the maximum table from 95080 to
  63784 bytes. Sharing loop state reduces generated code from 25488 to 23048
  bytes. Non-arithmetic sequences retain their full tables; no approximation is
  made to instruction words.
- It still fails final tensor placement: tile 552 cannot place an 82944-byte
  standard allocation into the remaining tensor range 531792..949168.
  The run stops after 51.52 seconds, without a fallback search. See
  `artifacts/baseline-local-planner/hoisted27/run.log` and its memory profile.
  The mid estimator's total peak is 419840 bytes, already above that final
  417376-byte tensor range; its provisional support reservation is insufficient
  here. This is not evidence that the full model is impossible to fit.

The baseline policy still needs calibration; this replacement does not claim
that every previously representable graph now has a feasible starting plan.

## Subsequent support-storage work

See [support storage consolidation](SUPPORT_STORAGE_2026_09_11.md): 31008 bytes
per tile recovered, unused GEMM entry points removed, and final one-layer hardware
validation at 970428 cropped cycles. The 27-layer baseline still fails tensor
placement despite the improved support footprint.
