# Optimized materialized ViT exchange frontiers

Capture build: resident 27-layer batch-1 ViT, FP8 weights (scale -4), fused QKV,
automatic attention, 8 local optimization steps, compact stream size 256.
The unmodified production scheduler is used to capture transfers. The capture
retains its selected ordinary/paired widths and all Repeat source addresses.

Source revision: 2eda5dd (production binary). Plot tooling: 518a62e.

The build repeats the optimization recipe of the 9.625460 ms validated run in
`../baseline-local-planner/parallel-pairwise-balanced-full27/`. Only one host
invocation is packaged for this capture (the validated run used three).
The capture is from a fresh build, not an extraction from that old executable.
`current-profile-exchanges.json` describes the earlier hardware profile.

`plots/` compares every captured phase under 19 direct scheduling configurations.
Eight configurations run concurrently, with one Rayon worker each. Compiler
wall times therefore include contention; scheduled device cycles do not.
The black ring is a fresh S256/B1024 policy replay at captured addresses, which
can differ from a reused recipe in production.

Rows are measured before cross-phase sharing and Repeat patch support. Individual
maximum-row sizes cannot be summed to infer the full package's SRAM footprint.

Completed: 1,539 measurements across all 80 phases, including 19 paired alternatives
for phase 43. Open `combined/index.html`. Six hardware payload replays passed.
The new compact policy passed the full model at 9.582696 ms, with a 37,384-byte
exchange table and FP32 cosine 0.994168165. See `small-streams-full27/model.html`.
Production now adds S64 under the original S256 maximum/total row caps; the plots
retain the old S256/B1024 reference. Detailed compiler/runtime comparisons are in
`../../docs/EXCHANGE_FRONTIERS_MATERIALIZED_2026_09_11.md`.
