# BS2 search, 2026-09-12

Pinned compiler: 922453b. Started using run.sh, unified exec session 52619,
compiler PID 485059. Resumes fragmentation-20260912/state.json at attempt 2,
with a 96-candidate invocation budget and 24 Rayon threads. Keep running.
Output log: run.log. Checkpoint: state.json. Final package: model.ipuexe.
Successful builds run two resident FP32-reference inference checks.

Incumbent: 27-layer SigLIP BS2, fused QKV, FP8 projection/MLP multiplies,
FP16 accumulation and encoder FlashAttention. Weight layouts have one replica.
Main GEMMs cast activations before distribution. Encoder GEMM compute tiles:
QKV 1440, attention output 1440, MLP up 1458, MLP down 1440.
The MAP attention is materialized.

First shortlist has 93 distinct lowered programs after cycle screening.
Alternatives include operator plans, producer/consumer boundary changes,
early/late casting, packing distribution, and reduction grouping. Both cast
storage policies are offered. No actual shifted cast donations have appeared
in the debug log as of 18:36 UTC. Large late casts still combine packing.

At 18:36 UTC, no accepted challenger yet. First four completed candidates failed
placement. Four unique failed tile-0 dumps: peak-live deficits 16328, 10696,
10696 bytes, and one with 5688 bytes aggregate headroom. Aggregate capacity is
not sufficient to guarantee bank/lifetime/contiguity feasibility.

summarize.py compares candidate operation sequences with the baseline and writes
candidates.json. First shortlist coarse tensor+support peaks: 514680–577144 B;
baseline 514680 B. These are not physical post-support placement measurements.

While the pinned run continued, cleanup 52d1a90 removed 90 non-test lines from
placement/profiling and cast handling. Full compiler suite: 256 passed, 5 ignored;
targeted cast tests also passed after the final local simplification.
