# Review 99b7e4d: assessment and fixes

The review identifies real allocation and search-bookkeeping defects. Its
"correctness" label on repeated candidate validation describes wasted search
budget, not incorrect device arithmetic. Several other findings describe the
same underlying duplication from different logging perspectives. The FP8 fusion
code has moved to `mid/output_fusion.rs` since the reviewed commit.

## Fusion safety and pricing (findings 1, 8, 13, 14)

Add/GeLU and Add/LayerNorm fusion discarded proven in-place alias chains when
those aliases were represented by primitive input edges rather than an already
unified storage group. Those edges can point to trailing dependency inputs,
not just the arithmetic operand indices. Fusion now follows the aliased value
through both operations and translates it to the replacement's input index.
Intervening copies must not read or overwrite the affected storage groups.

The whole fusion pass checks its recomputed memory peaks against the configured
standard reservation and tensor-memory budget before accepting a cycle saving.
A faster fusion that fails that check leaves the original program intact. This
uses the existing capacity model; it does not guarantee aligned physical
placement or predict exact exchange-table storage.

The regression checks both preserved alias-chain memory and a fresh-output
case where fusion genuinely increases the live peak. The latter is accepted
with space available and rejected when the original program exactly fits the
configured budget.

A shared operation-sequence cost helper and fusion decision replace the
independent arithmetic at elementwise, residual/statistics, producer-owned FP8,
and consumer-owned FP8 fusion sites. Each decision reports the same separate
and fused costs and verdict. Logs identify the fusion family instead of printing
an entire primitive and its operand windows as a "kernel".

## Search bookkeeping (findings 2, 3, 9, 10)

Implicit cast-storage policy is normalized to its configured Boolean value,
including recipes in old checkpoints. Raw open-boundary proposals and their
completed recipes are remembered together.

Equal mid programs are grouped before shortlisting, even when they are not
adjacent. Each group carries its equivalent recipes. Only groups whose results
are consumed by the ordered search are recorded as visited: discarded aliases
of a truncated or cancelled candidate must remain available. This avoids the
review's suggested pitfall of marking all deduplicated recipes visited before
their representative is actually evaluated.

After a fully rejected shortlist, the unchanged incumbent exits without an
extra lowering round. A budget-truncated tail remains eligible on checkpoint
resume. The resume regression now verifies that two successive rejected
candidates have different low programs, rather than merely different recipes.
Existing uninterrupted/resumed and parallel ordered-selection tests still apply.

## Diagnostics and shared change reporting (findings 4–7, 11, 12, 15)

Candidate memory-profile failures warn and preserve the validated incumbent.
The regression replaces the diagnostic directory with a regular file after
baseline validation, then checks that search still returns the valid package.
The initial baseline diagnostic remains fallible: there is no validated
incumbent at that point.

Screening has a round/proposal span inherited by fusion logs; validation records
the proposal index too. One Recipe change helper covers changed/removed plans,
boundary and cast changes, and scalar settings in both screening and acceptance
logs. Mid costs are named estimates, distinct from final placed-program costs.
Round summaries count invalid, visited, non-improving, deduplicated, truncated
and shortlisted proposals, including empty shortlists. Formatting is checked
normally after removing the duplicated long tracing expressions.

These fixes do not establish that every rejected layout would be unprofitable.
The separate limitations in proposal diversity and the strict estimate gate
remain search-policy questions.

## Validation

Artifacts: `artifacts/review-99b7e4d-20260914/`.

- Focused alias/memory regression and eight local-search tests pass.
- Full codegen suite: 294 passed, five ignored; no failures.
- Workspace check, targeted rustfmt check and `git diff --check` pass.
- All six pretrained hardware cases pass. Their raw embedding outputs are
  byte-for-byte identical to the previous run; minimum FP32-reference cosine
  remains 0.994567066.
- The replay uses the existing full 27-layer calibrated BS1 FP8-QK checkpoint,
  original weights, and all six image fixtures. Its outputs are saved separately
  from the previous experiment.
