# Invariant audit, 2026-09-12

Implemented in six commits on `main`, from `5f34a28` through `2d4e69d`.
Three parallel audits examined layout semantics, planning/memory invariants,
and Repeat/storage aliases. The parent audit also compiled and ran a fresh
two-layer SigLIP package, which exposed an exchange failure missed by unit tests.

## Confirmed failures and fixes

| Area | Counterexample | Change |
| --- | --- | --- |
| Cost memoization | A 16 KiB rearrangement costs 2,336 cycles through the underlying model, but 149,504 through its cache on a 64-tile device. | Removed the additional device/active-tile multiplier. The underlying estimator already accounts for busiest-owner work. |
| Memory estimates | An early Repeat yield was freed before the body/backedge finished; an identity graph reported no initial live memory. | Keep yields live through the boundary and observe initial allocations. The small yield regression changes its peak from 2,048 to 2,560 bytes. |
| Pareto comparison | Different raw metrics with identical saturated objective vectors could mutually dominate. | Require strict improvement in the actual compared objectives. |
| Output/storage bindings | Reused subviews could escape as unallocated placeholders, or be rebound to writable storage without moving their data. | Enforce canonical storage for writable/reduction/Repeat bindings. Export full identity copies as explicit aliases; materialize exported subviews. |
| Borrowed reads | Conversion paths used a placeholder's strides; a scalar borrowed from a wider tensor lost its broadcast shape. | Normalize reads before physical copy planning and kernel construction; preserve canonical semantic shape for broadcast. Removed four scattered source-ID substitutions. |
| Repeat state | Shared initial values were overwritten; valid sequence prefixes were rejected; indirect readers of donated state were missed. | Isolate shared carried inputs, consume only the requested sequence prefix, and follow aliases when checking donation. |
| Repeat addresses | A Sum alias of an iterated parameter used an absolute first-iteration address in both local copies and exchanges. | Share storage-chain traversal across kernels, copies, and exchange relocation, retaining signed alias offsets. Removed duplicated local-copy address assembly. |
| Fragment ownership | A primitive cast on a tile owning multiple row fragments repeatedly read its first fragment. | Match coordinates for same-shaped non-product operands. This also covers LayerNorm feature inputs. |
| Mapped padding | A factored copy with an entirely padded output owner either rejected the empty mapping or left SRAM contents in its output. | Skip empty logical intersections and initialize fresh outputs with no incoming mappings. Ordinary and factored copies share this handling. |
| Exchange legality | The fresh ViT build failed while encoding a partial-word span. | Gather unaligned source spans locally before word exchange, using existing copy kernels and ordinary staging allocations. Identical source slices share their gather. |
| Redundant staging | Aligned row-major receivers used an intermediate row-major buffer followed by a copy into the final allocation. | Receive directly into the final buffer; retain staging for receiver misalignment or packed output conversion. |
| Profiling | An active tile with an empty Repeat body had an unwritten/missing sample. | Profile the empty loop as one idle interval with both timestamps. |

## Validation

- Final `cargo test --workspace --lib`: **362 passed, 0 failed, 7 ignored**.
  The codegen portion has 281 passing tests.
- Final `cargo check --workspace`: passed.
- Independent byte interpreter: **4,096 randomized copy programs**, each
  executed before and after composition, for **8,192 executions**. It compares
  actual placed transfer bytes, local copies and fills with a logical-coordinate
  reference. Cases include crops, forward/inverse factor views, padding,
  replicas, linear fragments and rotated ownership. All passed after the fix.
  The permanent test uses 256 programs to keep ordinary testing inexpensive.
- Four initial view/storage regressions were run against the previous revision:
  all four failed there and passed after the fixes. Other targeted tests cover
  primitive fragments, Repeat pointer relocation, negative pointer offsets,
  source packing and shared carried state.
- Newly compiled **SigLIP So400m, batch 1, two layers**, FP8 GEMMs with scale -4,
  fused QKV, capacity baseline, B1024 exchanges, no local optimization steps:
  **two successive inferences passed against FP32**, uploading parameters once.
  Minimum cosine similarity was **0.996969216** on each inference.
- Rendered the resulting cycle profile successfully. Exact placement HTML was
  generated by the same package build.

The destination-staging simplification changed the detailed profiling sample
count from 555,728 to 545,041. The final placed cycle estimate changed from
2,442,278 to 2,435,916. These are sample counts and compiler estimates, not a
measurement of device speedup.

## Evidence

- [Final test log](invariants-workspace-tests-20260912.log)
- [Workspace check](invariants-workspace-check-final-20260912.log)
- [Final hardware log](invariants-vit2-hardware-final-20260912.log)
- [Rendered cycle profile](invariants-20260912/vit2-final.html)
- [Exact memory placement](invariants-20260912/memory-final/placement-532187.html)
- [Mapping failures before the fix](layout-mapping-before-20260912.log)
- [8,192-execution mapping sweep](layout-mapping-broad-20260912.log)
- [Four storage/view failures before the fix](layout-boundary-before-20260912.log)
- [Initial hardware-build failure](invariants-vit2-hardware-20260912.log)

## Scope and limits

This pass grew the source tree overall, chiefly through regression tests and
the independent byte interpreter. It consolidated address resolution and copy
handling, but it is not another net source-size reduction.

The alias-lifetime audit found no additional allocator overlap/residency
counterexample: alias groups retain the union of their ranges/lifetimes and a
parameter member pins the complete group permanently. This is evidence from
inspection and tests, not a proof of the allocator.

Nested finalized Repeat and cross-coupled carried updates requiring separate
copy-back storage remain unsupported. The latter now rejects instead of
overwriting an indirectly live value. Partial-word *total payloads* still need
a separate representation; the new source gather repairs fragmented or
unaligned reads whose complete payload is word-sized.

The 27-layer searches and real-weight accuracy evaluation were not rerun in
this pass. Cost-model corrections can change the plans selected by a future
search. This audit does not establish that no other bugs remain.

## Commits

- `748125b` Profile empty Repeat bodies with a complete idle interval
- `3d3776e` Correct cost memoization, region memory accounting, and Pareto dominance
- `e7c6af3` Resolve borrowed copy storage at execution boundaries and preserve Repeat state
- `b2a0c65` Preserve Repeat pointer relocation through storage aliases
- `cc2c2bc` Respect copy padding and per-tile fragment coordinates
- `2d4e69d` Gather partial-word source spans and avoid redundant row-major staging
