# Shared storage geometry — in progress

Base: 6b41664a. The preceding goal turn made progress: committed checked local-copy family binding, removed late selection, passed workspace/SDK/hardware gates. The larger compiler-structure goal remains active and source reduction is unmet.

Current evidence: ExpansionCache owns normalized local-copy keys plus separately hand-hashed destination mappings. GeometryAnalysis reconstructs normalized view keys and traversal pairs during detailed costing. The existing expansion benchmark does not actually make a warm-cache expansion and reports entry counts without retained payload. First add that measurement, then capture disabled/cold/warm on MLP B2 and representative ViT/SigLIP plans before replacing these caches.

Memory counters estimate table/vector/Arc payload from allocated capacities; shared traversal bodies may be counted more than once, and allocator bookkeeping is excluded. Do not call those exact RSS. RSS is recorded separately.

Intended endpoint: neutral storage geometry owns normalized view/traversal/pair facts, shared by movement and detailed costing. Low binds geometry to actual storage/alias identity; movement owns physical selection; exchange costs own protocol timing/footprint calculations. Do not retain the old cache as a facade over a new duplicate one. Preserve symbolic spans and exact identity/order/padding/alias semantics. Cache lifetime for search must remain separate from per-candidate graph state and on-disk compiled kernels.

Baseline results so far (8 pinned CPU cores):
- MLP B2 Repeat3: cold expansion 1421 ms; warm 953 ms. Recost 306/305 ms; footprint 345/353 ms. Retained payload: copy recipes 67,112 bytes; destination facts 34,509,968; detailed geometry 36,382,112. Costing interns 60,051 views / 117,900 pairs, so blindly carrying the old 32,768-entry limit into a common pool would immediately undersize it.
- Materialized attention B1: cold expansion 478 ms, warm 418 ms; recost 99/117 ms; footprint 202/206 ms. Retained payload: recipes 84,608 bytes, destination facts 22,709,472, detailed geometry 2,311,008.
- Expansion-cache-disabled variants retain the per-candidate costing cache, as before. MLP expansion 1903 ms; attention 704 ms. Record that distinction instead of claiming every cache was disabled.
- Codegen release tests: 336 passed, 6 ignored. Full SigLIP baseline measurement follows after those tests finish, avoiding contention during timed sections.

Design implication for the actual consolidation:
1. A neutral storage geometry cache should intern normalized views/traversals and canonical source/destination span pairs. Source IDs, ownership, signed alias locations and actual placement remain low bindings. Costing consumes paired rows and applies exchange encoding sizes/timing itself; do not move IPU control heuristics into generic storage geometry.
2. Destination population facts should use those same interned views/pairs for alignment, exact coverage and fragment counts, replacing the separate raw extents hashing/equality/owned-key construction in ExpansionCache.
3. Local copy cache entries currently contain a selected launch policy: low/copy.rs::row_copy_pattern uses six workers and a 512-byte cutoff. Moving that cache unchanged to storage would preserve misplaced policy. The copy family should choose launches from cached geometric row groups. The current recipe cache is only ~67–85 KB, so reevaluate whether it is worth retaining at all once shared pairs avoid rebuilding traversals/spans; do not add a new policy-keyed cache automatically.
4. Prefer existing [StridedSpan; 2] as neutral affine pair facts, with full equality/order/padding preserved. Keep exact canonical pairs for exchange fragment costing; destination-sorted/recoalesced local copy geometry is a separate ordering choice and must honor backing aliases. Do not use local helper grouping to silently change exchange fragment counts.
5. Use ordinary foldhash maps unless the new benchmark proves borrowed/prehashed key machinery necessary. Retain limits and inspect actual retained bytes; the existing per-candidate view/pair population already exceeds the old search-cache entry limit. Keep per-candidate low graphs and on-disk kernel artifacts outside this cache.

Baseline measurement is committed as 9756530f, all jobs are terminal. Full SigLIP results: disabled/cold/warm expansion 8272/9982/6696 ms; warm recost 931 ms, footprint 1172 ms. Retained recipes 1.149 MB, destination geometry 149.719 MB, costing geometry 31.316 MB. See verification.md and baseline-summary.json. The cold destination cache construction is costly here; normalized common keys are more urgent than retaining the tiny selected-copy cache. No cache has yet been replaced. Next work must do the actual consolidation, not another measurement-only pass.

Additional point to validate during copy-family grouping: same-buffer currently prevents destination sorting, but coalescing can still create a strided worker call. Check cross-row alias dependencies before assuming preservation of incoming row order alone proves safety. This is an unverified concern, not a demonstrated corruption.
