# Typed attention results and workspaces

Softmax declares probability, FP32 maximum/denominator, FP32 segmented reduction workspace, and (only for masked FP8 output) FP16 tail results. The attention family shares their row geometry between mid construction and local call validation. Kernels receive explicit pointers, with no required adjacency or distinct memory elements. Merge consumes only statistics, so it no longer specializes on probability precision or key width. PV uses the probability result directly. The probability-prefix copy and attention-name exceptions in finite-padding reuse are removed.

The existing fragment emit/compute constructors now take complete result lists and indexed reuse requests. There is no additional workspace IR or new compatibility constructor. Initial FP16 merge has no previous-accumulator graph operand; its family fills the unused pointer slot in the concrete ABI. Binding checks the row strides the merge assembly actually uses.

Validation:

- Release workspace: 439 passed, 0 failed, 8 ignored, including doctests. The count falls by one because the obsolete name-blacklist test was removed. All-target checking, formatting and diff whitespace checks pass without warnings.
- SDK saved-search replay passes, covering uninterrupted/resumed search, budgets, visited choices, placement and scheduling.
- Attention specialization tests bind actual tensor views instead of fabricated one-buffer calls. Distributed PV expansion checks that padded groups cannot transfer FP32 statistics as probabilities.
- Standalone softmax: 640 hardware cases across full/masked blocks, FP16/FP8 output, whole/split rows, random/constant rows and extreme finite inputs. The explicit buffers occupy different memory elements. Checks cover numerical probabilities, FP32 statistics, zero padding and guards around every result/workspace. The checker now runs one implementation against the numerical reference; its old paired ABI harness and binary are archived in before/ for the one-time comparison.
- Whole/split pre-change FP16 measurements are in before-whole/ and before-split/. Representative short-row calls save 12–42 cycles. Longer-row timings vary in both directions with the new storage/instruction arrangement; this is not evidence of a general speedup. The issue model removes the obsolete address arithmetic and its regression compares estimates with measured samples, allowing the existing bank/skew approximation.
- Small ViT, FP8 projections/probabilities, Repeat3: hardware PASS, minimum cosine 0.996164615, maximum absolute error 0.275879.
- Streaming attention, 729 keys, two heads and 64 working tiles: hardware PASS, maximum absolute error 0.000031. This covers reuse of probabilities/statistics across key blocks and online merge state.
- Real pretrained 27-layer SigLIP, six images, FP32 reference: hardware PASS; minimum cosine 0.994567066. All six output files are byte-identical to the pre-change exchange-relocation run. Final build plus six inferences/reference checking took 102.431 s, peak host RSS 6,296,220 KiB; this is not a controlled compiler benchmark.
- Default FP16 MLP Repeat3 package remains byte-identical to scoped-grouping. Attention packages intentionally change because their buffer and argument contracts change.

The full refactor remains active. Resolved read binding, local-copy binding, remaining movement-policy ownership, measured cache consolidation and the overall size/extension audit are unfinished. This slice does not establish the requested overall non-test code reduction.
