# Next: explicit attention probabilities, statistics and workspace

The grouping stage is committed as 454155c9. This is investigation/planning, not an implemented attention change.

Observed source:

- planner/attention.rs allocates nominal F16 [heads, queries, key_block + 16] probabilities, or F8 [heads, queries, key_block + 64]. It explicitly copies the key_block prefix before distributed PV to avoid treating trailing FP32 bits as probabilities.
- attention_softmax_f16.S places maxima and denominator arrays after the entire packed probability matrix, not after each physical row. Address = probability_base + rows * key_block * element_bytes; each array has rows FP32 words.
- attention_softmax_split.inc adds three FP32 segment maxima and three segment sums after the two persistent arrays. Total FP32 workspace is eight planes of rows words.
- attention_softmax_fp8.inc uses an additional 16-half (32-byte) masked-tail temporary per row after those eight planes. Complete FP8 panels do not use it.
- attention_merge_f16.S only reads the first two FP32 planes from the softmax buffer. Its probability precision and key_block arguments exist to derive that pointer; it does not read probabilities.
- The running merge accumulator is already genuinely F32 (values plus two scalars), so its extended width is not the same mixed-dtype problem.
- low/initialization.rs exempts AttentionSoftmax and AttentionMerge by kernel name because nominal F16 storage hides FP32 writes.
- kernel/gemm.rs::gemm_rows multiplies all non-column physical extents. Softmax runs once per low shard, using flattened heads/query rows. A field-major F32 tensor [8, heads, queries] therefore describes the plane layout naturally. Preserve the existing head/query distribution and physical padding when projecting that layout. The existing project_tiling helper in planner/fragments.rs may provide the right axis projection and replication behavior.

Preferred direction to evaluate:

1. Softmax emits a pure probability tensor with key_block columns plus an explicit F32 statistics/workspace result. For masked FP8, expose a separate F16 tail workspace too (or eliminate that scratch in the kernel if a simple implementation exists). The eight F32 planes comprise two persistent fields and six private work planes; keep that distinction explicit in the family contract and have merge read the first two planes. Do not merely attach a flag to the old fake F16 tensor.
2. Extend the fragment construction helper to construct the complete multi-output call, including per-result reuse. Use existing Compute::Kernel results/output_aliases and ordinary mid values; do not add a new selected-operator or temporary-program representation.
3. Give assembly explicit state/workspace pointers. Keep current worker frame slot indices 0=probabilities, 1=scores, 2=rows, 3=valid keys; append state/workspace slots to avoid gratuitous worker rewrites. Update supervisor frame requirements and cost inputs. The masked FP8 path would still fit the supervisor argument registers with three outputs, one input, and three scalars.
4. Merge consumes the explicit F32 statistics pointer. Remove its now-unnecessary key_block_columns / probability-precision specialization parameters and address arithmetic. This should also consolidate merge code variants; measure rather than assume a speedup.
5. Distributed PV can consume the probability value directly, removing the statistics-stripping copy. Remove name-based finite-padding exceptions once the write/storage types make the exclusion explicit.

Separate pointers appear sufficient for these ordinary loads/stores; there is no observed hardware requirement for the old adjacency. Preserve packed placement if an actual kernel requirement emerges. Verify numerical behavior, stack/argument ABI, multi-output ownership, workspace padding, F16/F8 full and masked softmax, whole-row/split-row schedules, Flash merge initial/intermediate/final paths, and materialized PV. Check device performance and placement on the existing model fixtures; do not assume byte-identical packages after this ABI/layout change.

Still outstanding beyond this stage: common resolved read bindings and removal of borrowed-placeholder repair; complete local-copy family binding; remaining movement/routing-policy audit; cache measurements/consolidation; requirement/operation-extension/non-test-size audit for the full compiler refactor.
