Whole-phase staging with unchanged recipient sets and payload bytes. Synthetic addresses; no placement or hardware validation. Copy cycles are optimistic throughput floors, excluding barriers and setup. Affine loops are representability estimates, not generated code.
| Phase | Staging | Transfers | Max endpoint fragments | Max scratch KiB/tile | Copy cycle floor | Affine loops, all tiles | Provenance |
|---|---|---|---|---|---|---|---|
| 50 | original | 62,482 | 2,109 | 0.0 | 0 | 0 | op Some(18) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 50 | source | 62,320 | 2,108 | 35.8 | 4,576 | 2,471 | op Some(18) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 50 | destination | 5,441 | 113 | 91.5 | 11,712 | 122,021 | op Some(18) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 50 | both | 5,441 | 113 | 127.2 | 16,288 | 124,492 | op Some(18) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 55 | original | 303,042 | 1,913 | 0.0 | 0 | 0 | op Some(21) OperatorInputs: TensorShape([4304, 1152]) F8F143 { scale_exponent: -4 } Amp(TransposedLeft) -> TensorShape([4304, 1152]) BlockMajor(Matrix { row_block: 192, column_block: 16 }) |
| 55 | source | 4,142 | 162 | 30.5 | 3,904 | 1,696 | op Some(21) OperatorInputs: TensorShape([4304, 1152]) F8F143 { scale_exponent: -4 } Amp(TransposedLeft) -> TensorShape([4304, 1152]) BlockMajor(Matrix { row_block: 192, column_block: 16 }) |
| 55 | destination | 4,142 | 162 | 27.0 | 3,456 | 1,609 | op Some(21) OperatorInputs: TensorShape([4304, 1152]) F8F143 { scale_exponent: -4 } Amp(TransposedLeft) -> TensorShape([4304, 1152]) BlockMajor(Matrix { row_block: 192, column_block: 16 }) |
| 55 | both | 3,062 | 21 | 57.5 | 7,360 | 3,305 | op Some(21) OperatorInputs: TensorShape([4304, 1152]) F8F143 { scale_exponent: -4 } Amp(TransposedLeft) -> TensorShape([4304, 1152]) BlockMajor(Matrix { row_block: 192, column_block: 16 }) |
| 122 | original | 1,471 | 1,471 | 0.0 | 0 | 0 | op Some(41) LayoutRearrangement: TensorShape([2, 1, 4304]) F16 RowMajor -> TensorShape([2, 1, 4304]) RowMajor |
| 122 | source | 1,471 | 1,471 | 0.0 | 2 | 1,471 | op Some(41) LayoutRearrangement: TensorShape([2, 1, 4304]) F16 RowMajor -> TensorShape([2, 1, 4304]) RowMajor |
| 122 | destination | 1,471 | 1,471 | 16.8 | 2,150 | 1 | op Some(41) LayoutRearrangement: TensorShape([2, 1, 4304]) F16 RowMajor -> TensorShape([2, 1, 4304]) RowMajor |
| 122 | both | 1,471 | 1,471 | 16.8 | 2,152 | 1,472 | op Some(41) LayoutRearrangement: TensorShape([2, 1, 4304]) F16 RowMajor -> TensorShape([2, 1, 4304]) RowMajor |
| 16 | original | 54,502 | 1,420 | 0.0 | 0 | 0 | op Some(4) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 16 | source | 54,358 | 1,419 | 32.2 | 4,128 | 2,300 | op Some(4) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 16 | destination | 4,786 | 93 | 68.6 | 8,784 | 106,503 | op Some(4) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 16 | both | 4,786 | 93 | 100.9 | 12,912 | 108,803 | op Some(4) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |