Whole-phase staging with unchanged recipient sets and payload bytes. Synthetic addresses; no placement or hardware validation. Copy cycles are optimistic throughput floors, excluding barriers and setup. Affine loops are representability estimates, not generated code.
| Phase | Staging | Transfers | Max endpoint fragments | Max scratch KiB/tile | Copy cycle floor | Affine loops, all tiles | Provenance |
|---|---|---|---|---|---|---|---|
| 62 | original | 147,712 | 2,020 | 0.0 | 0 | 0 | op Some(21) LayoutRearrangement: TensorShape([2, 577, 4096]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 577, 4096]) Amp(Left) |
| 62 | source | 147,712 | 2,020 | 8.0 | 1,024 | 577 | op Some(21) LayoutRearrangement: TensorShape([2, 577, 4096]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 577, 4096]) Amp(Left) |
| 62 | destination | 8,655 | 113 | 55.1 | 7,056 | 138,480 | op Some(21) LayoutRearrangement: TensorShape([2, 577, 4096]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 577, 4096]) Amp(Left) |
| 62 | both | 8,655 | 113 | 63.1 | 8,080 | 139,057 | op Some(21) LayoutRearrangement: TensorShape([2, 577, 4096]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 577, 4096]) Amp(Left) |
| 101 | original | 1,023 | 1,023 | 0.0 | 0 | 0 | op Some(41) LayoutRearrangement: TensorShape([2, 1, 4096]) F16 RowMajor -> TensorShape([2, 1, 4096]) RowMajor |
| 101 | source | 1,023 | 1,023 | 0.0 | 2 | 1,023 | op Some(41) LayoutRearrangement: TensorShape([2, 1, 4096]) F16 RowMajor -> TensorShape([2, 1, 4096]) RowMajor |
| 101 | destination | 1,023 | 1,023 | 16.0 | 2,046 | 1 | op Some(41) LayoutRearrangement: TensorShape([2, 1, 4096]) F16 RowMajor -> TensorShape([2, 1, 4096]) RowMajor |
| 101 | both | 1,023 | 1,023 | 16.0 | 2,048 | 1,024 | op Some(41) LayoutRearrangement: TensorShape([2, 1, 4096]) F16 RowMajor -> TensorShape([2, 1, 4096]) RowMajor |
| 58 | original | 45,375 | 1,000 | 0.0 | 0 | 0 | op Some(18) OperatorInputs: TensorShape([2, 577, 1024]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 577, 1024]) Amp(Left) |
| 58 | source | 41,406 | 972 | 37.0 | 4,736 | 4,801 | op Some(18) OperatorInputs: TensorShape([2, 577, 1024]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 577, 1024]) Amp(Left) |
| 58 | destination | 4,886 | 87 | 59.0 | 7,552 | 80,530 | op Some(18) OperatorInputs: TensorShape([2, 577, 1024]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 577, 1024]) Amp(Left) |
| 58 | both | 4,256 | 80 | 96.0 | 12,288 | 83,102 | op Some(18) OperatorInputs: TensorShape([2, 577, 1024]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 577, 1024]) Amp(Left) |
| 84 | original | 82,504 | 749 | 0.0 | 0 | 0 | op Some(35) OperatorInputs: TensorShape([2, 577, 1024]) F16 RowMajor -> TensorShape([16, 577, 128]) RowMajor |
| 84 | source | 82,504 | 749 | 11.7 | 1,500 | 9,172 | op Some(35) OperatorInputs: TensorShape([2, 577, 1024]) F16 RowMajor -> TensorShape([16, 577, 128]) RowMajor |
| 84 | destination | 82,399 | 749 | 22.0 | 2,816 | 5,364 | op Some(35) OperatorInputs: TensorShape([2, 577, 1024]) F16 RowMajor -> TensorShape([16, 577, 128]) RowMajor |
| 84 | both | 80,479 | 719 | 33.5 | 4,316 | 14,536 | op Some(35) OperatorInputs: TensorShape([2, 577, 1024]) F16 RowMajor -> TensorShape([16, 577, 128]) RowMajor |