Whole-phase staging with unchanged recipient sets and payload bytes. Synthetic addresses; no placement or hardware validation. Copy cycles are optimistic throughput floors, excluding barriers and setup. Affine loops are representability estimates, not generated code.
| Phase | Staging | Transfers | Max endpoint fragments | Max scratch KiB/tile | Copy cycle floor | Affine loops, all tiles | Provenance |
|---|---|---|---|---|---|---|---|
| 57 | original | 104,922 | 154 | 0.0 | 0 | 0 | op Some(22) OperatorInputs: TensorShape([2, 729, 1152]) F16 Amp(Left) -> TensorShape([2, 729, 1152]) RowMajor |
| 57 | source | 11,666 | 26 | 2.5 | 320 | 11,666 | op Some(22) OperatorInputs: TensorShape([2, 729, 1152]) F16 Amp(Left) -> TensorShape([2, 729, 1152]) RowMajor |
| 57 | destination | 104,922 | 154 | 4.5 | 576 | 1,610 | op Some(22) OperatorInputs: TensorShape([2, 729, 1152]) F16 Amp(Left) -> TensorShape([2, 729, 1152]) RowMajor |
| 57 | both | 11,666 | 26 | 7.0 | 896 | 13,276 | op Some(22) OperatorInputs: TensorShape([2, 729, 1152]) F16 Amp(Left) -> TensorShape([2, 729, 1152]) RowMajor |
| 17 | original | 5,710 | 10 | 0.0 | 0 | 0 | op Some(4) OperatorInputs: TensorShape([2, 729, 3456]) F16 Amp(Left) -> TensorShape([4736]) RowMajor |
| 17 | source | 5,710 | 10 | 27.8 | 3,552 | 1,450 | op Some(4) OperatorInputs: TensorShape([2, 729, 3456]) F16 Amp(Left) -> TensorShape([4736]) RowMajor |
| 17 | destination | 5,710 | 10 | 37.0 | 4,736 | 2,830 | op Some(4) OperatorInputs: TensorShape([2, 729, 3456]) F16 Amp(Left) -> TensorShape([4736]) RowMajor |
| 17 | both | 5,710 | 10 | 64.8 | 8,288 | 4,280 | op Some(4) OperatorInputs: TensorShape([2, 729, 3456]) F16 Amp(Left) -> TensorShape([4736]) RowMajor |
| 74 | original | 128 | 4 | 0.0 | 0 | 0 | op Some(35) OperatorInputs: TensorShape([32, 1, 72]) F16 Amp(Left) -> TensorShape([32, 1, 80]) Amp(Left) |
| 74 | source | 128 | 4 | 0.0 | 4 | 128 | op Some(35) OperatorInputs: TensorShape([32, 1, 72]) F16 Amp(Left) -> TensorShape([32, 1, 80]) Amp(Left) |
| 74 | destination | 128 | 4 | 0.1 | 14 | 32 | op Some(35) OperatorInputs: TensorShape([32, 1, 72]) F16 Amp(Left) -> TensorShape([32, 1, 80]) Amp(Left) |
| 74 | both | 128 | 4 | 0.1 | 18 | 160 | op Some(35) OperatorInputs: TensorShape([32, 1, 72]) F16 Amp(Left) -> TensorShape([32, 1, 80]) Amp(Left) |
| 53 | original | 729 | 1 | 0.0 | 0 | 0 | op Some(19) LayoutRearrangement: TensorShape([2, 729, 4304]) F16 RowMajor -> TensorShape([2, 729, 4304]) RowMajor |
| 53 | source | 729 | 1 | 8.4 | 1,076 | 729 | op Some(19) LayoutRearrangement: TensorShape([2, 729, 4304]) F16 RowMajor -> TensorShape([2, 729, 4304]) RowMajor |
| 53 | destination | 729 | 1 | 8.4 | 1,076 | 729 | op Some(19) LayoutRearrangement: TensorShape([2, 729, 4304]) F16 RowMajor -> TensorShape([2, 729, 4304]) RowMajor |
| 53 | both | 729 | 1 | 8.4 | 2,152 | 1,458 | op Some(19) LayoutRearrangement: TensorShape([2, 729, 4304]) F16 RowMajor -> TensorShape([2, 729, 4304]) RowMajor |
| 58 | original | 729 | 1 | 0.0 | 0 | 0 | op Some(22) LayoutRearrangement: TensorShape([2, 729, 1152]) F16 RowMajor -> TensorShape([2, 729, 1152]) RowMajor |
| 58 | source | 729 | 1 | 2.2 | 288 | 729 | op Some(22) LayoutRearrangement: TensorShape([2, 729, 1152]) F16 RowMajor -> TensorShape([2, 729, 1152]) RowMajor |
| 58 | destination | 729 | 1 | 2.2 | 288 | 729 | op Some(22) LayoutRearrangement: TensorShape([2, 729, 1152]) F16 RowMajor -> TensorShape([2, 729, 1152]) RowMajor |
| 58 | both | 729 | 1 | 2.2 | 576 | 1,458 | op Some(22) LayoutRearrangement: TensorShape([2, 729, 1152]) F16 RowMajor -> TensorShape([2, 729, 1152]) RowMajor |