Whole-phase staging with unchanged recipient sets and payload bytes. Synthetic addresses; no placement or hardware validation. Copy cycles are optimistic throughput floors, excluding barriers and setup. Scheduler results use ordinary B1024 transfers and synthetic placement; rows are per-phase, before sharing. Affine loops are representability estimates, not generated code.
| Phase | Staging | Transfers | Max endpoint fragments | Max scratch KiB/tile | Copy cycle floor | Affine loops, all tiles | Max row bytes | Exchange cycles (model) | Provenance |
|---|---|---|---|---|---|---|---|---|---|
| 50 | original | 62,482 | 2,109 | 0.0 | 0 | 0 | 17908 | 33758 | READ/WRITE DEPENDENCIES: staging results invalid. op Some(18) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 50 | source | 62,482 | 2,109 | 35.8 | 4,576 | 2,471 | — | — | READ/WRITE DEPENDENCIES: staging results invalid. op Some(18) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 50 | destination | 5,603 | 114 | 91.5 | 11,712 | 122,021 | — | — | READ/WRITE DEPENDENCIES: staging results invalid. op Some(18) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 50 | both | 5,603 | 114 | 127.2 | 16,288 | 124,492 | — | — | READ/WRITE DEPENDENCIES: staging results invalid. op Some(18) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 55 | original | 303,042 | 1,913 | 0.0 | 0 | 0 | 7936 | 14672 | READ/WRITE DEPENDENCIES: staging results invalid. op Some(21) OperatorInputs: TensorShape([4304, 1152]) F8F143 { scale_exponent: -4 } Amp(TransposedLeft) -> TensorShape([4304, 1152]) BlockMajor(Matrix { row_block: 192, column_block: 16 }) |
| 55 | source | 4,326 | 163 | 30.5 | 3,904 | 1,696 | — | — | READ/WRITE DEPENDENCIES: staging results invalid. op Some(21) OperatorInputs: TensorShape([4304, 1152]) F8F143 { scale_exponent: -4 } Amp(TransposedLeft) -> TensorShape([4304, 1152]) BlockMajor(Matrix { row_block: 192, column_block: 16 }) |
| 55 | destination | 4,326 | 163 | 27.0 | 3,456 | 1,609 | — | — | READ/WRITE DEPENDENCIES: staging results invalid. op Some(21) OperatorInputs: TensorShape([4304, 1152]) F8F143 { scale_exponent: -4 } Amp(TransposedLeft) -> TensorShape([4304, 1152]) BlockMajor(Matrix { row_block: 192, column_block: 16 }) |
| 55 | both | 3,246 | 22 | 57.5 | 7,360 | 3,305 | — | — | READ/WRITE DEPENDENCIES: staging results invalid. op Some(21) OperatorInputs: TensorShape([4304, 1152]) F8F143 { scale_exponent: -4 } Amp(TransposedLeft) -> TensorShape([4304, 1152]) BlockMajor(Matrix { row_block: 192, column_block: 16 }) |
| 16 | original | 54,502 | 1,420 | 0.0 | 0 | 0 | 12056 | 29826 | READ/WRITE DEPENDENCIES: staging results invalid. op Some(4) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 16 | source | 54,502 | 1,420 | 32.2 | 4,128 | 2,300 | — | — | READ/WRITE DEPENDENCIES: staging results invalid. op Some(4) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 16 | destination | 4,930 | 94 | 68.6 | 8,784 | 106,503 | — | — | READ/WRITE DEPENDENCIES: staging results invalid. op Some(4) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 16 | both | 4,930 | 94 | 100.9 | 12,912 | 108,803 | — | — | READ/WRITE DEPENDENCIES: staging results invalid. op Some(4) OperatorInputs: TensorShape([2, 729, 1152]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 1152]) Amp(Left) |
| 54 | original | 196,830 | 1,374 | 0.0 | 0 | 0 | 11348 | 10220 | op Some(21) LayoutRearrangement: TensorShape([2, 729, 4304]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 4304]) Amp(Left) |
| 54 | source | 196,830 | 1,374 | 8.4 | 1,076 | 1,458 | 11348 | 10220 | op Some(21) LayoutRearrangement: TensorShape([2, 729, 4304]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 4304]) Amp(Left) |
| 54 | destination | 17,496 | 208 | 34.5 | 4,416 | 139,968 | 1520 | 9238 | op Some(21) LayoutRearrangement: TensorShape([2, 729, 4304]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 4304]) Amp(Left) |
| 54 | both | 16,767 | 115 | 42.9 | 5,492 | 141,426 | 1512 | 9239 | op Some(21) LayoutRearrangement: TensorShape([2, 729, 4304]) F8F143 { scale_exponent: -4 } Amp(Left) -> TensorShape([2, 729, 4304]) Amp(Left) |
| 56 | original | 586,368 | 1,206 | 0.0 | 0 | 0 | 5168 | 16724 | op Some(21) OperatorInputs: TensorShape([2, 729, 1152]) F16 Amp(Left) -> TensorShape([25344]) RowMajor |
| 56 | source | 32,576 | 67 | 50.6 | 6,480 | 32,576 | 716 | 15103 | op Some(21) OperatorInputs: TensorShape([2, 729, 1152]) F16 Amp(Left) -> TensorShape([25344]) RowMajor |
| 56 | destination | 586,368 | 1,206 | 51.8 | 6,624 | 1,488 | 5140 | 15109 | op Some(21) OperatorInputs: TensorShape([2, 729, 1152]) F16 Amp(Left) -> TensorShape([25344]) RowMajor |
| 56 | both | 32,576 | 67 | 102.4 | 13,104 | 34,064 | 716 | 15103 | op Some(21) OperatorInputs: TensorShape([2, 729, 1152]) F16 Amp(Left) -> TensorShape([25344]) RowMajor |
| 70 | original | 139,808 | 1,013 | 0.0 | 0 | 0 | 8524 | 11465 | op Some(35) OperatorInputs: TensorShape([2, 729, 1152]) F16 RowMajor -> TensorShape([32, 729, 80]) RowMajor |
| 70 | source | 139,808 | 1,013 | 13.3 | 1,706 | 24,271 | 8420 | 9521 | op Some(35) OperatorInputs: TensorShape([2, 729, 1152]) F16 RowMajor -> TensorShape([32, 729, 80]) RowMajor |
| 70 | destination | 139,808 | 1,013 | 31.9 | 4,082 | 22,335 | 7756 | 9472 | op Some(35) OperatorInputs: TensorShape([2, 729, 1152]) F16 RowMajor -> TensorShape([32, 729, 80]) RowMajor |
| 70 | both | 129,600 | 919 | 44.9 | 5,788 | 46,606 | 7288 | 9427 | op Some(35) OperatorInputs: TensorShape([2, 729, 1152]) F16 RowMajor -> TensorShape([32, 729, 80]) RowMajor |