2026-09-21T20:51:50.454452Z PL:POPLIN 1125227.1125227 W: Selecting partial type HALF which is the only valid option for Convolution where the input is of type QUARTER 2026-09-21T20:51:50.454529Z PL:POPLIN 1125227.1125227 W: Selecting partial type HALF which is the only valid option for Convolution where the input is of type QUARTER 2026-09-21T20:51:50.454532Z PL:POPLIN 1125227.1125227 W: Selecting partial type HALF which is the only valid option for Convolution where the input is of type QUARTER 2026-09-21T20:51:50.454603Z PL:POPLIN 1125227.1125227 W: Selecting partial type HALF which is the only valid option for Convolution where the input is of type QUARTER 2026-09-21T20:51:50.454605Z PL:POPLIN 1125227.1125227 W: Selecting partial type HALF which is the only valid option for Convolution where the input is of type QUARTER 2026-09-21T20:51:50.454607Z PL:POPLIN 1125227.1125227 W: Selecting partial type HALF which is the only valid option for Convolution where the input is of type QUARTER 2026-09-21T20:51:50.454616Z PL:POPLIN 1125227.1125227 D: Planning convolution with a per-tile memory limit of 383385.6 bytes across 1472 tiles. 2026-09-21T20:51:50.454638Z PL:POPLIN 1125227.1125227 D: Creating non-joint plan (NONE_MATMUL pass)... 2026-09-21T20:51:50.454648Z PL:POPLIN 1125227.1125227 D: Creating plan with objective { minimise cycles : tile temp memory bound = 383385B } 2026-09-21T20:51:50.831180Z PL:POPLIN 1125227.1125227 D: Found new best candidate plan using {"type":"AMP","convUnits":"16","convInputLoadElems":"8"}: Cost{cycles=183123, memory=286624, tiles=1440} 2026-09-21T20:51:51.032975Z PL:POPLIN 1125227.1125227 D: Found new best candidate plan using {"type":"AMP","convUnits":"16","convInputLoadElems":"8"}: Cost{cycles=107147, memory=156192, tiles=1440} 2026-09-21T20:51:51.061371Z PL:POPLIN 1125227.1125227 D: Found new best candidate plan using {"type":"AMP","convUnits":"16","convInputLoadElems":"8"}: Cost{cycles=54393, memory=91872, tiles=1440} 2026-09-21T20:51:51.176308Z PL:POPLIN 1125227.1125227 D: Evaluated a total of 18862707 constraints to find the best plan 2026-09-21T20:51:51.176379Z PL:POPLIN 1125227.1125227 D: Found best plan using {"type":"AMP","convUnits":"16","convInputLoadElems":"8"}: Cost{cycles=54393, memory=91872, tiles=1440}. 2026-09-21T20:51:51.176388Z PL:POPLIN 1125227.1125227 D: for input {1}x(729x1x1152), kernel {1}, output = {1}x(729x1x4304), pass=NONE_MATMUL 2026-09-21T20:51:51.176391Z PL:POPLIN 1125227.1125227 D: breakdown of memory and cycle estimates: 2026-09-21T20:51:51.176394Z PL:POPLIN 1125227.1125227 D: - total parallel split: 1440 2026-09-21T20:51:51.176396Z PL:POPLIN 1125227.1125227 D: - total serial split: 1 2026-09-21T20:51:51.176398Z PL:POPLIN 1125227.1125227 D: - broadcast operands before loop: 0 copy cycles, 0 exchange cycles, 0 bytes 2026-09-21T20:51:51.176401Z PL:POPLIN 1125227.1125227 D: - rearrangement before slice: 0 cycles, 0 bytes (0 overhead, 0 per-loop iteration) 2026-09-21T20:51:51.176403Z PL:POPLIN 1125227.1125227 D: - dynamic slice: 0 cycles, unknown bytes 2026-09-21T20:51:51.176405Z PL:POPLIN 1125227.1125227 D: - transform: 4404 copy cycles, 1395 exchange cycles, 8800 bytes (input 0, weights 1178, output 4400) 2026-09-21T20:51:51.176408Z PL:POPLIN 1125227.1125227 D: - exchange: 15497 cycles, n/a bytes. (Input 6458, Weight 2151, Reduce 6888 + 0) 2026-09-21T20:51:51.176410Z PL:POPLIN 1125227.1125227 D: - tile level transform: 0 cycles, 0 bytes 2026-09-21T20:51:51.176412Z PL:POPLIN 1125227.1125227 D: - inputs cast: 0 cycles, 0 bytes 2026-09-21T20:51:51.176414Z PL:POPLIN 1125227.1125227 D: - compute: 27472 cycles, 91872 bytes 2026-09-21T20:51:51.176415Z PL:POPLIN 1125227.1125227 D: - reduction: 5625 cycles, 0 bytes 2026-09-21T20:51:51.176417Z PL:POPLIN 1125227.1125227 D: - dynamic update: 0 cycles, unknown bytes 2026-09-21T20:51:51.176419Z PL:POPLIN 1125227.1125227 D: - add in-place: 0 cycles, 0 bytes 2026-09-21T20:51:51.176420Z PL:POPLIN 1125227.1125227 D: - output cast: 0 cycles, 0 bytes 2026-09-21T20:51:51.176432Z PL:POPLIN 1125227.1125227 D: - total: 54393 cycles, 91872 bytes 2026-09-21T20:51:51.176451Z PL:POPLIN 1125227.1125227 D: Plan: transform #0 Transform: extraFieldDims 1 dilatePostConv {} swapOperands true expandDims {} outChanFlattenDims {} flattenDims {0,1,2} combineConvGroupsFactor 1 partition #0 Partition: fieldSplit {1,15} batchSplit 1 outChanSplit.serial 1 outChanSplit.parallel 16 kernelSplit {1,1} inChanSplit.serial 1 inChanSplit.parallel 6 convGroupSplit 1 fieldAxisGrainSize {1,1} inChanGrainSize 32 outChanGrainSize 16 types #0 Types: partialType half resultType half transform #1 Transform: extraFieldDims 0 dilatePostConv {} swapOperands false expandDims {} outChanFlattenDims {} flattenDims {} combineConvGroupsFactor 1 types #1 Types: partialType half resultType half convGroupsPerGroup 1 inChansPerGroup 32 partialChansPerGroup 16 method {"type":"AMP","convUnits":"16","convInputLoadElems":"8"} isJointPlan false startTile 0 linearizeTileDirection ASCENDING totalTiles 1440 broadcastInputBeforeLoop false 2026-09-21T20:51:51.176461Z PL:POPLIN 1125227.1125227 D: for params: Params: inputType quarter outputType half batchSize 729 numConvGroups 1 inputFieldShape {1} kernelShape {1} inputChannelsPerConvGroup 1152 outputChannelsPerConvGroup 4304 inputTruncationLower {0} inputTruncationUpper {0} inputDilation {1} inputPaddingLower {0} inputPaddingUpper {0} flipInput {0} kernelTruncationLower {0} kernelTruncationUpper {0} kernelDilation {1} kernelPaddingLower {0} kernelPaddingUpper {0} flipKernel {0} outputTruncationLower {0} outputTruncationUpper {0} stride {1} outputPaddingLower {0} outputPaddingUpper {0} outputFieldShape {1} 2026-09-21T20:51:51.176496Z PL:POPLIN 1125227.1125227 D: Options: availableMemoryProportion 0.6 pass NONE_MATMUL partialsType half interTilePartialsType half interIpuPartialsType half use128BitConvUnitLoad 0 planConstraints {} planConstraintsOutputFilename enableMultiStageReduce 1 enableFastReduce 0 remapOutputTensor 1 enableConvDithering 0 disableTransformations 0 insertTransformsCycleCountProgs 0 experimental.convTransformsEstimates 0 gatherConvOutput 0 experimental.slicVmac16 0 disableSRForAMPVertices 0 enableTileLevelExpandDims 0 Estimated cost {cycles 54393, temporary bytes 91872} Plan: transform #0 Transform: extraFieldDims 1 dilatePostConv {} swapOperands true expandDims {} outChanFlattenDims {} flattenDims {0,1,2} combineConvGroupsFactor 1 partition #0 Partition: fieldSplit {1,15} batchSplit 1 outChanSplit.serial 1 outChanSplit.parallel 16 kernelSplit {1,1} inChanSplit.serial 1 inChanSplit.parallel 6 convGroupSplit 1 fieldAxisGrainSize {1,1} inChanGrainSize 32 outChanGrainSize 16 types #0 Types: partialType half resultType half transform #1 Transform: extraFieldDims 0 dilatePostConv {} swapOperands false expandDims {} outChanFlattenDims {} flattenDims {} combineConvGroupsFactor 1 types #1 Types: partialType half resultType half convGroupsPerGroup 1 inChansPerGroup 32 partialChansPerGroup 16 method {"type":"AMP","convUnits":"16","convInputLoadElems":"8"} isJointPlan false startTile 0 linearizeTileDirection ASCENDING totalTiles 1440 broadcastInputBeforeLoop false