# Low-cost reranking

The default planner now assembles its eight best retained complete DP paths,
lowers each through the normal passes, and selects by pre-scheduling low cycles
(mid memory peak breaks ties). Expansion uses the configured target, tile map,
diagnostic mode and cast reuse setting. Geometry is shared across evaluations;
each expanded tile graph is discarded after costing. Packaging still lowers
the selected MidGraph normally. No placement or exchange scheduling runs during
reranking, and lowering errors propagate rather than silently falling back.

`SearchLimits.low_cost_finalists` controls the number. Zero retains mid-only
selection for the exhaustive tests of the DP's mid objective. This is a final
selection step, not a change to catalogue generation or intermediate dominance.
Candidates already removed by coarse shortlisting or DP cannot be recovered.

The randomized reranking test constructs executable alternatives with different
scratch/compute tradeoffs, prices each independently through lowering, then
deliberately reverses their mid ordering. Reranking must select the low minimum
and report its low compute/exchange estimate. Existing exact-DP tests retain
their original objective instead of silently changing their oracle.

## Coarse pruning and the SDK solver

The useful distinction is representation and search order, not a different
constant in our current shortlist:

- `../poplibs/lib/poplin/ConvPlan.cpp::choosePlan` builds a constraint model and
  minimizes cycles then temporary bytes (or another requested objective).
- `ConvModel.cpp::constructModel` defines partition variables and constraints;
  it tightens cycle, temporary-memory and tile-count limits using the best cost
  already found. The surrounding planner enumerates discrete transforms and
  kernel families, not every completed tensor graph.
- The SDK solver interface supports variable domains and variable-priority
  groups. Its objective is still a model: optimum model cost is not proof of
  optimum measured runtime.

For this planner, a corresponding direction is to search compact GEMM choices
before constructing MidGraphs. Branch first over algorithm/orientation and
partition geometry; derive local extents, padding, work and memory requirements
from partial choices. Reject impossible capacity/alignment combinations early.
Use valid lower bounds to prune branches against an incumbent, and approximate
prices to order the remaining exploration. Do not promote an approximate byte
estimate into a proof of dominance. Full graphs are needed only for promising
complete assignments, followed by low evaluation.

This requires bounds for partially specified choices, not merely applying the
current complete-plan model earlier. Where strong bounds are unavailable,
evaluation remains explicitly budgeted search. Preserve alternatives across
memory budgets and boundary representations instead of using cycle-only
crowding to approximate their diversity. Measure retained-winner regret and
planning cost against broader offline catalogues before selecting a new rule.

No coarse-pruning changes are included here.

## Validation

303 unit tests and the doctest passed. The full-size FP16 batch-one MLP
reranking replay evaluated:

| Mid rank | Mid cycles | Low cycles |
| --- | ---: | ---: |
| 0 | 237,432 | 275,574 |
| 1 | 237,672 | 275,814 |
| 2 | 237,727 | 287,089 |
| 3 | 237,983 | 287,345 |
| 4 | 238,157 | 289,320 |
| 5 | 238,371 | 289,576 |
| 6 | 238,471 | 275,356 |
| 7 | 238,698 | 283,219 |

Rank 6 wins. Production low evaluations took 3.07–5.89 seconds each, 37.74
seconds total; complete mid construction/search/reranking took 92.4 seconds.
These timings supersede the much older subsecond expansion measurements for
this particular plan family. The diagnostic test binary is slower because it
independently checks cached kernel prices against uncached calls.

Hardware passed with maximum absolute error 0.000977 and 276,012
renderer-normalized cycles. The previous recorded native-resident FP16 run was
275,562 cycles; this is not a demonstrated speedup. Its estimator also predates
the latest copy-cost correction, so that comparison does not isolate reranking.
No new hardware measurement of the current mid-only winner was made.

Artifacts: `artifacts/low-rerank-20260923/` contains the diagnostic replay,
production build/run logs, package, profile and normalized summary. Temporary
benchmark instrumentation was removed.
