# Explicit compiler driver

`compile.rs` now owns both ordered whole-program search and complete evaluation of one candidate. Evaluation shows expansion/screening, provisional placement and scheduling, kernel compilation, package support sizing, final placement and scheduling, the best cheap address alternative, scheduled costing, and final emission in that order. There is no finalization callback from search into package code.

Package support sizing returns concrete code/row/descriptor reservations and auxiliary requests. It does not place tensors or schedule exchanges. Package emission consumes final physical work and checks its actual capacities. `ScheduledPlan` and the separate `BuiltApplication` are replaced by one evaluated result with final low work, placement, exchanges, application, cost and schedule cache. Provisional placement/phases remain local to evaluation and are dropped after sizing. Accepted address alternatives now return their cache; failed/slower mapping or address alternatives cannot replace the incumbent cache. Search's mapping estimate now sees the actual final placement, rather than the old provisional member.

Proposals and checkpoint I/O move from package into planner; compiler configuration and diagnostic benchmark entry points are owned by compile. Mid's graph builder, family choices/configuration and fragment construction still require their planned migration. Scoped owner maps and other proposal requirements remain unfinished. This structural split adds 135 production source lines including explicit imports, documentation and the concrete support reservation record; it is not claimed as a line-count reduction.

Validation:

- Workspace release tests: 419 passed, 0 failed, 8 ignored.
- All-target workspace check and `git diff --check`: clean.
- The separately enabled SDK test `saved_search_preserves_budget_aliases_and_order` passes. It compiles actual packages (no device): four uninterrupted search steps match two steps resumed for two more, across different Rayon thread counts. It checks final low work, placement, exchanges, modeled cost, alias normalization, attempt bounds, zero-step replay and an infeasible baseline. Six former mock-finalizer tests are replaced by this real driver test; the independent memoization/configuration tests are retained under planner.
- Small FP16 MLP Repeat 3: device PASS, maximum error 0.000183, modeled total/exchange cycles 32,613 / 18,087.
- Small FP8 ViT with fused QKV and Repeat 3: device PASS, maximum error 0.189453, minimum cosine 0.997459656, modeled total/exchange cycles 331,874 / 120,708.
- Real-weight SigLIP So400m, 27 layers, BS1, six real images: device PASS. All six output files are bit-identical to the preceding tree; minimum cosine against FP32 reference is 0.994567066. Modeled total/exchange cycles remain 13,301,152 / 5,422,804. The historical search checkpoint is resumed with zero optimization steps.

Commands, profiles, logs and full-model outputs are adjacent. Workspace logs are `/tmp/ipu-compile-workspace-tests.log`, `/tmp/ipu-compile-final-check.log`, and `/tmp/ipu-compile-replay-test.log`. The hardware binary predates only final removal of redundant reference borrows; the final all-target check covers those source simplifications. No hardware speedup is claimed.
