Anonymous supplementary materials

Learning When to Reason

Predictive Timing for Compute-Constrained Embodied Agents

Same calls. Better timing. More events covered.

PRS learns pre-event opportunity scores from recent state history and compact bird's-eye-view context. We evaluate when requests are placed by comparing constrained call sets with exactly the same number of invocations.

Prediction, selection, and continuous control

PRS opportunity scoring, constrained call selection, continuous fast-policy control, and independent reactive recovery
Opportunity scores guide constrained selection; the fast policy maintains continuous control. Reactive recovery is accounted for separately. The Town10HD images provide context from a separate visualization run.

Same call count, more event windows

165calls per method at the 2% budget
23 / 63event windows covered by the selected PRS checkpoint
~6 / 63covered by Random-k at the same budget

On frozen CARLA Town10HD trajectories, the UWC@0.5 s difference is 26.71 percentage points (95% bootstrap interval: 12.64 to 40.97). Random-k is the strongest non-oracle baseline at this operating point.

Primary equal-call coverage frontier and paired fresh-scenario comparisons
Current manuscript Figure 2. The primary checkpoint comparison and the fresh-scenario selection-procedure evaluation are distinct experiments.

A separately frozen development-selection procedure yields a normalized log-budget UWC-AUC gain of 0.2392 over Random-k on fresh Town10HD scenarios, with positive differences at all five budgets. These are selected-model results, not a claim that arbitrary training seeds perform consistently.

UWC measures request placement on frozen trajectories, independently of backend answer quality. Exact call parity controls invocation count, not FLOPs, tokens, energy, or elapsed compute time. Online commitment and downstream utility are evaluated separately.

Supplementary Video S1

Illustrative Town10HD timing replay, 1920 x 1080, 20 FPS, 11.5 seconds. The PRS marker precedes onset by 0.30 s; the History-MLP reference follows it by 0.30 s. Both overlays use the same separately recorded scene stream.

Ten corrupted sensor frames in the event segment use previous-valid-frame hold, as disclosed on screen. This visualization rerun is not a bit-exact reproduction of the quantitative trajectory and is not used for reported metrics. The embedded video is an H.264 compatibility export; the original download is retained unchanged. Full caption and provenance summary.

Poster, video, and numerical evidence

MaterialContents
Poster PDF / HTMLCurrent framework, primary result, and evaluation scope.
Anonymous supplementary dataPrimary and fresh-scenario tables, all-seed comparisons, target and matched-search studies, sensitivity results, and lead-time diagnostic records.
Lead-time RGB frames / cases 01--02 and cases 03--04Original RGB inputs used by the auxiliary human-review viewer, split into two downloads for static hosting.
Single-reviewer observations / CSVAll 21 existing observations, corrections, missing fields, and provisional associations. A viewer with the original RGB inputs is included under lead_time/human_review/ in the data archive. These observations are not used to compute Q/T/C.
Video S1Original supplementary visualization, preserved without re-encoding.
Material guideVersion information, included files, and interpretation of the video.

All-seed results and the complete matched-search contrasts remain in the manuscript and data export. The downloadable data are an evidence release, not a complete training environment or checkpoint distribution.