autosetup: preprocessing watchdog — cancel cloud jobs stuck in prover preprocessing - #84
autosetup: preprocessing watchdog — cancel cloud jobs stuck in prover preprocessing#84shellygr wants to merge 3 commits into
Conversation
Add a preprocessing watchdog to CloudProverRunner's wait loop. The prover's treeview carries an empty rules list for as long as the run is in preprocessing (scene construction, storage/pointer analyses, CVL typechecking); rules appear only once rule checking begins. A run whose preprocessing blows up therefore sits in RUNNING with no rules until the prover's global timeout, and autosetup burned its whole 150-minute per-job budget polling it. The watchdog piggy-backs on the existing 10s status poll: queue time is free, the clock starts on the first RUNNING observation, and after a grace period it probes the treeview at a bounded cadence. A non-empty rules list makes it dormant; probe transport errors self-disable it (restoring the old behavior); exceeding the budget cancels the job and classifies the result as a new distinct JobStatus.PREPROCESSING_TIMEOUT so that: - conf workarounds never retry an identical doomed job (on_job_problem guard) and the submission cache is dropped so a deliberate future run submits fresh, - sanity_summary.md shows a dedicated row instead of crashing on the missing job report of a cancelled job. Env knobs (0 budget disables): AUTOSETUP_PREPROCESSING_BUDGET_SECONDS (default 1800), AUTOSETUP_PREPROCESSING_GRACE_SECONDS (300), AUTOSETUP_PREPROCESSING_PROBE_SECONDS (90). Verified live against a production job stuck in preprocessing: cancelled and classified ~5.5 minutes after the budget window opened, with the job showing CANCELED on the cloud. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
| self._running_since = now | ||
|
|
||
| running_for = now - self._running_since | ||
| if running_for < self.grace_seconds: |
There was a problem hiding this comment.
to Claude: why do we need grace seconds? what if preprocessing is super quick? we're introducing a 5 minute delay with this
…y answered probe The probe interval already expresses "don't probe a job that just started running": seeding the last-probe timestamp when the clock starts puts the first probe one interval in, so the separate grace period bought nothing that a second time knob had to exist for. It also removes a misconfiguration trap, since grace and budget were tunable independently with no ordering check between them. _consecutive_probe_errors now resets after any probe the endpoint answers, not only after a JobNotFoundError. An endpoint that answers is not a broken endpoint, and the counter gates a permanent self-disable.
|
Claude answers: You're right that the knob shouldn't exist — removed it. But the reasoning is worth stating, because "5 minute delay" isn't quite what it was doing. It delayed nothing. The grace only suppressed treeview probes. The job, the 10s status poll, and the returned result were untouched, and nothing downstream consumes What it actually bought was probe-saving. Measured over 702 real cloud jobs in our fleet logs, p50 whole-job RUNNING time is 124s and ~70% finish under 300s, so most jobs never issued a single probe. Per-contract fastest successful run (an upper bound on that contract's preprocessing) is p50 39s, p95 282s. But that doesn't justify a second time knob, and it had a real footgun: grace and budget were independently env-tunable with no ordering check, so Two things that came out of digging into this:
16 tests pass, pyright clean. Note the branch is 40 commits behind master — say the word and I'll rebase before merge. |
Problem
A cloud job whose prover-side preprocessing (scene construction, storage/pointer analyses, CVL typechecking) blows up never produces rule results: it sits in
RUNNINGuntil the prover's global timeout (~2h). Autosetup's wait loop only polls the coarse job status, so it cannot distinguish this from a healthy long run and burns its entire 150-minute per-job budget — per job, several times per contract.Approach
The prover's
treeViewStatus.jsoncarries an emptyruleslist for the whole preprocessing phase; rules appear in it only once rule checking begins (verified against a live production job — note that the file itself is served early, so mere existence is not the signal). The newPreprocessingWatchdog(certora_autosetup/utils/preprocessing_watchdog.py) piggy-backs on the existing 10s status poll:RUNNINGobservation does;get_treeview_statusat a bounded cadence (default 90s), the first probe one interval intoRUNNING(a job that just started running is still fetching the jar and booting the JVM);rules⇒ dormant forever (zero further overhead);JobNotFoundError= "not served yet", not an error — verified live: the endpoint presigns the S3 object without an existence check and a missing object answers 404; 3 consecutive transport errors self-disable the watchdog, restoring the old behavior byte-for-byte, and any answered probe resets that count;PREPROCESSING_TIMEOUT.The result is classified as a new distinct
JobStatus.PREPROCESSING_TIMEOUT:on_job_problemskips conf workarounds for it (a preprocessing stall is not conf-fixable), so no identical job is resubmitted; the submission-cache entry is dropped so a deliberate future run submits fresh;sanity_summary.mdgets a dedicated row instead of crashing on the missing job report of a cancelled job (theget_job_reportcall is now also generally hardened).Env knobs:
AUTOSETUP_PREPROCESSING_BUDGET_SECONDS(default 1800;0disables),AUTOSETUP_PREPROCESSING_PROBE_SECONDS(90). No CLI-surface change, no new dependencies.Testing
tests/test_preprocessing_watchdog.py— 16 tests: state machine (queue-time free, first probe one interval in, probe cadence, dormancy, empty-rules-is-still-preprocessing, not-found vs error, self-disable only on consecutive errors, answered probe resets the count), wait-loop integration with mocked POU (stuck ⇒ cancel+classify; rules appear ⇒ never cancelled; zero-budget ⇒ watchdog off), retry suppression, serialization round-trip, reporter rows.CANCELEDon the cloud, andPREPROCESSING_TIMEOUTwas returned (vs ~90 min of doomed polling before).pyright: 0 errors on all changed files.🤖 Generated with Claude Code