Run and monitor scheduled jobs
Define a background job, test it, enable it, and handle runs that retry, fail into the dead letter state, or need cancelling.
Before you begin
Background work, such as metrics, alert scans, package scans and software bills of materials, runs through the shared job engine. Open Platform > Operations > Scheduled jobs for the definitions and Job runs for every run the engine accepted.
The first visit creates seven platform jobs:
| Job | Schedule | State |
|---|---|---|
platform.metrics_snapshot | every 5 minutes | Enabled |
platform.alert_scan | every 5 minutes | Enabled |
platform.package_scan | daily 03:00 | Enabled |
platform.sbom | weekly, Sunday 03:30 | Tested |
platform.prune | daily 04:00 | Enabled |
platform.restore_drill | on demand | Tested |
platform.upgrade_rehearsal | on demand | Tested |
The two drill jobs need a drill to run against. Running them by hand from the Jobs screen ends in the dead letter state with 'The run does not name a drill.' Start them from Restore drills or from a release.
Create a job
- A configuration owner presses New on Scheduled jobs.
- Enter the Code: lower-case words joined by dots, 2 to 6 parts (
sales.nightly_digest). It cannot change. - Enter the Name and choose the Handler from the registered list.
- Choose the Schedule: On demand, Every N minutes, Daily at a time, or Weekly on a day. Set Every (minutes) (1 to 10,080), At (HH:MM), Weekday and Time zone (default
Asia/Dubai). - Set Timeout (s) (5 to 21,600), Lease (s) (10 to 21,600), Attempts (1 to 10), Backoff (s) (1 to 3,600) and Concurrency (1 to 4).
- Enter a Payload as a flat JSON object of up to 4 KB with plain values only.
- Save. The job is a Draft.
Test and enable
- An operator presses Run now on the draft. It runs as a test. When it succeeds, the job becomes Tested.
- Press Run the scheduler now under Job runs to make the engine pick it up.
- Press Enable. A draft job cannot be enabled directly: 'A draft job cannot be enabled.' The job becomes Enabled with Next due set.
Pause stops new runs and clears Next due. A configuration owner presses Back to draft to edit.
How schedules are counted
- Interval jobs are aligned to 1 January 2026 00:00 UTC. A 60-minute job enabled at 10:17 Dubai (06:17 UTC) is next due at 11:00 Dubai (07:00 UTC), and its run key is
2026-10-02T07:00Z. - Daily and weekly jobs use the wall time in the job's zone. A daily 02:00 Asia/Dubai job enabled on 2 October after 02:00 is next due on 3 October at 02:00 Dubai, which is 22:00 UTC on 2 October.
- After downtime only the latest due slot is accepted. A 60-minute job that missed three hours runs once, not three times.
- The same run key is one run. Pressing Run now twice with the key
test-key-1returns the same run.
Retries and dead letters
A failing run waits and tries again. The wait is the backoff times 2 to the power of (attempt minus 1), capped at 6 hours.
Worked example. Attempts 3, backoff 60 seconds. Attempt 1 fails and waits 60 seconds. Attempt 2 fails and waits 120 seconds. Attempt 3 fails and the run becomes a dead letter: an event platform.job_failed.v1 is raised and the error shows with secrets masked. With a backoff of 3,600 seconds, the fourth wait would be 28,800 seconds but is capped at 21,600.
Open the run for the stage rail (Queued, Running, Succeeded, Dead letter or Cancelled), the worker, lease, next retry and recorded side effects.
- Retry a dead or cancelled run: it returns to Queued with the same run key, and side effects already recorded are not repeated. 'Only a dead or cancelled run is retried.'
- Cancel a queued run (cancelled at once) or a running one (it stops at its next check). 'This run has already finished.' for a finished run.
Check job health
Platform > Reporting > Job health shows each job's accepted runs split into queued, retrying, running, succeeded, dead and cancelled. Reconciles is yes when the parts add up to the accepted total.
Good to know
- Errors from the Run the scheduler now button are not always shown on screen; if nothing seems to happen, check Job runs.
- The platform system jobs can be edited by a configuration owner. Avoid changing their handlers.