Alerts and workload

Understand the shipped alert policies, evaluate and acknowledge alerts, create your own policy and read the server's latency figures.

Required permission: Operator (evaluate, acknowledge, pause); Configuration owner (create, activate policies)

Before you begin

An alert policy is a rule over one measurement. When the measurement crosses the threshold, the system opens an alert. Alerts appear in the app only: nobody is emailed or paged, so someone must look at Platform > Operations > Alerts or the Overview tile Open alerts.

The six shipped policies

CodeMeasurementThresholdSeverity
QUEUE-LAGOldest queued job (seconds)above 900, over 15 minutesCritical
P95p95 request latencyabove 2,000 msWarning
ERRORSServer error rateabove 5%Warning
DEAD-JOBSDead-letter jobsabove 0Warning
BACKUP-AGEHours since the last good backupabove 26Critical
DRILLSFailed restore drills in 30 daysabove 0Warning, answered by the release manager

Evaluate and acknowledge

  1. Open Platform > Operations > Alerts.
  2. Press Evaluate now to measure every active policy immediately. The platform.alert_scan job does the same every five minutes.
  3. Read the row: policy, severity, value, detail, up to five correlation ids, and when it opened.
  4. Press Acknowledge to say you are looking at it. The state changes to Acknowledged; the button disappears. An acknowledged alert still counts in Open alerts until it resolves.

Alerts resolve by themselves. When the measurement recovers, the next evaluation sets the alert to Resolved. There is no Resolve button. A second evaluation while the problem persists updates the same alert; it does not open a duplicate. An alert of a policy that is paused stays open until the policy is active and the measurement recovers.

Worked example: dead letters. A run ends as a dead letter. Evaluate now opens a DEAD-JOBS alert with the detail 'Dead-letter jobs is 1.0 (> 0.0) over 60 min.' and the correlation ids of the failing runs. You open the run, fix the cause and press Retry. After the next evaluation the count is 0 and the alert is Resolved.

Worked example: queue lag. The scheduler is off and the oldest queued run has waited 20 minutes. The measurement is 1,200 seconds, above 900, so a Critical alert opens. A lag of more than a week is reported as 604,800.

If you press Acknowledge or Evaluate now without the right platform role, the server refuses but the screen may show nothing. Check that you hold the Operator role.

Create your own policy

  1. A configuration owner opens Platform > Configuration > Alert policies and presses New.
  2. Enter a Code (capitals, digits, _ and -, starting with a letter) and Name.
  3. Choose the Measurement: p95 request latency (ms), Server error rate (%), Oldest queued job (seconds), Dead-letter jobs, Hours since the last good backup, or Failed restore drills (30 days).
  4. Choose When (>, >=, <, <=) and enter the Threshold.
  5. Set the Window (minutes), 1 to 1,440.
  6. Choose the Severity (Info, Warning, Critical) and Who answers it (one of the five platform roles).
  7. Write a Runbook of up to 400 characters: what to check first.
  8. Save, then press Activate.

A policy cannot be edited while active: 'Pause the policy before changing it.' An operator presses Pause, a configuration owner edits, then activates again. Paused policies are not evaluated.

Workload

Open Platform > Operations > Workload to see latency and errors of this server process. The summary line shows, since the server started, the requests, the p50, p95 and p99 latency and the server errors, then the p95 of the last 15 minutes, the queue lag and the dead letters. The table lists each route with requests, p50, p95, p99 and server errors (5xx). Press Measure again to refresh.

The figures are for this server process only and reset when it restarts.

Good to know

  • The Overview shows the tiles: Open alerts, Dead letters, Jobs queued and Last good backup. Each opens the matching list.
  • Messages and logs mask passwords, bearer tokens, API keys and long hex secrets as [redacted].