CronClawscheduled, verified, reported

RELIABILITY GUIDE · UPDATED 28 SEPTEMBER 2026

Cron job monitoring: catch missed runs and false success

Monitoring a cron job means checking that it ran on time and that the intended result exists. A process exit code or HTTP 200 is useful evidence, but it may not be the evidence your customer or operator actually needs.

Three questions every important job needs answered

  1. Did it start? Compare the expected schedule to actual starts, with a realistic grace period.
  2. Did it finish? Record an end event, exit status or response, duration, and error details.
  3. Did it produce the intended result? Check a fresh backup artifact, reconciled total, completed invoice batch, or another independent business fact.

A heartbeat sent only after success detects missing or late completions. A separate read-only verifier can catch the harder case: the job says it succeeded but the result is stale or incomplete. Keep evidence specific to the task.

Choose the right signal

SignalDetectsLimit
Scheduler or system logAttempted start and local errorsA stopped scheduler may produce no new log
Completion heartbeatMissing or late successful runsA heartbeat sent too early can be falsely green
Exit code or HTTP responseExplicit process or request failureSuccess may hide incomplete business work
Independent outcome checkMissing, stale, or wrong resultNeeds a safe proof source for that workflow

The simplest reliable setup often combines a completion heartbeat with one result-specific check. For a backup, check file existence and freshness; when feasible, perform a controlled restore test. For inventory sync, compare the expected batch or source cursor to what was committed.

Cron job monitoring checklist

  1. Write down the schedule, time zone, expected duration, and acceptable delay. Include what should happen during daylight saving changes.
  2. Name the owner and the consequence of a missed run. Send alerts to someone who can act.
  3. Emit a heartbeat only after the task reaches its success condition. Protect the heartbeat URL as a secret.
  4. Check an independent result for valuable jobs. Make the proof read-only and avoid returning customer data or credentials.
  5. Set an alert grace period that accounts for normal run time and network delay without hiding a real miss.
  6. Test one controlled failure and recovery. Record when the alert arrived and how the incident was resolved.
  7. Review noisy alerts and false-green cases regularly. Change thresholds based on evidence, not on alert fatigue.

Retries, duplicate work, and security

Retries can recover temporary failures, but a second attempt may repeat a payment, email, or file operation. Make work idempotent where possible and distinguish “request retried” from “business result verified.” Use authentication or signatures on trigger endpoints, rotate exposed secrets, limit what a job endpoint can do, and avoid placing secret tokens in public links or logs.

CronClaw’s signed-job model sends an authenticated HTTPS request to your application. Its heartbeat model can watch a scheduler you already use. Outcome checks add a result-level signal for workflows such as backups, syncs, and reconciliation. Compare these choices in the scheduling guide and see the sample evidence report.

Start with the job that matters most.

Define the proof, route the alert, and rehearse a failure before depending on the monitor.

Explore the verified-job pilot