Skip to content

Failures, retries & timeouts

A task can try again when it fails, and can be stopped when it runs too long. Both are set per task:

[tasks.publish-feed]
cron = "*/15 * * * *"
run = "/usr/local/bin/publish.sh"
retry_attempts = 3
retry_delay = "30s"
retry_backoff = "exponential"
timeout = "10m"

By default, a run fails when the command exits non-zero, times out, is stopped by log_on_full = "kill", or was still running when the daemon crashed. A scheduled run missed while the daemon was down counts too. A failed run turns red in the UI, sends notifications, and can be retried.

You can change this per task with failures, for example to treat only some exit codes as failures. The full rules are in What counts as a failure.

retry_attempts is the number of extra tries after the first one fails. It’s 0 by default, so nothing retries until you set it. The retry keys are per task; they can’t go in [defaults].

A run is retried only when it failed and the task’s failures list counts it as a failure. With failures = ["1-23"], exit code 42 isn’t a failure, so it isn’t retried either.

Some endings never retry:

  • stopped: someone stopped the run, or on_overlap = "kill" replaced it. This also ends the retry chain.
  • skipped: the previous run is still going.
  • missed: nothing ran, so there is nothing to retry.

Adding stopped or missed to failures doesn’t change this.

A retry uses exactly the same parameter values as the run it retries.

retry_delay is the wait before the first retry. retry_backoff decides whether later waits stay the same or grow:

[tasks.flaky-fetch]
retry_attempts = 4
retry_delay = "10s"
retry_backoff = "exponential" # waits 10s, 20s, 40s, 80s

No wait is ever longer than 5 minutes, whatever the backoff.

Every attempt is its own run, with its own exit code and log. The history numbers them with retry_attempt: 0 for the first try, 1 for the first retry, and so on. If an attempt fails and a later one succeeds, you can still read what the failed one printed.

Each failed attempt is a failure event of its own. Repeats of the same failure are grouped by coalesce_window, so a retry chain usually sends one message, not one per attempt.

timeout is the longest one attempt may run:

[tasks.heavy-job]
cron = "0 3 * * *"
run = "/usr/local/bin/heavy-job.sh"
timeout = "30m"

When the time is up, RunWisp sends the task’s stop_signal (SIGTERM by default) to the whole process group, waits graceful_stop (5s by default), then kills whatever is left. The run is recorded as timeout. Manual stops, on_overlap = "kill", and daemon shutdown stop runs the same way.

  • The limit is per attempt. Each retry gets the full time again, and the wait between attempts doesn’t count.
  • Without a timeout on the task or in [defaults], runs have no time limit. Set timeout = "0s" on a task to turn off a [defaults] timeout.

Services don’t retry; they restart. By default a service is started again after every exit, with a growing delay between restarts. If it keeps failing before it has been up for healthy_after, RunWisp stops trying and marks it FATAL. The keys are restart, restart_delay, restart_backoff, and restart_attempts; see When a service can’t start.

If you want a limited number of tries, use a task with retry_attempts. If you want a process kept alive indefinitely, use a service.