Skip to content

Health checks

A health check is a curl on a schedule. The settings around it decide how quickly you hear about an outage and how much noise a flaky service makes.

To check a service RunWisp runs itself, and restart it when the check keeps failing, use its health_check instead.

[tasks.healthcheck-api]
group = "Health checks"
description = "External probe of the public API every 5 minutes"
cron = "*/5 * * * *"
on_overlap = "skip"
timeout = "30s"
retry_attempts = 2
retry_delay = "30s"
retry_backoff = "linear"
keep_runs = 500
notify = ["slack-ops"]
run = """
curl --silent --show-error --fail-with-body \\
--max-time 25 --connect-timeout 5 \\
--user-agent 'runwisp-healthcheck/1.0' \\
https://api.example.com/health
"""
  • cron = "*/5 * * * *": you learn about an outage within five minutes. Use * * * * * for tighter checks or */15 * * * * for internal tools.
  • on_overlap = "skip": a slow probe doesn’t get a second one stacked on it. The dropped firing is recorded as skipped, which isn’t a failure and doesn’t alert.
  • timeout = "30s" with curl --max-time 25: curl gives up first, so the log shows curl’s own error message. The timeout is the backstop.
  • retry_attempts = 2, retry_delay = "30s", retry_backoff = "linear": a failed probe is tried again after 30s, then after 60s more. A short blip heals on its own and the task goes green again within about 90 seconds. Each failed attempt still alerts, but repeats are grouped (see below). See retry_attempts.
  • keep_runs = 500: five-minute checks make 288 runs a day, so this keeps well over a day of history. See keep_runs.
  • notify: alerts slack-ops on every failed run. During a long outage, repeats inside coalesce_window (default 1h) are grouped: the first alert goes out right away, then occasional check-ins with a count, instead of one message every five minutes.
  • --fail-with-body: makes curl exit non-zero on an HTTP 4xx or 5xx and print the response body into the log. Without it, a 500 counts as success. It needs curl 7.76 or newer; on older curl use --fail.

For a flaky service where you’d rather look than be messaged, drop the retries and notify. Failures then go only to the bell in the Web UI and TUI, grouped into one row with a count and a sparkline:

[tasks.healthcheck-api]
cron = "*/5 * * * *"
on_overlap = "skip"
timeout = "30s"
retry_attempts = 2
retry_delay = "30s"
retry_backoff = "linear"
keep_runs = 500
notify = ["slack-ops"]

This relies on global_notifiers keeping its default ["inapp"].

Probe something that uses the service, not just its port. A database that accepts connections but refuses queries passes a port check. For Postgres:

Terminal window
PGPASSWORD="$HEALTHCHECK_PG_PWD" psql --host=db.internal --username=healthcheck \
--dbname=app_production --command='SELECT 1' --quiet --no-psqlrc > /dev/null

pg_isready is cheaper if you only need to know the server accepts connections.

When all you can check is an open port (a device you can’t log in to), nc is fine. Just don’t treat it as more than that:

Terminal window
nc -zw5 host.internal 5432 || { echo "host.internal:5432 not reachable"; exit 1; }