Health checks
A health check is a curl on a schedule. The settings around it decide how
quickly you hear about an outage and how much noise a flaky service makes.
To check a service RunWisp runs itself, and restart it when the check keeps
failing, use its health_check
instead.
The task
Section titled “The task”[tasks.healthcheck-api]group = "Health checks"description = "External probe of the public API every 5 minutes"cron = "*/5 * * * *"on_overlap = "skip"timeout = "30s"retry_attempts = 2retry_delay = "30s"retry_backoff = "linear"keep_runs = 500notify = ["slack-ops"]
run = """curl --silent --show-error --fail-with-body \\ --max-time 25 --connect-timeout 5 \\ --user-agent 'runwisp-healthcheck/1.0' \\ https://api.example.com/health"""What each setting does
Section titled “What each setting does”cron = "*/5 * * * *": you learn about an outage within five minutes. Use* * * * *for tighter checks or*/15 * * * *for internal tools.on_overlap = "skip": a slow probe doesn’t get a second one stacked on it. The dropped firing is recorded asskipped, which isn’t a failure and doesn’t alert.timeout = "30s"withcurl --max-time 25: curl gives up first, so the log shows curl’s own error message. Thetimeoutis the backstop.retry_attempts = 2,retry_delay = "30s",retry_backoff = "linear": a failed probe is tried again after 30s, then after 60s more. A short blip heals on its own and the task goes green again within about 90 seconds. Each failed attempt still alerts, but repeats are grouped (see below). Seeretry_attempts.keep_runs = 500: five-minute checks make 288 runs a day, so this keeps well over a day of history. Seekeep_runs.notify: alertsslack-opson every failed run. During a long outage, repeats insidecoalesce_window(default1h) are grouped: the first alert goes out right away, then occasional check-ins with a count, instead of one message every five minutes.--fail-with-body: makes curl exit non-zero on an HTTP 4xx or 5xx and print the response body into the log. Without it, a500counts as success. It needs curl 7.76 or newer; on older curl use--fail.
Bell only, no messages
Section titled “Bell only, no messages”For a flaky service where you’d rather look than be messaged, drop the retries
and notify. Failures then go only to the bell in the Web UI and TUI, grouped
into one row with a count and a sparkline:
[tasks.healthcheck-api]cron = "*/5 * * * *"on_overlap = "skip"timeout = "30s"retry_attempts = 2retry_delay = "30s"retry_backoff = "linear"keep_runs = 500notify = ["slack-ops"]This relies on global_notifiers
keeping its default ["inapp"].
Probing internal services
Section titled “Probing internal services”Probe something that uses the service, not just its port. A database that accepts connections but refuses queries passes a port check. For Postgres:
PGPASSWORD="$HEALTHCHECK_PG_PWD" psql --host=db.internal --username=healthcheck \ --dbname=app_production --command='SELECT 1' --quiet --no-psqlrc > /dev/nullpg_isready is cheaper if you only need to know the server accepts connections.
When all you can check is an open port (a device you can’t log in to), nc is
fine. Just don’t treat it as more than that:
nc -zw5 host.internal 5432 || { echo "host.internal:5432 not reachable"; exit 1; }