Skip to content

[services.*]

A service is a long-running process. RunWisp restarts it when it exits or crashes, until you stop it. This page lists every key a [services.<name>] block accepts. Unset keys fall back to [defaults], then a built-in default.

type: string required

The text after services. is the service’s name: [services.api-worker] is a service called api-worker. The name is used in the CLI, API, Web UI, and log directory.

  • Letters, digits, and . _ - :, up to 100 characters.
  • Must be unique across [services.*] and [tasks.*].
  • Quote names that aren’t a bare TOML key ([services."api:v2"]).
[services.api-worker]
run = "/usr/local/bin/worker"

type: string required (unless compose_file is set)

The shell command to run. It runs as /bin/sh -e -c <command>, so the first failing command stops the run and its exit code becomes the run’s exit code (see Fail-fast). Set shell to bash for arrays or pipefail.

With compose_file, run is optional: it runs your command in the service’s container (see compose_mode).

[tasks.backup-db]
run = "pg_dump app | gzip > /backups/app.sql.gz"

Examples:

  • /usr/local/bin/backup.sh: a script on disk
  • pg_dump app | gzip > /backups/app.sql.gz: a shell pipeline
  • make -C /srv/app deploy: any command on your PATH

type: string

A short text shown in the Web UI and TUI. It doesn’t change how anything runs.

[tasks.backup-db]
description = "Nightly Postgres dump to S3"

type: string default: "Services"

The sidebar section it is listed under in the Web UI and TUI. Entries with the same group are listed together.

[tasks.weekly-report]
group = "Reports"

type: boolean default: true

Whether operators can start, stop, or restart it by hand, and pause its schedule. Set false to lock it to its schedule (or, on a service, to supervisor-only restarts): runwisp start/stop/restart/pause and the REST API return 403, the Web UI hides its controls, the station trigger is refused. Hooks still work.

[tasks.backup-db]
manual_trigger = false

On a service, this also gates the TUI’s s/r keys and runwisp start/stop/restart <name>. With manual_trigger = false, only a runwisp.toml edit plus runwisp reload, or a hook, changes its running state.

type: array of strings or tables

Tokens that let CI or a webhook control this unit over a hook without the dashboard password. Each token works for this unit only. Load them with ${VAR} or ${file:...} so they stay out of the file. Tokens are never shown in the API or Web UI.

[tasks.deploy]
run = "./deploy.sh"
hook_tokens = [
"${CI_DEPLOY_TOKEN}", # every action
{ token = "${ROLLBACK_TOKEN}", allow = ["stop"] }, # stop only
]

A plain string allows every action. A table with allow limits the token to the listed actions: run, start, stop, restart (services have no run).

Each token needs at least 32 characters and no whitespace. Generate one with openssl rand -hex 32. List more than one to rotate without downtime, or to give each caller its own. manual_trigger doesn’t apply to hooks.

type: integer default: 1 min: 1 max: 64

How many copies of the process run at once.

  • Each instance is its own run with an instance_index (0, 1, 2, …). The Web UI and TUI show them as name#1, name#2; a single instance is just name.
  • All instances share the config and the log view.
  • When one instance exits, only that one is restarted.
[services.api-worker]
instances = 3
run = "/usr/local/bin/worker"

For a command that should run once and exit, use a task.

type: enum default: "always"

When an instance is started again after it exits.

[services.worker]
restart = "on_failure"

Possible values:

  • always: restart after any exit
  • on_failure: restart only when failures counts the exit as a failure (and the reason is failed, timeout, crashed, log_overflow, start_failed, or unhealthy)
  • never: don’t restart; one exit is final

A manual Restart always starts the instance again, whatever this is set to.

type: duration default: 1s

The base wait before an automatic restart. restart_backoff decides how it grows. 0 is allowed and means no wait. A manual Restart never waits.

[services.flaky]
restart_delay = "2s"

type: duration default: 60s

How long an instance must stay up to count as healthy. Reaching it resets the restart backoff and the failed-start count used by restart_attempts.

With a health_check, it is the deadline for the first passing check instead, and 0 is rejected.

[services.api-worker]
healthy_after = "2m"

type: integer default: 3 min: 0 max: 100

How many failed starts in a row an instance may have before it goes FATAL and stops restarting. 0 gives up after the first failed start.

[services.flaky]
restart_attempts = 2

type: enum default: "exponential"

How restart_delay grows after each restart. The delay has an upper limit.

[services.flaky]
restart_backoff = "exponential"

Possible values: constant, linear, exponential (the same curves as retry_backoff).

A failed start is a failure exit before the instance became healthy: before it reached healthy_after, or, with a health_check, before its check first passed. A clean exit (success code) never counts.

  • After more than restart_attempts failed starts in a row, the instance goes FATAL: it stops restarting, a start_failed run is recorded, and a service.fatal notification goes to the in-app bell and your notifiers.
  • FATAL is kept in memory only. A daemon restart, or Restart in the UI or API, resets the count and tries again.
  • The start_failed runs stay in history.
[services.flaky]
run = "exit 1" # can never come up
healthy_after = "5s" # must stay up 5s to count as healthy
restart_attempts = 2 # give up after 2 failed starts in a row

type: table default: none

A command that decides when an instance is healthy, and stops it when it stops passing. Without it, uptime (healthy_after) is the only signal.

[services.api]
run = "./api-server"
[services.api.health_check]
run = "curl -fsS http://127.0.0.1:8080/healthz"
  • The check runs on its schedule next to every running instance. A non-zero exit fails it.
  • The first pass makes the instance healthy. Failures before that are ignored, but if nothing passes within healthy_after, the instance is stopped as unhealthy and counts as a failed start.
  • Once healthy, a failed check is retried. When the retries run out, the instance is stopped as unhealthy and restart decides what happens next.
  • Each failed check, the first pass, and the reason for a stop are written to the instance’s log. The check gets no runs of its own.

The table takes these task keys, with the same meaning:

Key Default
run required
cron "@every 10s"
timezone the daemon’s
timeout [defaults] timeout, else 5s; 0 is rejected
failures [defaults] failures, not the service’s
retry_attempts 2; 0 stops the instance on the first failed check
retry_delay, retry_backoff as on tasks
working_dir, shell, umask, env_base, user the service’s
env, secrets merged over the service’s, key by key
env_file, secrets_file none
compose_file, compose_service, compose_mode none

A check that times out is killed at once. Ticks that fall while a check is still running are skipped.

To check inside a running compose container, set compose_file and compose_service. Only compose_mode = "exec" is allowed, and the check then takes nothing from the service:

[services.api.health_check]
run = "curl -fsS http://127.0.0.1:8080/healthz"
compose_file = "./compose.yaml"
compose_service = "api"

type: boolean default: true

Whether the service starts when the daemon starts. With false, it stays stopped until you start it from the UI or API.

This is read from runwisp.toml at every boot, not saved: a manually started autostart = false service is stopped again after a daemon restart, and a manually stopped autostart = true service starts again.

[services.migration-runner]
autostart = false

type: integer default: 0

Start order at boot. Lower starts first; equal values start in name order. It doesn’t wait for a service to be ready. For that, use depends_on.

[services.db]
priority = 0
[services.api]
priority = 10

type: string[] default: none

Services that must be healthy (at least one instance up for its healthy_after, or passing its health_check) before this one starts at boot.

[services.db]
run = "postgres -D /var/lib/pg"
healthy_after = "3s"
[services.api]
run = "./api-server"
depends_on = ["db"] # start after db is healthy
  • If a dependency never becomes healthy, the service waits a limited time, then starts anyway and logs a warning.
  • It applies only to the automatic start at boot. Manual start and restart don’t wait.
  • At shutdown, dependent services stop first.
  • Names must be existing services (not tasks). Cycles are rejected at load.
  • There are no restart cascades or data passing between services. See Tasks vs Services.

Services have no on_overlap, max_concurrent, or max_queued. The number of running copies is set by instances.

type: duration default: none (no limit)

The longest an instance may run. After that it is stopped and recorded as timeout, then restart decides what happens next.

[services.report-server]
timeout = "24h"

type: string[] default: ["failed", "timeout", "crashed", "log_overflow", "start_failed", "unhealthy", "missed"]

Which outcomes count as a failure: they turn the run red, fire notify, and (with restart = "on_failure") cause a restart. Same syntax as on tasks; see What counts as a failure.

[services.web]
failures = ["-missed"]

type: duration default: 5s

How long a stopping run has to exit after stop_signal before RunWisp SIGKILLs it. Used on timeout, on_overlap = "kill", manual stop, and daemon shutdown. "0s" kills at once.

[tasks.queue-drain]
graceful_stop = "30s"
  • The signal goes to the whole process group. Every child process gets the same time, and anything still running after it is killed.
  • If it is longer than shutdown_timeout, the daemon warns at boot, because at shutdown the run may be killed early.

It applies to each instance, on manual stop, Restart, and daemon shutdown. At daemon shutdown, shutdown_timeout still limits it.

To stop cleanly, trap the stop signal in run and exit. An instance that exits after SIGTERM is recorded as stopped:

Terminal window
trap 'echo "shutting down"; exit 0' TERM INT
while true; do
# do work
done

type: enum default: "SIGTERM"

The first signal sent to stop a run. The SIG prefix is optional. Use SIGINT for tools that expect Ctrl-C, or SIGKILL to skip the grace time. Anything still running after graceful_stop is SIGKILLed.

[tasks.queue-drain]
stop_signal = "SIGINT"
graceful_stop = "10s"

Possible values: SIGTERM, SIGINT, SIGQUIT, SIGHUP, SIGKILL, SIGUSR1, SIGUSR2 (also without the SIG prefix).

type: integer min: 0 max: 1000000

Keep only the N most recent completed runs. Older runs and their logs are deleted by the hourly cleanup. 0 keeps no completed runs; unset means no count limit. With keep_for also set, whichever deletes more wins.

[tasks.backup-db]
keep_runs = 200

Every crash adds a run, so a service’s history can grow faster than a task’s. 200 is a good start for a service that restarts often.

type: duration

Delete runs older than this. With keep_runs also set, whichever deletes more wins.

[tasks.backup-db]
keep_for = "90d"

Examples: 48h (2 days), 14d, 2w, 90d.

type: size default: 100mb

The largest size of one run’s captured output on disk. Units: b, kb, mb, gb, tb. At the limit, log_on_full decides what happens. 0, negative, and invalid values are rejected at load.

[tasks.backup-db]
log_max_size = "500mb"

type: enum default: "drop_old"

What happens when a run’s log reaches log_max_size.

[tasks.backup-db]
log_max_size = "50mb"
log_on_full = "drop_old"

Possible values:

  • drop_old: keep the newest output, delete the oldest
  • drop_new: keep the first output, ignore the rest
  • kill: stop the run and record log_overflow

All instances get the same environment.

env and secrets reach the process the same way. The difference is who can see them:

  • env and env_file values are shown in the Web UI, TUI, and REST API to anyone logged in.
  • secrets and secrets_file values are never shown. At most, the secrets_file path is shown.
  • If a command prints a secret, RunWisp replaces the exact value with [redacted] in the stored and streamed output. This only matches the exact text: a secret that was changed (base64, URL-encoded) or split across lines can still appear.

Merge order at start (a later layer wins on the same key):

  1. The daemon’s own environment (see env_base).
  2. env_file, then env: first from [defaults], then from the task itself.
  3. secrets_file, then secrets: first from [defaults], then from the task itself.

So an inline value beats a file value, and a secret beats an env value with the same name.

Files:

  • Paths can be absolute, ~/-relative, or relative to the runwisp.toml directory.
  • Format is dotenv: one KEY=VALUE per line, # starts a comment, blank lines are skipped. Values are read as they are: no shell expansion and no ${...} substitution.

Limits, checked at load:

  • Keys must match ^[A-Za-z_][A-Za-z0-9_]*$.
  • Values can’t contain NUL bytes or be longer than 32 KiB.
  • At most 256 env + secrets entries in total.

type: map

Environment variables for the process. Their values are shown in the API and UI, so put passwords and tokens in secrets. To read from a file, use env_file.

[tasks.backup.env]
BACKUP_BUCKET = "s3://prod-backups"
DRY_RUN = "0"

type: map

Secret environment variables: passwords, tokens, API keys. They work like env, but their keys and values are never shown. To read from a file, use secrets_file.

[tasks.backup.secrets]
RESTIC_PASSWORD = "${file:~/.config/runwisp/restic.pass}"

type: string

A dotenv file merged below env (inline values win). Its values are shown in the API and UI, like env. Paths can be absolute, ~/-relative, or relative to the runwisp.toml directory. Lines are plain KEY=VALUE, with no shell expansion.

[tasks.backup]
env_file = "backup.env"

type: string

A dotenv file merged below secrets. Only the path is shown in the API and UI, never the contents. Same paths and format as env_file.

[tasks.backup]
secrets_file = "/etc/runwisp/backup-secrets.env"

All instances use the same directory, shell, environment base, user, and umask.

type: string

The directory the process runs in. Unset, it is the daemon’s directory.

  • A relative path is relative to the runwisp.toml directory.
  • ~ is the home directory of the user the run runs as.
  • The directory is checked when the run starts, not at load.
[tasks.deploy]
working_dir = "/srv/app"
run = "./bin/build"

type: string default: /bin/sh

The shell for run, as an absolute path. Use /bin/bash for arrays, [[ ]], or set -o pipefail. It is started as <shell> -e -c <script>, so fail-fast works for POSIX shells. RunWisp warns at boot when it can’t turn fail-fast on. Host processes only: it is an error on a compose unit.

[tasks.nightly-report]
shell = "/bin/bash"
run = "set -o pipefail; make report | tee out.log"

type: enum default: "inherit"

What the environment starts from, before env and secrets are added. Host processes only: it is an error on a compose unit.

[tasks.backup-db]
env_base = "clean"
run = "/usr/local/bin/nightly"

Possible values:

  • inherit: start from the daemon’s environment. Runs then depend on how the daemon was started.
  • clean: start with only PATH, SHELL, HOME, and USER/LOGNAME (like cron), so runs are repeatable. Set everything else in env.

type: string

Run as another OS user: user or user:group, by name or numeric id (the same form as chown and Compose). RunWisp looks up the account at start, sets HOME, USER, and LOGNAME, and applies its groups.

  • Works only when the daemon runs as root. Otherwise the run fails.
  • Empty keeps the daemon’s user.
  • Host processes only: it is an error on a compose unit.
[tasks.backup-db]
user = "reporter:reporter"

type: string

The octal file-creation mask for the run: "027" makes new files rw-r-----.

  • Write at least three digits ("022", not "22").
  • It applies only to this run’s process, not to other runs.
  • Empty keeps the daemon’s umask.
  • Host processes only: it is an error on a compose unit.
[tasks.backup-db]
umask = "027"

A service can run a service from an existing docker-compose.yml instead of a shell command:

[services.api]
compose_file = "./docker-compose.yml"
compose_service = "api"
notify = ["slack-prod"]

To import every service in a compose file, see [compose.*].

type: string

Path to a docker-compose file. The unit then runs the compose service named by compose_service instead of a shell command. Without run, each run is docker compose run --rm <service> with the service’s own command. With run, see compose_mode.

[tasks.nightly-backup]
cron = "0 3 * * *"
compose_file = "./docker-compose.yml"
compose_service = "backup"

type: string default: the service name

The service in compose_file to run. Requires compose_file.

[tasks.nightly-backup]
compose_file = "./docker-compose.yml"
compose_service = "backup"

type: enum default: "exec" when run is set, else "run"

How a compose_file unit runs.

[tasks.migrate]
compose_file = "./docker-compose.yml"
compose_service = "app"
compose_mode = "exec"
run = "rails db:migrate"

Possible values:

  • exec: run your run command inside the service’s running container (docker compose exec). Requires run.
  • run: start a new container for each run (docker compose run --rm). It runs your run command if set, else the service’s own command.

See Running a command in a service for the limits of exec.

type: string[]

Notifier ids to alert when a run counts as a failure (see failures). A missed run is a failure by default, so it alerts here too.

[tasks.backup-db]
notify = ["slack-ops"]
  • Each entry is a notifier declared in [notifiers.<id>], "inapp" (the bell), or "<id>:<target>" to reuse a notifier’s credentials with another target ("slack:#deploys", "tg:-1009998887", "mail:[email protected]"). See notifiers for which types accept a target.
  • [notify] global_notifiers (default ["inapp"]) is added to every list.
  • notify = [] is the same as leaving it out.
  • To alert on other outcomes (a success, for example), use a [[route]].

For a service, this fires when an instance ends with an outcome that failures counts as a failure, and when an instance goes FATAL.