Push-to-deploy for one important Linux box. A small Python library + CLI that turns a git ref with a Dockerfile into a running, routed, TLS'd, health-checked, resource-capped service behind the Caddy reverse proxy you already run — with automatic rollback on a failed health check, and without restarting the proxy or anything else on the box.
deploy-this apply hello
hello: resolving https://git.example.internal/platform/hello.git @ main
hello: starting release r7 (3f9c2a1b0d4e, requested by alice)
hello: building deploy-this/hello:r7
hello: starting deploy-this-hello-r7.service on 127.0.0.1:41003
hello r7 health gate (127.0.0.1:41003/healthz): healthy after 3 probe(s) (HTTP 200 OK b'{"ok": true}')
hello: cutting route hello.apps.example.internal over to 127.0.0.1:41003
hello: fragment written; live route replaced (srv0/routes/0)
hello through-proxy check (https://hello.apps.example.internal/healthz): healthy after 1 probe(s)
hello: retiring r6 (deploy-this-hello-r6.service)
hello: release r7 deployed (3f9c2a1b0d4e) in 41s
It is built for the "one box shared with services that must not be disturbed" case — no Kubernetes, no Swarm, no second proxy, no database, no new privileged daemon. MIT licensed.
- One engine, three entry points: the CLI, a pull-based queue daemon (Azure Service Bus first; webhook and git-poll adapters included), and an importable library.
- Zero-downtime, health-gated deploys: the new release starts next to the old one, must
pass an HTTP health gate, is cut over with a single surgical change to Caddy's running
config (no reload, no restart), is verified through the proxy, and only then is the old
release retired. Any failure leaves the old release serving.
rollbackre-runs the same flow with a previous image. - Your Caddy stays yours: deploy-this owns one directory of Caddyfile fragments and one marker-tagged route per app, nothing else. Operator snippets (wildcard TLS, forward-auth) are used as-is. Reloads and restarts by the operator keep working. See docs/design.md for the research and the evidence.
- Mandatory caps and anti-footgun isolation: memory/CPU/pids via
docker runflags,--cap-drop ALL, no-new-privileges, loopback-only ports, a private bridge per app, and default-deny egress with a per-app allowlist. - Preview instances: one throwaway copy per pull request with a TTL, at its own subdomain.
- Key Vault secret references:
secrets: {DATABASE_URL: keyvault://vault/name}resolved on the box at deploy time (alsofile://andenv://); rotation is picked up by the next apply. - Scheduled functions:
function.scheduleturns a function app into a systemd timer job. - Staging slot and swap:
apply --stage→ look →swap; the old release stays running in staging so swapping back is instant. Orstrategy: canaryfor weighted gradual cutover. - Build where it suits you: on the box (in a resource-capped builder) or in CI, deploying registry images by digest so the shared box never builds at all.
- Durable records: who deployed what, when, from which ref, with which image, and how it went — as JSON files (or SQLite) that survive reboots and disk swaps.
- Legible and undoable: every unit, fragment, env file and record is a plain file you can
read;
deploy-this doctor,status,routes showtell you what the box thinks. - One privileged component: a single stdlib-only root helper with exactly four verbs
(
install-unit,remove-unit,apply-route,remove-route). - Self-repairing: a one-minute
maintaintimer auto-heals unhealthy releases (health-check- restart, rate-limited) and reconciles routes and units with the recorded desired state.
- Signed envelopes and notifications: HMAC-signed deploy requests (reject the rest), and a webhook line per deploy event for Slack/Teams.
Runtime dependency: PyYAML. Optional extras: [servicebus] (azure-servicebus + azure-identity)
for the queue daemon, [keyvault] (azure-keyvault-secrets + azure-identity) for Key Vault references.
It is easy to mix up what goes where, so:
your service's repo the box (operator-owned) this project
─────────────────── ──────────────────────── ────────────
code + Dockerfile deploy-this + its daemon the tool both
azure-pipelines.yml ──{app, ref}──▶ queue ──▶ /etc/deploy-this/apps/hello.yml sides install
▲ /etc/deploy-this/config.yml
└────────── git fetch by sha ────── /etc/caddy/… (+ one import line)
- The deployable repo (examples/app-repo) contains the app, its
Dockerfile, and a pipeline whose last step enqueues{"app": "hello", "ref": "<sha>"}. It never connects to the box and never pushes an image. - The box (examples/operator) holds the operator config, one manifest per app (where the repo is, which host/auth policy, which caps), the Caddyfile snippets, the root helper and the daemon that consumes the queue and does the clone/build/deploy.
- deploy-this (this repo) is the library/CLI installed on the box. Its
examples/app-repodirectory is only a stand-in for a separate repository.
Assumes Ubuntu 22.04/24.04 with sudo. Replace *.apps.example.internal with a name your
wildcard certificate covers.
1. Prerequisites: Docker, git, Caddy.
sudo apt-get update && sudo apt-get install -y docker.io docker-buildx git python3-venv curl
# Caddy (official repo): https://caddyserver.com/docs/install#debian-ubuntu-raspbian
sudo apt-get install -y debian-keyring debian-archive-keyring apt-transport-https
curl -1sLf 'https://dl.cloudsmith.io/public/caddy/stable/gpg.key' | sudo gpg --dearmor -o /usr/share/keyrings/caddy-stable-archive-keyring.gpg
curl -1sLf 'https://dl.cloudsmith.io/public/caddy/stable/debian.deb.txt' | sudo tee /etc/apt/sources.list.d/caddy-stable.list
sudo apt-get update && sudo apt-get install -y caddy2. A deploy user that can use Docker and read the journal, but is otherwise unprivileged.
sudo useradd --system --create-home --shell /bin/bash --groups docker,systemd-journal deploy3. Install deploy-this from git into a venv, plus the root helper and its sudoers rule.
deploy-this is installed straight from its git repository (it is not on PyPI). Pin a tag:
sudo python3 -m venv /opt/deploy-this
sudo /opt/deploy-this/bin/pip install 'deploy-this[servicebus] @ git+https://github.com/jackhumbert/deploy-this-py@v0.1.0'
sudo install -m 0755 "$(/opt/deploy-this/bin/deploy-this helper-path)" /usr/local/bin/deploy-this-helper
printf 'Defaults!/usr/local/bin/deploy-this-helper !requiretty\ndeploy ALL=(root) NOPASSWD: /usr/local/bin/deploy-this-helper\n' \
| sudo tee /etc/sudoers.d/deploy-this >/dev/null && sudo chmod 0440 /etc/sudoers.d/deploy-this(examples/operator/sudoers in this repo is the same two lines.)
4. Config and directories.
sudo mkdir -p /etc/deploy-this/apps /var/lib/deploy-this /etc/caddy/deploy-this.d
sudo curl -sSLo /etc/deploy-this/config.yml https://raw.githubusercontent.com/jackhumbert/deploy-this-py/main/examples/operator/config.yml
sudo chown -R deploy:deploy /var/lib/deploy-this /etc/deploy-this/apps && sudo chmod 0750 /var/lib/deploy-thisThe defaults in examples/operator/config.yml match this layout; edit
caddy.tls_snippet / caddy.auth_policies to your snippet names.
5. Three lines in your Caddyfile (snippets before the import; see
examples/operator/Caddyfile for a complete one). The TLS snippet
assumes your wildcard certificate is at /etc/caddy/certs/; for a first try on a throwaway VM,
mint a self-signed one:
sudo mkdir -p /etc/caddy/certs && sudo openssl req -x509 -newkey rsa:2048 -nodes -days 3650 \
-keyout /etc/caddy/certs/wildcard.key -out /etc/caddy/certs/wildcard.pem \
-subj "/CN=*.apps.example.internal" -addext "subjectAltName=DNS:*.apps.example.internal"
sudo chown caddy:caddy /etc/caddy/certs/* && sudo chmod 0640 /etc/caddy/certs/wildcard.key(wildcard_tls) {
tls /etc/caddy/certs/wildcard.pem /etc/caddy/certs/wildcard.key
}
(auth_internal) {
forward_auth 127.0.0.1:9091 {
uri /api/verify
copy_headers X-User
}
}
import /etc/caddy/deploy-this.d/*.caddysudo systemctl reload caddy # one last reload; deploy-this never needs another6. Check the box.
sudo -u deploy /opt/deploy-this/bin/deploy-this doctor7. Describe an app and deploy it.
sudo -u deploy tee /etc/deploy-this/apps/hello.yml <<'YML'
runtime: web
source: {repo: https://github.com/jackhumbert/deploy-this-py, ref: main, subdir: examples/app-repo}
port: 8080
route: {host: hello.apps.example.internal, auth: public} # `internal` once auth_internal points at your real forward-auth service
health: {path: /healthz, timeout: 60s}
limits: {memory: 256m, cpus: 0.5, pids: 64}
egress: {allow: []}
YML
sudo -u deploy /opt/deploy-this/bin/deploy-this apply hello
curl -sk --resolve hello.apps.example.internal:443:127.0.0.1 https://hello.apps.example.internal/8. Operate it.
deploy-this status hello # release, unit, container, health, route (live + desired), history
deploy-this list # every app at a glance
deploy-this logs hello -f # journalctl -u deploy-this-hello-rN.service -f
deploy-this rollback hello # previous successful image, same health-gated flow
deploy-this remove hello # route, unit, network, images gone; data keptThen install the two timers (self-repair every minute, disk GC daily):
for u in deploy-this-maintain.service deploy-this-maintain.timer deploy-this-gc.service deploy-this-gc.timer; do
sudo curl -sSLo /etc/systemd/system/$u https://raw.githubusercontent.com/jackhumbert/deploy-this-py/main/examples/operator/$u
done
sudo systemctl daemon-reload && sudo systemctl enable --now deploy-this-maintain.timer deploy-this-gc.timerPrefer a script? scripts/ci-setup.sh provisions a fresh VM (it writes its own Caddyfile
with a self-signed wildcard cert, so do not run it on a box that already has one), and
scripts/e2e.sh then runs the full deploy → fail → rollback → remove proof while checking that
Caddy, Docker and a canary service never restart. That is exactly what CI runs.
There is no PyPI release; the box installs from the repository with a pinned tag, and upgrades
by installing a newer tag (a changed URL is enough for pip to rebuild — no --force-reinstall
needed):
# GitHub (public)
sudo /opt/deploy-this/bin/pip install 'deploy-this[servicebus] @ git+https://github.com/jackhumbert/deploy-this-py@v0.1.0'
# Azure Repos over SSH (deploy key / the deploy user's key)
sudo /opt/deploy-this/bin/pip install 'deploy-this[servicebus] @ git+ssh://git@ssh.dev.azure.com/v3/<org>/<project>/deploy-this@v0.1.0'
# Azure Repos over HTTPS with a PAT kept out of the command line
sudo git config --system credential.helper store # then put https://<user>:<PAT>@dev.azure.com in /root/.git-credentials (0600)
sudo /opt/deploy-this/bin/pip install 'deploy-this[servicebus] @ git+https://dev.azure.com/<org>/<project>/_git/deploy-this@v0.1.0'After upgrading, re-copy the root helper (it is a plain file, not a console script) and restart the daemon:
sudo install -m 0755 "$(/opt/deploy-this/bin/deploy-this helper-path)" /usr/local/bin/deploy-this-helper
sudo systemctl restart deploy-this-daemon
sudo -u deploy /opt/deploy-this/bin/deploy-this doctor # reports a helper/library version mismatch if you forgotTags follow semver and the CHANGELOG; deploy-this version prints what is installed.
Azure App Service's deployment slots, for one box:
deploy-this apply hello --stage # builds + health-gates r8, routes it at https://hello-staging.apps.example.internal
# look at it, run your smoke tests against it...
deploy-this swap hello # r8 takes the production route; r7 keeps running, parked in staging
deploy-this swap hello # changed your mind: instant swap back, no rebuild, nothing restarted
deploy-this unstage hello # done with it: stop whatever is in the staging slotA swap is one route change (verified through the proxy before the pointers move); nothing is
stopped, so both releases keep their ports and "swap again" is the zero-cost rollback. The
staging host defaults to <first label>-staging.<rest> (set route.staging_host to change it)
and uses the same auth policy and wildcard certificate. Envelopes can drive the same flow:
"action": "stage" then, after an approval, "action": "swap".
For gradual exposure without a second host, strategy: canary sends canary.weight percent of
production traffic to the new release (Caddy weighted_round_robin) for canary.bake, probing
it the whole time; a failing probe reverts the route and removes the release, a passing bake cuts
over fully and retires the old one.
With previews: {enabled: true, ttl: 3d, max: 5} in the manifest, deploy-this apply hello --instance pr-123 (or an envelope with "instance": "pr-123") deploys a throwaway copy at
https://hello-pr-123.apps.example.internal — same build, caps, auth and egress, its own unit,
port and records — and maintain removes it ttl after its last deploy. Every push to the PR
redeploys the same instance and extends the TTL; deploy-this instances lists them;
remove hello --instance pr-123 ends one early.
examples/app-repo/azure-pipelines-preview.yml
is the PR-triggered pipeline that drives it.
A function app with function.schedule (a systemd OnCalendar expression) gets a oneshot
service + persistent timer per release that runs deploy-this run <app> — the WebJobs /
timer-trigger equivalent. The timer survives reboots (missed runs happen at boot), every run is
the usual throwaway hardened container with the app's caps and timeout, and its output is in
journalctl -u deploy-this-<app>-rN.service. status shows the next and last run.
Two modes, per app or per deploy:
| mode | who builds | envelope / CLI | manifest |
|---|---|---|---|
| on the box (default) | the box clones the commit and runs docker build |
ref: <sha> / apply hello --ref |
source: |
| in CI | the pipeline builds, pushes to a registry and pins the digest; the box only pulls | image: myregistry.azurecr.io/hello@sha256:… / apply hello --image … |
image: {repository: myregistry.azurecr.io/hello} |
Building in CI is the production recommendation for a shared box: build CPU, memory and disk
stay on the agent, the artifact is immutable and scannable, and the box's job shrinks to
pull → health gate → cutover. Envelopes may only reference images under the manifest's
image.repository; tag references are resolved to a digest at deploy time and recorded.
If you do build on the box, give builds a cap: configure docker.builder (a dedicated
BuildKit builder created with docker buildx create --driver docker-container and memory/CPU
limits, plus a bounded build cache) — otherwise docker build runs uncapped inside the shared
daemon, which is exactly what the rest of deploy-this is designed to prevent. deploy-this doctor
warns when neither is set up.
Two things keep running between deploys, both from timers you install once:
- Auto-heal (
deploy-this heal, part ofmaintain): every current release is probed on its health path; afterhealth.heal_afterconsecutive failures (default 3) — or if its unit is not active — the unit is restarted, at mostheal.max_restarts_per_hourtimes, with a notification each time and a final "giving up" notification when the limit is hit. The Azure App Service health-check/auto-heal idea, for one box. - Reconcile (
deploy-this reconcile, part ofmaintain):routes syncplus units — a current release whose unit file vanished (e.g./etcrestored from an older backup) is reinstalled from the saved unit text;deploy-this-*units that belong to no current or in-progress release (a retire that failed) are removed. Apps with a deploy in progress are skipped.
Install examples/operator/deploy-this-maintain.timer
(every minute) and you have a self-repairing box; deploy-this status shows heal history.
Disk is reclaimed by deploy-this gc (old images beyond keep_releases, build cache beyond
build_cache_keep, stale build directories, records beyond keep_records); install
examples/operator/deploy-this-gc.timer to run it daily.
resolve ref ─▶ fingerprint (no-op if unchanged) ─▶ clone ─▶ docker build
─▶ route preflight ─▶ allocate loopback port ─▶ install unit (new release starts beside the old)
─▶ health gate: GET 127.0.0.1:<port>/healthz ✗ → remove new unit; old keeps serving
─▶ cutover: write fragment + splice one route into Caddy's running config (no reload)
─▶ verify through the proxy: GET https://<host>/healthz via 127.0.0.1:443 ✗ → revert route, remove new unit
─▶ retire old release (docker stop -t <stop_grace>) ─▶ record ─▶ prune old images
Each release is a systemd unit (deploy-this-<app>-rN.service) running
docker run --rm … -p 127.0.0.1:<port>:<app port> --memory … --cpus … --pids-limit … with a
per-release egress chain in DOCKER-USER. Units, fragments and env files are plain files; read
them.
- You add one
import /etc/caddy/deploy-this.d/*.caddyline after your snippets. - deploy-this writes one fragment per app (
<host> { import <tls>; import <auth>; vars deploy_this_app <app>; reverse_proxy 127.0.0.1:<port> }) through the root helper, which validates the whole Caddyfile and reverts on error. - For the live change it runs
caddy adapton your Caddyfile (read-only), takes the one route carrying thedeploy_this_appmarker, andPATCHes/PUTs just that route via the admin API withIf-Match. Nothing else in the running config is touched; Caddy swaps configs atomically and keeps listeners open, so no request is refused and in-flight requests finish. - Your own
systemctl reload caddykeeps working and keeps the apps routed (the fragment is in your Caddyfile).deploy-this routes syncreconciles live state with desired state whenever you want. - Operators with WebSocket services: set
stream_close_delayon thosereverse_proxyblocks — every Caddy config change (yours or ours) closes proxied WebSockets otherwise. See docs/design.md §3 for why and for the evidence.
The alternative backend route_backend: caddyfile writes the same fragment and runs a graceful
caddy reload instead of the splice.
One YAML file per app in apps_dir. Full reference: docs/manifest.md.
runtime: web # web | function | process
source: {repo: ..., ref: main, subdir: .} # and/or image: {repository: myregistry.azurecr.io/hello}
build: {dockerfile: Dockerfile, args: {}, target: null}
port: 8080 # what the container listens on (also exported as $PORT)
route: {host: hello.apps.example.internal, auth: internal, extra: [], staging_host: null}
health: {path: /healthz, timeout: 60s, interval: 2s, expect: [200]}
limits: {memory: 256m, cpus: 0.5, pids: 64} # required
egress: {allow: ["10.0.0.0/8:5432", "api.example.com:443"]} # default-deny; or `egress: unrestricted`
env: {LOG_LEVEL: info}
volumes: ["data:/data"] # persisted under <state_dir>/data/<app>/
security: {read_only: false, cap_drop: [ALL], cap_add: [], no_new_privileges: true}
strategy: blue-green # or canary (+ canary: {weight: 10, bake: 5m}), or in-place (+ fixed_port)
stop_grace: 30sSecret values never go in the manifest. Two ways in:
secrets: # references, resolved on the box at deploy time
DATABASE_URL: keyvault://kv-example/hello-database-url # Azure Key Vault (pip install 'deploy-this[keyvault]', managed identity)
API_KEY: keyvault://kv-example/hello-api-key/8f1c0a… # pinned version
TLS_KEY: file:///etc/deploy-this/secrets/hello-tls.key # root-managed file readable by the deploy user
SENTRY_DSN: env://HELLO_SENTRY_DSN # from the daemon's environmentor deploy-this secrets set hello DATABASE_URL=..., which stores the value in
<state_dir>/secrets/hello.env (0600) and overrides a reference of the same name. Either way
the values are injected with --env-file, never logged, and their hash is part of the release
fingerprint — so a rotated Key Vault secret is picked up by the next apply (secrets check
verifies references without printing anything secret).
| Command | What it does |
|---|---|
apply <app> [--ref REF | --image REF] [--force] |
build (or pull a CI-built image) + health-gated zero-downtime deploy; no-op if nothing changed |
rollback <app> [--to rN] |
redeploy a previous successful image, same gate |
apply <app> --instance pr-123 [--ttl 3d] / instances / remove <app> --instance pr-123 |
preview instances per pull request, expired by maintain |
apply <app> --stage / swap <app> / unstage <app> |
deploy into the staging slot; promote it (old release parks in staging — swap again to roll back); clear the slot |
remove <app> [--keep-images] |
route, unit, network, images removed; persistent data kept |
status <app> [--json] / list [--json] |
what is deployed, unit/container/health/route state, history |
logs <app> [-f] [-n N] |
journalctl -u of the current release |
run <app> [-- args] |
run a function app once in a throwaway hardened container |
| `secrets set | unset |
| `routes sync | list |
validate [app] |
check manifests |
maintain [--dry-run] / heal [app] / reconcile [--dry-run] |
auto-heal unhealthy releases; reinstall missing units, remove orphans, sync routes (run from a timer) |
gc [--dry-run] |
reclaim disk: old images, build cache, stale build dirs, old records |
| `daemon [--trigger servicebus | webhook |
doctor |
check config, docker, caddy, helper, snippets, ports |
helper-path / version |
where the bundled root helper is (to install it); installed version |
Global flags: -c/--config PATH (default $DEPLOY_THIS_CONFIG or /etc/deploy-this/config.yml),
-v for debug logging on stderr; apply/rollback/remove/swap/unstage accept
--requested-by for the records.
Exit codes: 1 generic failure, 2 config/manifest/secrets, 3 build, 4 helper, 5 route, 6 health gate, 7 not deployed, 8 locked, 9 command, 130 interrupted.
The box accepts no inbound connections, so the app repo's pipeline enqueues a small envelope
and the daemon on the box does the work. A complete pipeline for a deployable repo is in
examples/app-repo/azure-pipelines.yml; it has a
deployMode parameter selecting one of two ways to hand the commit over:
| mode | pipeline side | what the pipeline learns |
|---|---|---|
wait (recommended) |
agentless job (pool: server) running PublishToAzureServiceBus@2 with an Azure Resource Manager service connection (identity needs Azure Service Bus Data Sender on the queue — no keys) and waitForCompletion: true |
the stage waits for the deploy and succeeds/fails with it: the daemon streams the deploy log to the task and posts TaskCompleted back to Azure Pipelines — an outbound call from the box |
script |
any agent, python3 ci/sb-send.py … (stdlib only; a copy of scripts/sb-send.py) with a Send-only SAS key in a secret variable |
nothing — fire-and-forget; the result is in the box's records (deploy-this status hello) |
The envelope in both cases:
{"app": "hello", "ref": "3f9c2a1b…", "action": "apply", "requested_by": "ado:hello#123"}ref should be $(Build.SourceVersion) so the box deploys exactly the commit that passed the
tests. action can also be stage, swap, unstage, rollback (with ref = release id, e.g.
r7) or remove.
On the box, install examples/operator/deploy-this-daemon.service
with AZURE_SERVICEBUS_CONNECTION_STRING=… (a Listen-only policy) or
triggers.servicebus.namespace (an FQDN, in config.yml) + a managed identity, then
systemctl enable --now deploy-this-daemon. Handled messages are completed even when the
deploy failed (the failure is recorded and retrying would not help); infrastructure errors
abandon the message for redelivery; malformed envelopes are dead-lettered (and reported as
failed to a waiting pipeline task).
Signing. Anyone who can send to the queue can deploy anything, so envelopes can carry an
HMAC-SHA256 signature (plus issued_at and message_id) computed with a secret shared between
the pipeline (deploy-this-envelope-secret) and the box (DEPLOY_THIS_ENVELOPE_SECRET). The
example pipeline signs in its test stage; sb-send.py --secret signs too. With
triggers.signing.required: true the daemon dead-letters unsigned, tampered or stale
envelopes (and tells a waiting pipeline task); until then it accepts unsigned ones with a warning
but always rejects a bad signature.
Notifications. Set DEPLOY_THIS_NOTIFY_URL (Slack/Teams incoming webhook) or notify.url
to get one line per event — deployed, failed, rolled-back, removed, and later healed/swapped —
or notify.format: json for the full event. Best effort: a dead webhook never fails a deploy.
Approvals, environments, gates. Put the human gate where Azure DevOps already has one:
- Approvals and checks on the service connection (Project settings → Service connections →
deploy-this-servicebus→ Approvals and checks): every stage that uses it — i.e. everywait-mode deploy — pauses for the approvers. Zero pipeline changes. - Environment with approvals: for
scriptmode, make the enqueue job adeploymentjob withenvironment: production; the environment's approvals/checks (business hours, required template, Azure Monitor alerts) run before it. - Stage, approve, swap: enqueue
"action": "stage", let aManualValidation@0server job (or an environment approval on the next stage) wait for a human who looks athttps://hello-staging.…, then enqueue"action": "swap". Swapping back is another"swap"— no rebuild, nothing restarted.
The example pipeline has a commented approve job showing the stage → approve → swap shape.
Things that bite:
useDataContractSerializer: falseon the task — otherwise the body arrives wrapped in .NET serialisation XML. The daemon unwraps that too, but don't rely on it.- The box needs outbound AMQP on TCP 5671; if only 443 is open, set
triggers.servicebus.transport: websocket(AMQP over WebSockets). - Hosted agents and the agentless task reach the namespace from Azure-side addresses, so the namespace must allow public network access (or use a self-hosted agent and a firewall rule).
- The box's deploy user needs read access to the app repo (SSH key or PAT in its git credential store) to fetch the commit.
from deploy_this.config import load_config
from deploy_this.engine import Engine
engine = Engine(load_config()) # same engine the CLI and daemon use
result = engine.apply("hello", ref="main", requested_by="scheduler")
print(result.action, result.record.release, result.record.port)Pluggable interfaces: deploy_this.routing.RouteBackend, deploy_this.recorder.Recorder,
deploy_this.triggers.Trigger, deploy_this.privileged.Helper. Every external effect of the
engine is injectable, which is how the test-suite runs the whole flow without Docker.
/etc/deploy-this/config.yml, apps/*.yml operator config, manifests
/etc/caddy/deploy-this.d/<app>.caddy the durable route (one per app)
/etc/systemd/system/deploy-this-<app>-rN.service one unit per release (current release only)
/var/lib/deploy-this/records/<app>/rN.json deploy records (+ CURRENT pointer)
/var/lib/deploy-this/routes/<app>.json desired route state
/var/lib/deploy-this/releases/<app>/rN/env env file (0600), unit copy
/var/lib/deploy-this/secrets/<app>.env secrets (0600)
/var/lib/deploy-this/data/<app>/<volume>/ persistent volumes (never deleted)
/var/lib/deploy-this/heal/, instances/ auto-heal counters; preview instances + expiry
To undo everything by hand: systemctl disable --now 'deploy-this-*', delete the units, delete
/etc/caddy/deploy-this.d/*.caddy, systemctl reload caddy,
docker image ls 'deploy-this/*' -q | xargs -r docker image rm.
Trusted first-party code only. The deploy user is unprivileged but can run the four-verb root
helper via sudo; because install-unit installs system units, that user is root-equivalent
by design — the helper's value is legibility (one file, four verbs, every call in the
journal), not containment of a hostile deployer. Apps are isolated from each other and from
the box by Docker defaults plus --cap-drop ALL, no-new-privileges, caps, loopback-only
ports, private bridges and default-deny egress. Details and limits: docs/design.md §5–6.
- Caddy's admin endpoint. The default
localhost:2019is reachable by every local user, which makes every local user a proxy administrator. Move it to a group-owned unix socket: a drop-in forcaddy.servicewithRuntimeDirectory=caddy,{ admin unix//run/caddy/admin.sock|0660 }in the Caddyfile (the|0660permission suffix needs Caddy ≥ 2.7),usermod -aG caddy deploy, andcaddy.admin: unix:///run/caddy/admin.sockinconfig.yml—doctor, the splice and thecaddyfilebackend's reload all follow that setting. - The deploy user needs only these extra groups:
docker(it is root-equivalent through Docker anyway, see below),systemd-journal(fordeploy-this logs) and, if you moved the admin endpoint to a unix socket,caddy. Nothing else. - Backups. Put
/var/lib/deploy-thison the same backup schedule as/etc: records, desired routes and secrets live there, androutes syncrebuilds/etc/caddy/deploy-this.dfrom it. - Caddy
stream_close_delay. Set it on any operator route that proxies WebSockets; every Caddy config change (yours or deploy-this's) closes proxied WebSockets otherwise.
python3 -m venv .venv && . .venv/bin/activate
pip install -e '.[dev]'
ruff check . && ruff format --check . && mypy && pytest
DEPLOY_THIS_TEST_CADDY=/usr/local/bin/caddy pytest -m caddy # live tests against a real Caddy binaryCI runs lint, types, unit + live-Caddy tests on Python 3.10, 3.12, 3.13 and 3.14, then provisions an Ubuntu
runner (systemd + Docker + Caddy service) and executes scripts/e2e.sh: deploy, idempotent
re-apply, failed deploy with automatic rollback, manual rollback, operator reload, remove —
while asserting zero non-200 responses from a canary site and identical MainPID /
NRestarts / start timestamps for caddy.service, docker.service and the canary.
Commits follow Conventional Commits. See CONTRIBUTING.md, docs/releasing.md and, for the first deployment on a real box, docs/first-deploy.md.
MIT — see LICENSE.