Skip to content

Repository files navigation

deploy-this

Push-to-deploy for one important Linux box. A small Python library + CLI that turns a git ref with a Dockerfile into a running, routed, TLS'd, health-checked, resource-capped service behind the Caddy reverse proxy you already run — with automatic rollback on a failed health check, and without restarting the proxy or anything else on the box.

deploy-this apply hello
hello: resolving https://git.example.internal/platform/hello.git @ main
hello: starting release r7 (3f9c2a1b0d4e, requested by alice)
hello: building deploy-this/hello:r7
hello: starting deploy-this-hello-r7.service on 127.0.0.1:41003
hello r7 health gate (127.0.0.1:41003/healthz): healthy after 3 probe(s) (HTTP 200 OK b'{"ok": true}')
hello: cutting route hello.apps.example.internal over to 127.0.0.1:41003
hello: fragment written; live route replaced (srv0/routes/0)
hello through-proxy check (https://hello.apps.example.internal/healthz): healthy after 1 probe(s)
hello: retiring r6 (deploy-this-hello-r6.service)
hello: release r7 deployed (3f9c2a1b0d4e) in 41s

It is built for the "one box shared with services that must not be disturbed" case — no Kubernetes, no Swarm, no second proxy, no database, no new privileged daemon. MIT licensed.

What you get

  • One engine, three entry points: the CLI, a pull-based queue daemon (Azure Service Bus first; webhook and git-poll adapters included), and an importable library.
  • Zero-downtime, health-gated deploys: the new release starts next to the old one, must pass an HTTP health gate, is cut over with a single surgical change to Caddy's running config (no reload, no restart), is verified through the proxy, and only then is the old release retired. Any failure leaves the old release serving. rollback re-runs the same flow with a previous image.
  • Your Caddy stays yours: deploy-this owns one directory of Caddyfile fragments and one marker-tagged route per app, nothing else. Operator snippets (wildcard TLS, forward-auth) are used as-is. Reloads and restarts by the operator keep working. See docs/design.md for the research and the evidence.
  • Mandatory caps and anti-footgun isolation: memory/CPU/pids via docker run flags, --cap-drop ALL, no-new-privileges, loopback-only ports, a private bridge per app, and default-deny egress with a per-app allowlist.
  • Preview instances: one throwaway copy per pull request with a TTL, at its own subdomain.
  • Key Vault secret references: secrets: {DATABASE_URL: keyvault://vault/name} resolved on the box at deploy time (also file:// and env://); rotation is picked up by the next apply.
  • Scheduled functions: function.schedule turns a function app into a systemd timer job.
  • Staging slot and swap: apply --stage → look → swap; the old release stays running in staging so swapping back is instant. Or strategy: canary for weighted gradual cutover.
  • Build where it suits you: on the box (in a resource-capped builder) or in CI, deploying registry images by digest so the shared box never builds at all.
  • Durable records: who deployed what, when, from which ref, with which image, and how it went — as JSON files (or SQLite) that survive reboots and disk swaps.
  • Legible and undoable: every unit, fragment, env file and record is a plain file you can read; deploy-this doctor, status, routes show tell you what the box thinks.
  • One privileged component: a single stdlib-only root helper with exactly four verbs (install-unit, remove-unit, apply-route, remove-route).
  • Self-repairing: a one-minute maintain timer auto-heals unhealthy releases (health-check
    • restart, rate-limited) and reconciles routes and units with the recorded desired state.
  • Signed envelopes and notifications: HMAC-signed deploy requests (reject the rest), and a webhook line per deploy event for Slack/Teams.

Runtime dependency: PyYAML. Optional extras: [servicebus] (azure-servicebus + azure-identity) for the queue daemon, [keyvault] (azure-keyvault-secrets + azure-identity) for Key Vault references.

Three places, three roles

It is easy to mix up what goes where, so:

  your service's repo                         the box (operator-owned)             this project
  ───────────────────                         ────────────────────────             ────────────
  code + Dockerfile                           deploy-this + its daemon              the tool both
  azure-pipelines.yml ──{app, ref}──▶ queue ──▶ /etc/deploy-this/apps/hello.yml    sides install
           ▲                                   /etc/deploy-this/config.yml
           └────────── git fetch by sha ────── /etc/caddy/…  (+ one import line)
  • The deployable repo (examples/app-repo) contains the app, its Dockerfile, and a pipeline whose last step enqueues {"app": "hello", "ref": "<sha>"}. It never connects to the box and never pushes an image.
  • The box (examples/operator) holds the operator config, one manifest per app (where the repo is, which host/auth policy, which caps), the Caddyfile snippets, the root helper and the daemon that consumes the queue and does the clone/build/deploy.
  • deploy-this (this repo) is the library/CLI installed on the box. Its examples/app-repo directory is only a stand-in for a separate repository.

10-minute quickstart on a stock Ubuntu VM

Assumes Ubuntu 22.04/24.04 with sudo. Replace *.apps.example.internal with a name your wildcard certificate covers.

1. Prerequisites: Docker, git, Caddy.

sudo apt-get update && sudo apt-get install -y docker.io docker-buildx git python3-venv curl
# Caddy (official repo): https://caddyserver.com/docs/install#debian-ubuntu-raspbian
sudo apt-get install -y debian-keyring debian-archive-keyring apt-transport-https
curl -1sLf 'https://dl.cloudsmith.io/public/caddy/stable/gpg.key' | sudo gpg --dearmor -o /usr/share/keyrings/caddy-stable-archive-keyring.gpg
curl -1sLf 'https://dl.cloudsmith.io/public/caddy/stable/debian.deb.txt' | sudo tee /etc/apt/sources.list.d/caddy-stable.list
sudo apt-get update && sudo apt-get install -y caddy

2. A deploy user that can use Docker and read the journal, but is otherwise unprivileged.

sudo useradd --system --create-home --shell /bin/bash --groups docker,systemd-journal deploy

3. Install deploy-this from git into a venv, plus the root helper and its sudoers rule.

deploy-this is installed straight from its git repository (it is not on PyPI). Pin a tag:

sudo python3 -m venv /opt/deploy-this
sudo /opt/deploy-this/bin/pip install 'deploy-this[servicebus] @ git+https://github.com/jackhumbert/deploy-this-py@v0.1.0'
sudo install -m 0755 "$(/opt/deploy-this/bin/deploy-this helper-path)" /usr/local/bin/deploy-this-helper
printf 'Defaults!/usr/local/bin/deploy-this-helper !requiretty\ndeploy ALL=(root) NOPASSWD: /usr/local/bin/deploy-this-helper\n' \
  | sudo tee /etc/sudoers.d/deploy-this >/dev/null && sudo chmod 0440 /etc/sudoers.d/deploy-this

(examples/operator/sudoers in this repo is the same two lines.)

4. Config and directories.

sudo mkdir -p /etc/deploy-this/apps /var/lib/deploy-this /etc/caddy/deploy-this.d
sudo curl -sSLo /etc/deploy-this/config.yml https://raw.githubusercontent.com/jackhumbert/deploy-this-py/main/examples/operator/config.yml
sudo chown -R deploy:deploy /var/lib/deploy-this /etc/deploy-this/apps && sudo chmod 0750 /var/lib/deploy-this

The defaults in examples/operator/config.yml match this layout; edit caddy.tls_snippet / caddy.auth_policies to your snippet names.

5. Three lines in your Caddyfile (snippets before the import; see examples/operator/Caddyfile for a complete one). The TLS snippet assumes your wildcard certificate is at /etc/caddy/certs/; for a first try on a throwaway VM, mint a self-signed one:

sudo mkdir -p /etc/caddy/certs && sudo openssl req -x509 -newkey rsa:2048 -nodes -days 3650 \
  -keyout /etc/caddy/certs/wildcard.key -out /etc/caddy/certs/wildcard.pem \
  -subj "/CN=*.apps.example.internal" -addext "subjectAltName=DNS:*.apps.example.internal"
sudo chown caddy:caddy /etc/caddy/certs/* && sudo chmod 0640 /etc/caddy/certs/wildcard.key
(wildcard_tls) {
	tls /etc/caddy/certs/wildcard.pem /etc/caddy/certs/wildcard.key
}
(auth_internal) {
	forward_auth 127.0.0.1:9091 {
		uri /api/verify
		copy_headers X-User
	}
}
import /etc/caddy/deploy-this.d/*.caddy
sudo systemctl reload caddy    # one last reload; deploy-this never needs another

6. Check the box.

sudo -u deploy /opt/deploy-this/bin/deploy-this doctor

7. Describe an app and deploy it.

sudo -u deploy tee /etc/deploy-this/apps/hello.yml <<'YML'
runtime: web
source: {repo: https://github.com/jackhumbert/deploy-this-py, ref: main, subdir: examples/app-repo}
port: 8080
route: {host: hello.apps.example.internal, auth: public}   # `internal` once auth_internal points at your real forward-auth service
health: {path: /healthz, timeout: 60s}
limits: {memory: 256m, cpus: 0.5, pids: 64}
egress: {allow: []}
YML
sudo -u deploy /opt/deploy-this/bin/deploy-this apply hello
curl -sk --resolve hello.apps.example.internal:443:127.0.0.1 https://hello.apps.example.internal/

8. Operate it.

deploy-this status hello        # release, unit, container, health, route (live + desired), history
deploy-this list                # every app at a glance
deploy-this logs hello -f       # journalctl -u deploy-this-hello-rN.service -f
deploy-this rollback hello      # previous successful image, same health-gated flow
deploy-this remove hello        # route, unit, network, images gone; data kept

Then install the two timers (self-repair every minute, disk GC daily):

for u in deploy-this-maintain.service deploy-this-maintain.timer deploy-this-gc.service deploy-this-gc.timer; do
  sudo curl -sSLo /etc/systemd/system/$u https://raw.githubusercontent.com/jackhumbert/deploy-this-py/main/examples/operator/$u
done
sudo systemctl daemon-reload && sudo systemctl enable --now deploy-this-maintain.timer deploy-this-gc.timer

Prefer a script? scripts/ci-setup.sh provisions a fresh VM (it writes its own Caddyfile with a self-signed wildcard cert, so do not run it on a box that already has one), and scripts/e2e.sh then runs the full deploy → fail → rollback → remove proof while checking that Caddy, Docker and a canary service never restart. That is exactly what CI runs.

Installing and upgrading from git

There is no PyPI release; the box installs from the repository with a pinned tag, and upgrades by installing a newer tag (a changed URL is enough for pip to rebuild — no --force-reinstall needed):

# GitHub (public)
sudo /opt/deploy-this/bin/pip install 'deploy-this[servicebus] @ git+https://github.com/jackhumbert/deploy-this-py@v0.1.0'
# Azure Repos over SSH (deploy key / the deploy user's key)
sudo /opt/deploy-this/bin/pip install 'deploy-this[servicebus] @ git+ssh://git@ssh.dev.azure.com/v3/<org>/<project>/deploy-this@v0.1.0'
# Azure Repos over HTTPS with a PAT kept out of the command line
sudo git config --system credential.helper store   # then put https://<user>:<PAT>@dev.azure.com in /root/.git-credentials (0600)
sudo /opt/deploy-this/bin/pip install 'deploy-this[servicebus] @ git+https://dev.azure.com/<org>/<project>/_git/deploy-this@v0.1.0'

After upgrading, re-copy the root helper (it is a plain file, not a console script) and restart the daemon:

sudo install -m 0755 "$(/opt/deploy-this/bin/deploy-this helper-path)" /usr/local/bin/deploy-this-helper
sudo systemctl restart deploy-this-daemon
sudo -u deploy /opt/deploy-this/bin/deploy-this doctor     # reports a helper/library version mismatch if you forgot

Tags follow semver and the CHANGELOG; deploy-this version prints what is installed.

Staging slot, swap, canary

Azure App Service's deployment slots, for one box:

deploy-this apply hello --stage   # builds + health-gates r8, routes it at https://hello-staging.apps.example.internal
# look at it, run your smoke tests against it...
deploy-this swap hello            # r8 takes the production route; r7 keeps running, parked in staging
deploy-this swap hello            # changed your mind: instant swap back, no rebuild, nothing restarted
deploy-this unstage hello         # done with it: stop whatever is in the staging slot

A swap is one route change (verified through the proxy before the pointers move); nothing is stopped, so both releases keep their ports and "swap again" is the zero-cost rollback. The staging host defaults to <first label>-staging.<rest> (set route.staging_host to change it) and uses the same auth policy and wildcard certificate. Envelopes can drive the same flow: "action": "stage" then, after an approval, "action": "swap".

For gradual exposure without a second host, strategy: canary sends canary.weight percent of production traffic to the new release (Caddy weighted_round_robin) for canary.bake, probing it the whole time; a failing probe reverts the route and removes the release, a passing bake cuts over fully and retires the old one.

Preview instances per pull request

With previews: {enabled: true, ttl: 3d, max: 5} in the manifest, deploy-this apply hello --instance pr-123 (or an envelope with "instance": "pr-123") deploys a throwaway copy at https://hello-pr-123.apps.example.internal — same build, caps, auth and egress, its own unit, port and records — and maintain removes it ttl after its last deploy. Every push to the PR redeploys the same instance and extends the TTL; deploy-this instances lists them; remove hello --instance pr-123 ends one early. examples/app-repo/azure-pipelines-preview.yml is the PR-triggered pipeline that drives it.

Scheduled functions

A function app with function.schedule (a systemd OnCalendar expression) gets a oneshot service + persistent timer per release that runs deploy-this run <app> — the WebJobs / timer-trigger equivalent. The timer survives reboots (missed runs happen at boot), every run is the usual throwaway hardened container with the app's caps and timeout, and its output is in journalctl -u deploy-this-<app>-rN.service. status shows the next and last run.

Where images get built

Two modes, per app or per deploy:

mode who builds envelope / CLI manifest
on the box (default) the box clones the commit and runs docker build ref: <sha> / apply hello --ref source:
in CI the pipeline builds, pushes to a registry and pins the digest; the box only pulls image: myregistry.azurecr.io/hello@sha256:… / apply hello --image … image: {repository: myregistry.azurecr.io/hello}

Building in CI is the production recommendation for a shared box: build CPU, memory and disk stay on the agent, the artifact is immutable and scannable, and the box's job shrinks to pull → health gate → cutover. Envelopes may only reference images under the manifest's image.repository; tag references are resolved to a digest at deploy time and recorded.

If you do build on the box, give builds a cap: configure docker.builder (a dedicated BuildKit builder created with docker buildx create --driver docker-container and memory/CPU limits, plus a bounded build cache) — otherwise docker build runs uncapped inside the shared daemon, which is exactly what the rest of deploy-this is designed to prevent. deploy-this doctor warns when neither is set up.

Two things keep running between deploys, both from timers you install once:

  • Auto-heal (deploy-this heal, part of maintain): every current release is probed on its health path; after health.heal_after consecutive failures (default 3) — or if its unit is not active — the unit is restarted, at most heal.max_restarts_per_hour times, with a notification each time and a final "giving up" notification when the limit is hit. The Azure App Service health-check/auto-heal idea, for one box.
  • Reconcile (deploy-this reconcile, part of maintain): routes sync plus units — a current release whose unit file vanished (e.g. /etc restored from an older backup) is reinstalled from the saved unit text; deploy-this-* units that belong to no current or in-progress release (a retire that failed) are removed. Apps with a deploy in progress are skipped.

Install examples/operator/deploy-this-maintain.timer (every minute) and you have a self-repairing box; deploy-this status shows heal history.

Disk is reclaimed by deploy-this gc (old images beyond keep_releases, build cache beyond build_cache_keep, stale build directories, records beyond keep_records); install examples/operator/deploy-this-gc.timer to run it daily.

How a deploy works

resolve ref ─▶ fingerprint (no-op if unchanged) ─▶ clone ─▶ docker build
  ─▶ route preflight ─▶ allocate loopback port ─▶ install unit (new release starts beside the old)
  ─▶ health gate: GET 127.0.0.1:<port>/healthz        ✗ → remove new unit; old keeps serving
  ─▶ cutover: write fragment + splice one route into Caddy's running config (no reload)
  ─▶ verify through the proxy: GET https://<host>/healthz via 127.0.0.1:443   ✗ → revert route, remove new unit
  ─▶ retire old release (docker stop -t <stop_grace>) ─▶ record ─▶ prune old images

Each release is a systemd unit (deploy-this-<app>-rN.service) running docker run --rm … -p 127.0.0.1:<port>:<app port> --memory … --cpus … --pids-limit … with a per-release egress chain in DOCKER-USER. Units, fragments and env files are plain files; read them.

The routing contract (short version)

  • You add one import /etc/caddy/deploy-this.d/*.caddy line after your snippets.
  • deploy-this writes one fragment per app (<host> { import <tls>; import <auth>; vars deploy_this_app <app>; reverse_proxy 127.0.0.1:<port> }) through the root helper, which validates the whole Caddyfile and reverts on error.
  • For the live change it runs caddy adapt on your Caddyfile (read-only), takes the one route carrying the deploy_this_app marker, and PATCHes/PUTs just that route via the admin API with If-Match. Nothing else in the running config is touched; Caddy swaps configs atomically and keeps listeners open, so no request is refused and in-flight requests finish.
  • Your own systemctl reload caddy keeps working and keeps the apps routed (the fragment is in your Caddyfile). deploy-this routes sync reconciles live state with desired state whenever you want.
  • Operators with WebSocket services: set stream_close_delay on those reverse_proxy blocks — every Caddy config change (yours or ours) closes proxied WebSockets otherwise. See docs/design.md §3 for why and for the evidence.

The alternative backend route_backend: caddyfile writes the same fragment and runs a graceful caddy reload instead of the splice.

Manifests

One YAML file per app in apps_dir. Full reference: docs/manifest.md.

runtime: web                      # web | function | process
source: {repo: ..., ref: main, subdir: .}   # and/or image: {repository: myregistry.azurecr.io/hello}
build: {dockerfile: Dockerfile, args: {}, target: null}
port: 8080                        # what the container listens on (also exported as $PORT)
route: {host: hello.apps.example.internal, auth: internal, extra: [], staging_host: null}
health: {path: /healthz, timeout: 60s, interval: 2s, expect: [200]}
limits: {memory: 256m, cpus: 0.5, pids: 64}      # required
egress: {allow: ["10.0.0.0/8:5432", "api.example.com:443"]}   # default-deny; or `egress: unrestricted`
env: {LOG_LEVEL: info}
volumes: ["data:/data"]           # persisted under <state_dir>/data/<app>/
security: {read_only: false, cap_drop: [ALL], cap_add: [], no_new_privileges: true}
strategy: blue-green              # or canary (+ canary: {weight: 10, bake: 5m}), or in-place (+ fixed_port)
stop_grace: 30s

Secret values never go in the manifest. Two ways in:

secrets:                                                    # references, resolved on the box at deploy time
  DATABASE_URL: keyvault://kv-example/hello-database-url    # Azure Key Vault (pip install 'deploy-this[keyvault]', managed identity)
  API_KEY: keyvault://kv-example/hello-api-key/8f1c0a…      # pinned version
  TLS_KEY: file:///etc/deploy-this/secrets/hello-tls.key    # root-managed file readable by the deploy user
  SENTRY_DSN: env://HELLO_SENTRY_DSN                        # from the daemon's environment

or deploy-this secrets set hello DATABASE_URL=..., which stores the value in <state_dir>/secrets/hello.env (0600) and overrides a reference of the same name. Either way the values are injected with --env-file, never logged, and their hash is part of the release fingerprint — so a rotated Key Vault secret is picked up by the next apply (secrets check verifies references without printing anything secret).

CLI

Command What it does
apply <app> [--ref REF | --image REF] [--force] build (or pull a CI-built image) + health-gated zero-downtime deploy; no-op if nothing changed
rollback <app> [--to rN] redeploy a previous successful image, same gate
apply <app> --instance pr-123 [--ttl 3d] / instances / remove <app> --instance pr-123 preview instances per pull request, expired by maintain
apply <app> --stage / swap <app> / unstage <app> deploy into the staging slot; promote it (old release parks in staging — swap again to roll back); clear the slot
remove <app> [--keep-images] route, unit, network, images removed; persistent data kept
status <app> [--json] / list [--json] what is deployed, unit/container/health/route state, history
logs <app> [-f] [-n N] journalctl -u of the current release
run <app> [-- args] run a function app once in a throwaway hardened container
`secrets set unset
`routes sync list
validate [app] check manifests
maintain [--dry-run] / heal [app] / reconcile [--dry-run] auto-heal unhealthy releases; reinstall missing units, remove orphans, sync routes (run from a timer)
gc [--dry-run] reclaim disk: old images, build cache, stale build dirs, old records
`daemon [--trigger servicebus webhook
doctor check config, docker, caddy, helper, snippets, ports
helper-path / version where the bundled root helper is (to install it); installed version

Global flags: -c/--config PATH (default $DEPLOY_THIS_CONFIG or /etc/deploy-this/config.yml), -v for debug logging on stderr; apply/rollback/remove/swap/unstage accept --requested-by for the records.

Exit codes: 1 generic failure, 2 config/manifest/secrets, 3 build, 4 helper, 5 route, 6 health gate, 7 not deployed, 8 locked, 9 command, 130 interrupted.

Pull-based CI (Azure DevOps)

The box accepts no inbound connections, so the app repo's pipeline enqueues a small envelope and the daemon on the box does the work. A complete pipeline for a deployable repo is in examples/app-repo/azure-pipelines.yml; it has a deployMode parameter selecting one of two ways to hand the commit over:

mode pipeline side what the pipeline learns
wait (recommended) agentless job (pool: server) running PublishToAzureServiceBus@2 with an Azure Resource Manager service connection (identity needs Azure Service Bus Data Sender on the queue — no keys) and waitForCompletion: true the stage waits for the deploy and succeeds/fails with it: the daemon streams the deploy log to the task and posts TaskCompleted back to Azure Pipelines — an outbound call from the box
script any agent, python3 ci/sb-send.py … (stdlib only; a copy of scripts/sb-send.py) with a Send-only SAS key in a secret variable nothing — fire-and-forget; the result is in the box's records (deploy-this status hello)

The envelope in both cases:

{"app": "hello", "ref": "3f9c2a1b…", "action": "apply", "requested_by": "ado:hello#123"}

ref should be $(Build.SourceVersion) so the box deploys exactly the commit that passed the tests. action can also be stage, swap, unstage, rollback (with ref = release id, e.g. r7) or remove.

On the box, install examples/operator/deploy-this-daemon.service with AZURE_SERVICEBUS_CONNECTION_STRING=… (a Listen-only policy) or triggers.servicebus.namespace (an FQDN, in config.yml) + a managed identity, then systemctl enable --now deploy-this-daemon. Handled messages are completed even when the deploy failed (the failure is recorded and retrying would not help); infrastructure errors abandon the message for redelivery; malformed envelopes are dead-lettered (and reported as failed to a waiting pipeline task).

Signing. Anyone who can send to the queue can deploy anything, so envelopes can carry an HMAC-SHA256 signature (plus issued_at and message_id) computed with a secret shared between the pipeline (deploy-this-envelope-secret) and the box (DEPLOY_THIS_ENVELOPE_SECRET). The example pipeline signs in its test stage; sb-send.py --secret signs too. With triggers.signing.required: true the daemon dead-letters unsigned, tampered or stale envelopes (and tells a waiting pipeline task); until then it accepts unsigned ones with a warning but always rejects a bad signature.

Notifications. Set DEPLOY_THIS_NOTIFY_URL (Slack/Teams incoming webhook) or notify.url to get one line per event — deployed, failed, rolled-back, removed, and later healed/swapped — or notify.format: json for the full event. Best effort: a dead webhook never fails a deploy.

Approvals, environments, gates. Put the human gate where Azure DevOps already has one:

  • Approvals and checks on the service connection (Project settings → Service connections → deploy-this-servicebus → Approvals and checks): every stage that uses it — i.e. every wait-mode deploy — pauses for the approvers. Zero pipeline changes.
  • Environment with approvals: for script mode, make the enqueue job a deployment job with environment: production; the environment's approvals/checks (business hours, required template, Azure Monitor alerts) run before it.
  • Stage, approve, swap: enqueue "action": "stage", let a ManualValidation@0 server job (or an environment approval on the next stage) wait for a human who looks at https://hello-staging.…, then enqueue "action": "swap". Swapping back is another "swap" — no rebuild, nothing restarted.

The example pipeline has a commented approve job showing the stage → approve → swap shape.

Things that bite:

  • useDataContractSerializer: false on the task — otherwise the body arrives wrapped in .NET serialisation XML. The daemon unwraps that too, but don't rely on it.
  • The box needs outbound AMQP on TCP 5671; if only 443 is open, set triggers.servicebus.transport: websocket (AMQP over WebSockets).
  • Hosted agents and the agentless task reach the namespace from Azure-side addresses, so the namespace must allow public network access (or use a self-hosted agent and a firewall rule).
  • The box's deploy user needs read access to the app repo (SSH key or PAT in its git credential store) to fetch the commit.

Library use

from deploy_this.config import load_config
from deploy_this.engine import Engine

engine = Engine(load_config())           # same engine the CLI and daemon use
result = engine.apply("hello", ref="main", requested_by="scheduler")
print(result.action, result.record.release, result.record.port)

Pluggable interfaces: deploy_this.routing.RouteBackend, deploy_this.recorder.Recorder, deploy_this.triggers.Trigger, deploy_this.privileged.Helper. Every external effect of the engine is injectable, which is how the test-suite runs the whole flow without Docker.

Where things live

/etc/deploy-this/config.yml, apps/*.yml          operator config, manifests
/etc/caddy/deploy-this.d/<app>.caddy             the durable route (one per app)
/etc/systemd/system/deploy-this-<app>-rN.service one unit per release (current release only)
/var/lib/deploy-this/records/<app>/rN.json       deploy records (+ CURRENT pointer)
/var/lib/deploy-this/routes/<app>.json           desired route state
/var/lib/deploy-this/releases/<app>/rN/env       env file (0600), unit copy
/var/lib/deploy-this/secrets/<app>.env           secrets (0600)
/var/lib/deploy-this/data/<app>/<volume>/        persistent volumes (never deleted)
/var/lib/deploy-this/heal/, instances/           auto-heal counters; preview instances + expiry

To undo everything by hand: systemctl disable --now 'deploy-this-*', delete the units, delete /etc/caddy/deploy-this.d/*.caddy, systemctl reload caddy, docker image ls 'deploy-this/*' -q | xargs -r docker image rm.

Security model, briefly

Trusted first-party code only. The deploy user is unprivileged but can run the four-verb root helper via sudo; because install-unit installs system units, that user is root-equivalent by design — the helper's value is legibility (one file, four verbs, every call in the journal), not containment of a hostile deployer. Apps are isolated from each other and from the box by Docker defaults plus --cap-drop ALL, no-new-privileges, caps, loopback-only ports, private bridges and default-deny egress. Details and limits: docs/design.md §5–6.

Hardening the box (recommended)

  • Caddy's admin endpoint. The default localhost:2019 is reachable by every local user, which makes every local user a proxy administrator. Move it to a group-owned unix socket: a drop-in for caddy.service with RuntimeDirectory=caddy, { admin unix//run/caddy/admin.sock|0660 } in the Caddyfile (the |0660 permission suffix needs Caddy ≥ 2.7), usermod -aG caddy deploy, and caddy.admin: unix:///run/caddy/admin.sock in config.yml — doctor, the splice and the caddyfile backend's reload all follow that setting.
  • The deploy user needs only these extra groups: docker (it is root-equivalent through Docker anyway, see below), systemd-journal (for deploy-this logs) and, if you moved the admin endpoint to a unix socket, caddy. Nothing else.
  • Backups. Put /var/lib/deploy-this on the same backup schedule as /etc: records, desired routes and secrets live there, and routes sync rebuilds /etc/caddy/deploy-this.d from it.
  • Caddy stream_close_delay. Set it on any operator route that proxies WebSockets; every Caddy config change (yours or deploy-this's) closes proxied WebSockets otherwise.

Development

python3 -m venv .venv && . .venv/bin/activate
pip install -e '.[dev]'
ruff check . && ruff format --check . && mypy && pytest
DEPLOY_THIS_TEST_CADDY=/usr/local/bin/caddy pytest -m caddy     # live tests against a real Caddy binary

CI runs lint, types, unit + live-Caddy tests on Python 3.10, 3.12, 3.13 and 3.14, then provisions an Ubuntu runner (systemd + Docker + Caddy service) and executes scripts/e2e.sh: deploy, idempotent re-apply, failed deploy with automatic rollback, manual rollback, operator reload, remove — while asserting zero non-200 responses from a canary site and identical MainPID / NRestarts / start timestamps for caddy.service, docker.service and the canary.

Commits follow Conventional Commits. See CONTRIBUTING.md, docs/releasing.md and, for the first deployment on a real box, docs/first-deploy.md.

License

MIT — see LICENSE.

About

Push-to-deploy for one important Linux box: Docker apps behind your existing Caddy — health-gated, zero-downtime, resource-capped, with rollback. No control plane.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages