Skip to content

Deploy with Docker / Podman

Danbyte ships a production container stack - Postgres, Redis, the Django backend (gunicorn), a separate WebSocket process (daphne), a background worker pool, and nginx serving the built SPA. It runs the same way under Docker and Podman. (Addresses issue #19.)

Bare metal vs containers

For a single VM, the Installation script (systemd units + host nginx) is still the smoothest path. Containers suit hosts where you already run Docker/Podman, or an orchestrator.

Quick start

git clone https://github.com/danbyte-net/danbyte.git
cd danbyte
cp deploy/docker/.env.example .env

Edit .env - at minimum set the two secrets and the DB password. Generate each secret with:

python -c "import secrets; print(secrets.token_urlsafe(64))"

Then bring the stack up:

docker compose -f docker-compose.prod.yml --env-file .env up -d --build
podman-compose -f docker-compose.prod.yml --env-file .env up -d --build

The backend container runs migrations, the idempotent bootstrap, and collectstatic on first start, then serves. When the web container is healthy, open http://localhost:8080 (the HTTP_PORT you set).

Create the first admin (or set the DJANGO_SUPERUSER_* vars in .env and the matching lines in the compose file to have bootstrap do it):

docker compose -f docker-compose.prod.yml exec backend python manage.py createsuperuser
podman-compose -f docker-compose.prod.yml exec backend python manage.py createsuperuser

What runs

Service Image / stage Role
postgres postgres:17 Database (named volume postgres_data)
redis redis:7 Queues + channels layer
backend app (runtime) gunicorn WSGI + one-time migrate/bootstrap/static
ws app (runtime) daphne ASGI - WebSockets only
workers app (runtime) rqworker-pool (RQ_WORKERS processes; ICMP-capable)
scheduler app (runtime) the periodic beat - the container's systemd timers
frontend node (frontend) the SPA server (vite preview) - SSR build
web nginx (web) proxies the SPA + /api /ws, serves /static /media; HTTP :80 + HTTPS :443

WebSockets run as a separate daphne process, never channels-in-runserver - putting ASGI in front of all HTTP wedges plain requests. The frontend is a TanStack Start SSR build, so web proxies / to the frontend node server rather than serving files. Collected static and uploaded media live on shared volumes the backend writes and nginx serves.

The scheduler

Nothing in Danbyte polls on its own: the checks, digests, discovery and retention all have to be triggered. A bare-metal install gets that from systemd timers; a container has no init, so the scheduler service runs one process that reads the same table (core/schedule.py) and calls the same management commands. Without it the stack looks configured and measures nothing - assignments never expand into checks, so nothing ever dispatches and every check sits there having never reported.

What it runs, at the same cadence as the timers:

Cadence Work
every minute dispatch due checks, SNMP drift, Outpost work, alert escalation
every 5 min materialise check assignments, discover subnets
every 15/30 min prefix utilisation, hardware health
daily ACME renewal, retention, stale-IP cleanup, link check, certificate expiry, digest

manage.py run_scheduler --list prints the table. Auto-upgrade is the one timer a container does not run: the image is the unit of upgrade, so you deploy a new tag instead.

Occurrences are claimed in Redis, so restarting the container will not re-send this morning's digest and a second replica cannot double-send it. Keep it to one replica anyway - it buys nothing.

To run one pass by hand (a cron-driven install, or when debugging):

docker compose -f docker-compose.prod.yml exec scheduler python manage.py run_scheduler --once

ICMP

The worker container sets net.ipv4.ping_group_range so the ICMP monitor's unprivileged pings work - icmplib opens datagram sockets, which is cheaper than running as root or granting NET_RAW. Podman honours the sysctl too.

Unprivileged LXC containers cannot set this

Inside an unprivileged LXC (Proxmox and friends), writing net.ipv4.ping_group_range fails with EIO even in the container's own network namespace, so Docker refuses to start the worker or the sysctl silently does not apply. ICMP checks then report unknown rather than up/down, because the socket cannot be opened at all. Run the Docker host in a privileged LXC or on a VM/bare metal, or use a TCP/HTTP check instead of ICMP for those targets.

web listens on :80 (HTTP_PORT, default 8080) and :443 (HTTPS_PORT, default 8443) with a self-signed cert baked into the image - so browsers that force HTTPS still connect (one-time cert warning). Put a real TLS terminator in front for production and set DANBYTE_HTTPS=True.

Environment

All configuration is in .env (see deploy/docker/.env.example for the full, commented list). The essentials:

Variable Purpose
DJANGO_SECRET_KEY Django secret; unique per deployment.
MONITORING_SECRET_KEY Encrypts stored credentials; required with DEBUG=False, never change once secrets are stored.
DB_PASSWORD Postgres password (db + app).
ALLOWED_HOSTS Hosts/IPs served (no scheme/port).
CSRF_TRUSTED_ORIGINS Origins allowed to POST (scheme + host [+ port]).
DANBYTE_HTTPS True when TLS terminates in front - enables secure cookies + HSTS.
DANBYTE_TRUSTED_PROXY_DEPTH Proxies appending to X-Forwarded-For, counting the stack's own nginx: 1 direct (default), 2 with one proxy in front. Finds the client address for the login lockout.
HTTP_PORT Published host port (default 8080).
RQ_WORKERS Worker pool size.

TLS

The web container speaks plain HTTP on :80 (published as HTTP_PORT) and HTTPS with a self-signed cert on :443 (HTTPS_PORT). Terminate real TLS in front of it - a host reverse proxy, a cloud load balancer, or a caddy/traefik sidecar - then set DANBYTE_HTTPS=True and add your external URL to ALLOWED_HOSTS and CSRF_TRUSTED_ORIGINS.

Behind your own nginx

A host nginx that owns the public certificate and forwards to the stack. Point it at the HTTPS port (8443): the container's nginx sets X-Forwarded-Proto from its own listener, so forwarding to the plain :8080 listener tells Django the request was http - with DANBYTE_HTTPS=True that is a redirect loop, and without it CSRF rejects every POST from an https origin. The self-signed hop on 127.0.0.1 is fine; nginx does not verify upstream certificates by default.

# /etc/nginx/sites-available/danbyte.conf  (host nginx, in front of compose)
upstream danbyte_stack { server 127.0.0.1:8443; }

server {
    listen 80;
    server_name danbyte.example.com;
    return 301 https://$host$request_uri;
}

server {
    listen 443 ssl;
    http2 on;
    server_name danbyte.example.com;

    ssl_certificate     /etc/letsencrypt/live/danbyte.example.com/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/danbyte.example.com/privkey.pem;

    client_max_body_size 100m;

    # WebSockets (presence, SSH terminal). The Upgrade/Connection headers
    # must survive this hop, or daphne only ever sees a plain GET.
    location /ws/ {
        proxy_pass https://danbyte_stack;
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection "upgrade";
        proxy_set_header Host $host;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_read_timeout 3600s;
    }

    location / {
        proxy_pass https://danbyte_stack;
        proxy_http_version 1.1;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
    }
}

Host $host is what makes Django see the public name, so .env needs:

ALLOWED_HOSTS=danbyte.example.com
CSRF_TRUSTED_ORIGINS=https://danbyte.example.com
DANBYTE_HTTPS=True
# The host nginx is a second hop in X-Forwarded-For, in front of the stack's.
DANBYTE_TRUSTED_PROXY_DEPTH=2

DANBYTE_TRUSTED_PROXY_DEPTH tells the login lockout where the client's address sits in X-Forwarded-For. It has to match the real number of proxies: too low and every user shares the front proxy's address, too high and a client can name any address it likes. Failures are counted per client and account, plus a higher per-account ceiling across all addresses, so a wrong value is never a way to lock everyone out or to brute-force one account.

WebSocket troubleshooting

docker compose logs ws shows daphne's handshake log. What it says narrows the cause:

Log line Meaning Look at
no WSCONNECTING at all, browser gets 404/400 The upgrade never reached daphne The location /ws/ block above on every proxy hop
WSREJECT The handshake arrived and the app closed it The close code in the browser (DevTools → Network → WS)
Python traceback The consumer crashed on connect Redis / channel layer reachability from the ws container

Close codes the presence socket uses: 4401 - no authenticated session reached daphne (the session cookie was not forwarded, or a different DJANGO_SECRET_KEY on the ws service); 4400 - the session has no active tenant, or the page opened the socket without an object to watch. Hitting https://<host>:8443/ directly, bypassing the front proxy, tells the two layers apart.

Upgrading

A container can't upgrade itself - do it from the host

The in-app upgrade (Settings → Updates) is disabled on Docker/Podman deployments, and the page shows the commands below instead. A process inside a container can't rebuild its own image or recreate the container it runs in, so a self-upgrade can only ever half-apply: it may migrate the database and even swap files inside the running container, but the app processes keep executing the old code. That mismatch - new schema, old code - surfaces as confusing errors like "a required field was left empty (is_uplink)" on an otherwise healthy-looking box. Always upgrade a container deployment from the host, with the commands here.

git -C /opt/danbyte fetch --tags
git -C /opt/danbyte checkout <version>   # e.g. v0.12.0
docker compose -f docker-compose.prod.yml --env-file .env build
docker compose -f docker-compose.prod.yml --env-file .env stop scheduler workers fastlane ws
docker compose -f docker-compose.prod.yml --env-file .env up -d

build is what makes this a real upgrade: it rebuilds the images from the checked-out source, so the recreated containers actually run the new code. The backend migrates on start, before it serves; stopping the scheduler, workers, fast lane and websockets first keeps the previous release's processes from running against the database while it does (a plain up -d --build recreates them one by one, some after the migration has started). up -d then starts them all on the new image. The named volumes keep your data.

Take a backup first: Settings → Backups → Back up now writes an encrypted archive to the backups volume (/app/backups in backend and workers). A restore later runs inside workers without restarting any container - see Backup and restore.

up -d also creates services added since your last pull - check that scheduler is among them, because a stack upgraded from before it existed has never run any periodic work:

docker compose -f docker-compose.prod.yml ps scheduler

Confirm the upgrade took

After it settles, the running version should match the tag you checked out:

curl -s localhost/api/health/    # {"status":"ok",...,"version":"0.12.0"}

If version still shows the old number, the app processes weren't replaced - usually a missing --build, or containers that weren't recreated. Re-run the command above.

Podman specifics

The stack is rootless-friendly and works with podman-compose. A few notes:

  • Rootless ports < 1024: a non-root Podman can't bind :80/:443. Keep the default HTTP_PORT=8080 (or higher) and front it with a host proxy.
  • SELinux volumes: on SELinux hosts, if a bind mount is ever added, append :Z. The stack uses named volumes, which Podman labels automatically - no change needed.
  • podman play kube: podman-compose is the simplest path; if you prefer Kubernetes YAML, generate it from the running pod with podman generate kube.

Prebuilt images (ghcr.io)

Tagging a release (v*) publishes the three images to GitHub Container Registry via .github/workflows/container.yml:

ghcr.io/danbyte-net/danbyte-app:<version>       # gunicorn / daphne / workers
ghcr.io/danbyte-net/danbyte-web:<version>       # nginx + TLS
ghcr.io/danbyte-net/danbyte-frontend:<version>  # vite preview (SSR)

To run from the registry instead of building locally, set the image: fields in docker-compose.prod.yml to the ghcr paths (and drop the build: blocks), or keep a small override file. latest tracks the newest release.

Where to host

ghcr.io is the default - it ships with the GitHub repo, authenticates with the built-in GITHUB_TOKEN, and is free for public images (make the package public in the repo's Packages settings). Docker Hub, Quay, or a self-hosted Harbor work identically - change the registry/image prefix in the workflow. For fully air-gapped installs, prefer the offline tarball from the release workflow over a registry.

Development

For a lightweight dev backend (auto-reload, DEBUG=True, source bind-mounted, no nginx/frontend container) use docker-compose.dev.yml instead - see the Dev workflow.