Deploy with Docker / Podman¶
Danbyte ships a production container stack - Postgres, Redis, the Django backend (gunicorn), a separate WebSocket process (daphne), a background worker pool, and nginx serving the built SPA. It runs the same way under Docker and Podman. (Addresses issue #19.)
Bare metal vs containers
For a single VM, the Installation script (systemd units + host nginx) is still the smoothest path. Containers suit hosts where you already run Docker/Podman, or an orchestrator.
Quick start¶
Edit .env - at minimum set the two secrets and the DB password. Generate each
secret with:
Then bring the stack up:
The backend container runs migrations, the idempotent bootstrap, and
collectstatic on first start, then serves. When the web container is healthy,
open http://localhost:8080 (the HTTP_PORT you set).
Create the first admin (or set the DJANGO_SUPERUSER_* vars in .env and the
matching lines in the compose file to have bootstrap do it):
What runs¶
| Service | Image / stage | Role |
|---|---|---|
postgres |
postgres:17 |
Database (named volume postgres_data) |
redis |
redis:7 |
Queues + channels layer |
backend |
app (runtime) |
gunicorn WSGI + one-time migrate/bootstrap/static |
ws |
app (runtime) |
daphne ASGI - WebSockets only |
workers |
app (runtime) |
rqworker-pool (RQ_WORKERS processes; ICMP-capable) |
scheduler |
app (runtime) |
the periodic beat - the container's systemd timers |
frontend |
node (frontend) |
the SPA server (vite preview) - SSR build |
web |
nginx (web) |
proxies the SPA + /api /ws, serves /static /media; HTTP :80 + HTTPS :443 |
WebSockets run as a separate daphne process, never channels-in-runserver -
putting ASGI in front of all HTTP wedges plain requests. The frontend is a
TanStack Start SSR build, so web proxies / to the frontend node server
rather than serving files. Collected static and uploaded media live on shared
volumes the backend writes and nginx serves.
The scheduler¶
Nothing in Danbyte polls on its own: the checks, digests, discovery and
retention all have to be triggered. A bare-metal install gets that from
systemd timers; a container has no init, so the scheduler service runs one
process that reads the same table (core/schedule.py) and calls the same
management commands. Without it the stack looks configured and measures
nothing - assignments never expand into checks, so nothing ever dispatches and
every check sits there having never reported.
What it runs, at the same cadence as the timers:
| Cadence | Work |
|---|---|
| every minute | dispatch due checks, SNMP drift, Outpost work, alert escalation |
| every 5 min | materialise check assignments, discover subnets |
| every 15/30 min | prefix utilisation, hardware health |
| daily | ACME renewal, retention, stale-IP cleanup, link check, certificate expiry, digest |
manage.py run_scheduler --list prints the table. Auto-upgrade is the one timer
a container does not run: the image is the unit of upgrade, so you deploy a
new tag instead.
Occurrences are claimed in Redis, so restarting the container will not re-send this morning's digest and a second replica cannot double-send it. Keep it to one replica anyway - it buys nothing.
To run one pass by hand (a cron-driven install, or when debugging):
ICMP¶
The worker container sets net.ipv4.ping_group_range so the ICMP monitor's
unprivileged pings work - icmplib opens datagram sockets, which is cheaper
than running as root or granting NET_RAW. Podman honours the sysctl too.
Unprivileged LXC containers cannot set this
Inside an unprivileged LXC (Proxmox and friends), writing
net.ipv4.ping_group_range fails with EIO even in the container's own
network namespace, so Docker refuses to start the worker or the sysctl
silently does not apply. ICMP checks then report unknown rather than
up/down, because the socket cannot be opened at all. Run the Docker host in
a privileged LXC or on a VM/bare metal, or use a TCP/HTTP check instead of
ICMP for those targets.
web listens on :80 (HTTP_PORT, default 8080) and :443
(HTTPS_PORT, default 8443) with a self-signed cert baked into the image -
so browsers that force HTTPS still connect (one-time cert warning). Put a real
TLS terminator in front for production and set DANBYTE_HTTPS=True.
Environment¶
All configuration is in .env (see deploy/docker/.env.example for the full,
commented list). The essentials:
| Variable | Purpose |
|---|---|
DJANGO_SECRET_KEY |
Django secret; unique per deployment. |
MONITORING_SECRET_KEY |
Encrypts stored credentials; required with DEBUG=False, never change once secrets are stored. |
DB_PASSWORD |
Postgres password (db + app). |
ALLOWED_HOSTS |
Hosts/IPs served (no scheme/port). |
CSRF_TRUSTED_ORIGINS |
Origins allowed to POST (scheme + host [+ port]). |
DANBYTE_HTTPS |
True when TLS terminates in front - enables secure cookies + HSTS. |
DANBYTE_TRUSTED_PROXY_DEPTH |
Proxies appending to X-Forwarded-For, counting the stack's own nginx: 1 direct (default), 2 with one proxy in front. Finds the client address for the login lockout. |
HTTP_PORT |
Published host port (default 8080). |
RQ_WORKERS |
Worker pool size. |
TLS¶
The web container speaks plain HTTP on :80 (published as HTTP_PORT) and
HTTPS with a self-signed cert on :443 (HTTPS_PORT). Terminate real TLS in
front of it - a host reverse proxy, a cloud load balancer, or a
caddy/traefik sidecar - then set DANBYTE_HTTPS=True and add your external
URL to ALLOWED_HOSTS and CSRF_TRUSTED_ORIGINS.
Behind your own nginx¶
A host nginx that owns the public certificate and forwards to the stack. Point
it at the HTTPS port (8443): the container's nginx sets
X-Forwarded-Proto from its own listener, so forwarding to the plain :8080
listener tells Django the request was http - with DANBYTE_HTTPS=True that
is a redirect loop, and without it CSRF rejects every POST from an https
origin. The self-signed hop on 127.0.0.1 is fine; nginx does not verify
upstream certificates by default.
# /etc/nginx/sites-available/danbyte.conf (host nginx, in front of compose)
upstream danbyte_stack { server 127.0.0.1:8443; }
server {
listen 80;
server_name danbyte.example.com;
return 301 https://$host$request_uri;
}
server {
listen 443 ssl;
http2 on;
server_name danbyte.example.com;
ssl_certificate /etc/letsencrypt/live/danbyte.example.com/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/danbyte.example.com/privkey.pem;
client_max_body_size 100m;
# WebSockets (presence, SSH terminal). The Upgrade/Connection headers
# must survive this hop, or daphne only ever sees a plain GET.
location /ws/ {
proxy_pass https://danbyte_stack;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_set_header Host $host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_read_timeout 3600s;
}
location / {
proxy_pass https://danbyte_stack;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
}
}
Host $host is what makes Django see the public name, so .env needs:
ALLOWED_HOSTS=danbyte.example.com
CSRF_TRUSTED_ORIGINS=https://danbyte.example.com
DANBYTE_HTTPS=True
# The host nginx is a second hop in X-Forwarded-For, in front of the stack's.
DANBYTE_TRUSTED_PROXY_DEPTH=2
DANBYTE_TRUSTED_PROXY_DEPTH tells the login lockout where the client's
address sits in X-Forwarded-For. It has to match the real number of
proxies: too low and every user shares the front proxy's address, too high
and a client can name any address it likes. Failures are counted per client
and account, plus a higher per-account ceiling across all addresses, so a
wrong value is never a way to lock everyone out or to brute-force one
account.
WebSocket troubleshooting¶
docker compose logs ws shows daphne's handshake log. What it says narrows the
cause:
| Log line | Meaning | Look at |
|---|---|---|
no WSCONNECTING at all, browser gets 404/400 |
The upgrade never reached daphne | The location /ws/ block above on every proxy hop |
WSREJECT |
The handshake arrived and the app closed it | The close code in the browser (DevTools → Network → WS) |
| Python traceback | The consumer crashed on connect | Redis / channel layer reachability from the ws container |
Close codes the presence socket uses: 4401 - no authenticated session
reached daphne (the session cookie was not forwarded, or a different
DJANGO_SECRET_KEY on the ws service); 4400 - the session has no active
tenant, or the page opened the socket without an object to watch. Hitting
https://<host>:8443/ directly, bypassing the front proxy, tells the two
layers apart.
Upgrading¶
A container can't upgrade itself - do it from the host
The in-app upgrade (Settings → Updates) is disabled on Docker/Podman
deployments, and the page shows the commands below instead. A process
inside a container can't rebuild its own image or recreate the container it
runs in, so a self-upgrade can only ever half-apply: it may migrate the
database and even swap files inside the running container, but the app
processes keep executing the old code. That mismatch - new schema, old
code - surfaces as confusing errors like "a required field was left empty
(is_uplink)" on an otherwise healthy-looking box. Always upgrade a
container deployment from the host, with the commands here.
git -C /opt/danbyte fetch --tags
git -C /opt/danbyte checkout <version> # e.g. v0.12.0
docker compose -f docker-compose.prod.yml --env-file .env build
docker compose -f docker-compose.prod.yml --env-file .env stop scheduler workers fastlane ws
docker compose -f docker-compose.prod.yml --env-file .env up -d
build is what makes this a real upgrade: it rebuilds the images from the
checked-out source, so the recreated containers actually run the new code. The
backend migrates on start, before it serves; stopping the scheduler, workers,
fast lane and websockets first keeps the previous release's processes from
running against the database while it does (a plain up -d --build recreates
them one by one, some after the migration has started). up -d then starts
them all on the new image. The named volumes keep your data.
Take a backup first: Settings → Backups → Back up now writes an encrypted
archive to the backups volume (/app/backups in backend and workers).
A restore later runs inside workers without restarting any container - see
Backup and restore.
up -d also creates services added since your last pull - check that
scheduler is among them, because a stack upgraded from before it existed has
never run any periodic work:
Confirm the upgrade took
After it settles, the running version should match the tag you checked out:
If version still shows the old number, the app processes weren't replaced -
usually a missing --build, or containers that weren't recreated. Re-run the
command above.
Podman specifics¶
The stack is rootless-friendly and works with podman-compose. A few notes:
- Rootless ports < 1024: a non-root Podman can't bind
:80/:443. Keep the defaultHTTP_PORT=8080(or higher) and front it with a host proxy. - SELinux volumes: on SELinux hosts, if a bind mount is ever added, append
:Z. The stack uses named volumes, which Podman labels automatically - no change needed. podman play kube:podman-composeis the simplest path; if you prefer Kubernetes YAML, generate it from the running pod withpodman generate kube.
Prebuilt images (ghcr.io)¶
Tagging a release (v*) publishes the three images to GitHub Container
Registry via .github/workflows/container.yml:
ghcr.io/danbyte-net/danbyte-app:<version> # gunicorn / daphne / workers
ghcr.io/danbyte-net/danbyte-web:<version> # nginx + TLS
ghcr.io/danbyte-net/danbyte-frontend:<version> # vite preview (SSR)
To run from the registry instead of building locally, set the image: fields
in docker-compose.prod.yml to the ghcr paths (and drop the build: blocks),
or keep a small override file. latest tracks the newest release.
Where to host
ghcr.io is the default - it ships with the GitHub repo, authenticates
with the built-in GITHUB_TOKEN, and is free for public images (make the
package public in the repo's Packages settings). Docker Hub, Quay, or a
self-hosted Harbor work identically - change the registry/image
prefix in the workflow. For fully air-gapped installs, prefer the
offline tarball from the release workflow over a registry.
Development¶
For a lightweight dev backend (auto-reload, DEBUG=True, source bind-mounted,
no nginx/frontend container) use docker-compose.dev.yml instead - see the
Dev workflow.