For a Software Engineer or DevOps, the moment when running the
docker-compose up -d command in production is often accompanied by anxiety. By default, Docker Compose will shutdown (stop) the old container before building and starting the new container. This time gap of a few seconds—starting from the process stop, start, until the application is actually ready to receive traffic—will result in downtime. Your users will be greeted by a 502 Bad Gateway error from Nginx. For enterprise scale, the standard solution to this problem is to use orchestration such as Kubernetes (K8s) or Docker Swarm. However, bringing the complexity of K8s into a single VPSenvironment (single-node server) is often a stepoverkill which actually increases the burden of infrastructure maintenance.
This is where Docker Rollout comes in as a game-changer solution. Docker Rollout is a Docker CLI plugin that brings the Kubernetes-style concept of rolling updates into the simple Docker Compose ecosystem.
This article will thoroughly explore how to use Docker Rollout, complete with API deployment case studies and a Senior Engineer's "How to Solve" playbook for solving various edge-case scenarios in production.
Dissecting the Concept: How Does Docker Rollout Work?
Before getting into implementation, an engineer must understand what's going on behind the scenes. Docker Rollout does not modify Docker's daemon core, but rather performs intelligent command orchestration.
The following is a comparison of the deployment process:
| Phase | Default docker compose up -d | Using docker rollout |
| Phase 1 | Sends a SIGTERM signal to the old container. | Create a new container (2nd Replica). |
| Phase 2 | The old container is down (Downtime begins). | Waiting for the new container's Healthcheck status to become Healthy. |
| Phase 3 | Create and run a new container. | Reverse proxy (automatically) redirects traffic to the new container. |
| Phase 4 | New container booting (Downtime still occurs). | Sends a SIGTERM signal to the old container for shutdown. |
| Phase 5 | The new container is ready to receive traffic. | Old container deleted. Zero Downtime achieved. |
An absolute requirement for the above scheme to work is: You must not do binding port statically to the host (eg:
ports: - "8080:8080") and you required to use Reverse Proxy that detects container IP changes dynamically (such as Traefik or nginx-proxy).Stage 1: Docker Rollout Plugin Installation
Docker Rollout is installed as a CLI plugin of Docker, not as a standalone binary. Make sure you are using Docker Compose V2 (the
docker compose command, not the older docker-compose). Run the following command on your production server:
Bash
# 1. Create a directory for the Docker CLI plugin if it doesn't already exist
mkdir -p ~/.docker/cli-plugins
# 2. Download the latest version of the docker-rollout script from GitHub
curl https://raw.githubusercontent.com/wowu/docker-rollout/main/docker-rollout -o ~/.docker/cli-plugins/docker-rollout
# 3. Give the script execution (executable) permissions
chmod +x ~/.docker/cli-plugins/docker-rollout
To verify the installation, run:
Bash
docker rollout --help
If successful, you will see a CLI options wizard from Docker Rollout.
Phase 2: Case Study - Deployment API (Golang + PostgreSQL)
Let's use a case study that is relevant to the modern ecosystem. We will deploy a Backend API for a Point of Sale (POS) system written using Golang. This API is behind Nginx.
1. Preparation of Reverse Proxy (nginx-proxy)
We won't use vanilla Nginx because standard Nginx does cache DNS resolution. When the new API container gets a new internal IP, vanilla Nginx will keep sending traffic to the old IP until Nginx is re-reload.
The best solution is to use
jwilder/nginx-proxy.The system will listen for event from Docker socket and automatically rewrite the Nginx configuration when a new container is up.Create file
docker-compose.proxy.yml (Run once on server):YAML
services:
nginx-proxy:
image: nginxproxy/nginx-proxy
ports:
- "80:80"
- "443:443"
volumes:
- /var/run/docker.sock:/tmp/docker.sock:ro
networks:
- web-network
networks:
web-network:
external: true
Run with:
docker network create web-network && docker compose -f docker-compose.proxy.yml up -d2. Refactor Application Configuration Files
Now, let's look at the
docker-compose.yml file for our Golang API. There are some mandatory adjustments from the standard configuration.YAML
# docker-compose.yml (For Applications)
services:
api:
image: my-registry.com/pos-api:latest
# ANTI-PATTERN DOCKER ROLLOUT:
# container_name: pos-api-1 (DELETE THIS)
# ports: (DELETE THIS)# - "8080:8080"
expose:
- "8080"
environment:
- VIRTUAL_HOST=api.notopos.id
- VIRTUAL_PORT=8080
- DB_HOST=postgres
networks:
- web-network
- db-network
# MANDATORY: Healthcheck
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8080/health"]
interval: 5s
timeout: 3s
retries: 3
start_period: 10s
postgres:
image: postgres:15-alpine
environment:
- POSTGRES_PASSWORD=secret
networks:
- db-network
volumes:
- pgdata:/var/lib/postgresql/data
networks:
web-network:
external: true
db-network:
volumes:
pgdata:
Senior Engineer Analysis of the above Configuration:
-
Deletion
container_name: Docker Rollout must run 2 containers at once during transition. If the name is hardcoded, a conflict will occur. Docker Rollout will automatically rename the containers to[project]-api-1and[project]-api-2. -
Removal of
portsand Use ofexpose: Just like names, host ports cannot be used simultaneously. We only exposeexpose port 8080 into internal networkDocker. -
Environment
VIRTUAL_HOST: This is magic variable captured bynginx-proxy. -
Presence
healthcheck: This is the "soul" of zero downtime. Docker Rollout will hold the execution of deleting old containers until endpoint/healthconsistently returns an HTTP status of 200.
3. Deployment Execution with Docker Rollout
When there is a new version of code that has been build to image
[my-registry.com/pos-api:latest](https://my-registry.com/pos-api:latest), the deployment process which previously used docker compose up -d is now changed to:Bash
# 1. Pull the latest image from the registry
docker compose pull api
# 2. Do a rolling update specifically for the "api" service
docker rollout api
Log Output you will see:
Plaintext
Scaling service api to 2 instances...
Waiting for container project-api-2 to become healthy...
Container project-api-2 is healthy!
Stopping old container project-api-1...
Removing old container project-api-1...
Rollout completed successfully!
At this point, your application has updated without a single HTTP request getting an error message from Nginx.
Stage 3: "How to Solve" - Solving Problems (Troubleshooting Playbook)
In practice in production, zero downtime does not always run smoothly with just the command above.As a Senior Engineer, you are required to understand how to mitigate various legacy problems. Here is a troubleshooting playbook for complex scenarios.
Problem 1: Request Disconnects Suddenly in the Middle of Transition (Active Connection Drop)
Symptoms:
Even though the new container is up, there are some users who requests (e.g. upload files or queryheavy report) suddenly failed when the transition occurred.
Root Cause:
Docker Rollout kills old containers as soon as new containers healthy. However, Reverse Proxy (Nginx) may take 1-2 seconds to notice routing changes, and legacy containers may still be processing hundreds of request (in-flight requests).
How to Solve: Implement Container Draining with Pre-stop Hook
We have to "cheat" the system. Before the old container is shut down, we have to intentionally change its state to unhealthy. The goal is for Nginx to stop sending NEW traffic to the old container, but let the ONGOING traffic finish processing.
Add instruction pre-stop-hook via label in
docker-compose.yml: YAML
labels:
docker-rollout.pre-stop-hook: "touch /tmp/drain && sleep 15"
Then, change the
/health endpoint logic in your Golang application source code:Go
http.HandleFunc("/health", func(w http.ResponseWriter, r *http.Request) {
// If drain file exists, return HTTP 503 (Unhealthy)
if _, err := os.Stat("/tmp/drain"); err == nil {
w.WriteHeader(http.StatusServiceUnavailable)
return
}
w.WriteHeader(http.StatusOK)
w.Write([]byte("OK"))
})
How this works:
When
docker rollout finishes setting up the new container, it executes pre-stop-hook on the old container.The file /tmp/drain is created. Docker's Healthcheck now detects old containers as unhealthy. Nginx automatically removes old containers from upstream routing. The sleep 15 command gives the old container 15 seconds to complete all in-flight requests before it is completely stopped (SIGTERM). This is absolute Zero Downtime.Issue 2: Database Migration Locking
Symptoms:
The new version of the application uses a changed database table structure (schema). If the old and new containers run simultaneously during the rollout phase, the old container will crash or produce corrupt data because the database schema has changed, while the codebases don't support it yet.
Root Cause:
A tie between the lifecycle of the application and the lifecycle of the database. Running Auto-Migrate (such as GORM
AutoMigrate or Prisma migrate deploy) when the boot application is an anti-pattern on the Zero architecture Downtime.How to Solve: Backward-Compatible Migrations & Separation of Concerns
-
Separate execution: Database migration should not be executed inside the application container (via entrypoint). Use a separate container or run it before rollout.
->Bashdocker compose run --rm api migrate up docker rollout api -
Backward Compatibility Rules: Your database schema should always support 2 code versions simultaneously (N and N-1).
-
If you want to delete a column, release new code that stops reading/writing that column first (Deploy 1).
-
After the old container dies, then do a migration release that drop that column (Deploy 2).
-
Never rename a column directly.Create a new column, copy the data, read/write from 2 columns, then gradually delete the old column.
-
Issue 3: Rollout Stuck (Timeout / Infinite Wait)
Symptoms:
Your terminal is stuck at
Waiting for container project-api-2 to become healthy... for tens of minutes, or fails with a Timeout error. Root Cause:
-
Healthcheck URL is incorrect or requires authentication.
-
Application crashes when booting (e.g. due to loss of Environment New variable), so it never reaches the Healthy state.
How to Solve: Debugging & Custom Timeout
-
New Container Investigation: Do not stop the rollout process yet. Open a new terminal and check the log of the container being booting:
->Bashdocker logs project-api-2 -fLook for stack trace or errors that cause the API to fail to run. -
Set Parameters
start_period: If your application is a giant Spring Boot Java-based application or a large PHP monolith that takes 30-40 seconds for initial initialization, increase thestart_periodin yourdocker-compose.ymlfile (e.g.start_period: 60s). -
Configure Rollout Timeout: By default,parameter
docker rolloutwaits 60 seconds. You can extend the time tolerance via the-t:
->Bashdocker rollout api -t 120
Problem 4: Port Conflict Due to Hardcoded Ports in Codebase
Symptoms:
Docker responds with
Error starting userland proxy: listen tcp4 0.0.0.0:8080: bind: address already in use. Root Cause:
Even though you have removed the
ports declaration in the docker-compose.yml file, your containerized application aggressively tries to open conflicting connections at the network stack, or you are running the application in host network mode.How to Solve:
Make sure
network_mode: host is off. Docker Rollout requires network isolation (default bridge mode). Make sure Reverse proxy acts as a mediator bridging external port 80/443 with your application's internal port 8080.Conclusion
Switching schema Hard Stop (
docker-compose up -d) to Zero Downtime Deployment with docker rollout is a massive leap in infrastructure maturity. Adopting this technology not only increases the SLA (Service Level Agreement) of the system you build, but also provides peace of mind (peace of mind) to the Engineering team when they have to deploy new features during busy working hours. The main key to successful use of Docker Rollout lies in 3 architectural foundations:
-
Stateless Containers: Avoid storing asset files or sessions inside file system container (use S3 bucket or Redis).
-
Dynamic Reverse Proxy: Traefik or Nginx-Proxy is a mandatory companion so that new container internal IPs can be detected automatically.
-
Accurate Healthcheck: Don't just rely on checking
ping.Make sure your application's endpoint/healthactually checks for database connection readiness and cache before responding with the status200 OK.
By mastering the tooling and troubleshooting described above, you are ready to bring deployment reliability to the enterprise without having to worry about complicated cluster management. Happy refactoring and enjoy deploy on a drama-free Friday!

Sigit Wasis Subekti
Software Engineer & Tech Educator
Software Engineer and Tech Educator sharing insights on web development and software architecture.