Backend••9 min read•14 views

How to Use Docker Rollout for Zero Downtime Deployment

Still frequently facing the 502 Bad Gateway error when deploying applications on VPS? It's time to bring the concept of Kubernetes-style rolling updates to the simplicity of Docker Compose. Learn about zero downtime architecture, integration case studies, and troubleshooting here.

For a Software Engineer or DevOps, the moment when running the docker-compose up -d command in production is often accompanied by anxiety. By default, Docker Compose will shutdown (stop) the old container before building and starting the new container. This time gap of a few seconds—starting from the process stop, start, until the application is actually ready to receive traffic—will result in downtime. Your users will be greeted by a 502 Bad Gateway error from Nginx.

For enterprise scale, the standard solution to this problem is to use orchestration such as Kubernetes (K8s) or Docker Swarm. However, bringing the complexity of K8s into a single VPSenvironment (single-node server) is often a stepoverkill which actually increases the burden of infrastructure maintenance.

This is where Docker Rollout comes in as a game-changer solution. Docker Rollout is a Docker CLI plugin that brings the Kubernetes-style concept of rolling updates into the simple Docker Compose ecosystem.

This article will thoroughly explore how to use Docker Rollout, complete with API deployment case studies and a Senior Engineer's "How to Solve" playbook for solving various edge-case scenarios in production.

Dissecting the Concept: How Does Docker Rollout Work?

Before getting into implementation, an engineer must understand what's going on behind the scenes. Docker Rollout does not modify Docker's daemon core, but rather performs intelligent command orchestration.

The following is a comparison of the deployment process:

Phase Default docker compose up -d Using docker rollout
Phase 1 Sends a SIGTERM signal to the old container. Create a new container (2nd Replica).
Phase 2 The old container is down (Downtime begins). Waiting for the new container's Healthcheck status to become Healthy.
Phase 3 Create and run a new container. Reverse proxy (automatically) redirects traffic to the new container.
Phase 4 New container booting (Downtime still occurs). Sends a SIGTERM signal to the old container for shutdown.
Phase 5 The new container is ready to receive traffic. Old container deleted. Zero Downtime achieved.
An absolute requirement for the above scheme to work is: You must not do binding port statically to the host (eg: ports: - "8080:8080") and you required to use Reverse Proxy that detects container IP changes dynamically (such as Traefik or nginx-proxy).

Stage 1: Docker Rollout Plugin Installation

Docker Rollout is installed as a CLI plugin of Docker, not as a standalone binary. Make sure you are using Docker Compose V2 (the docker compose command, not the older docker-compose).

Run the following command on your production server:

Bash
# 1. Create a directory for the Docker CLI plugin if it doesn't already exist
mkdir -p ~/.docker/cli-plugins

# 2. Download the latest version of the docker-rollout script from GitHub
curl https://raw.githubusercontent.com/wowu/docker-rollout/main/docker-rollout -o ~/.docker/cli-plugins/docker-rollout

# 3. Give the script execution (executable) permissions
chmod +x ~/.docker/cli-plugins/docker-rollout
->
To verify the installation, run:

Bash
docker rollout --help
->
If successful, you will see a CLI options wizard from Docker Rollout.

Phase 2: Case Study - Deployment API (Golang + PostgreSQL)

Let's use a case study that is relevant to the modern ecosystem. We will deploy a Backend API for a Point of Sale (POS) system written using Golang. This API is behind Nginx.

1. Preparation of Reverse Proxy (nginx-proxy)

We won't use vanilla Nginx because standard Nginx does cache DNS resolution. When the new API container gets a new internal IP, vanilla Nginx will keep sending traffic to the old IP until Nginx is re-reload.

The best solution is to use jwilder/nginx-proxy.The system will listen for event from Docker socket and automatically rewrite the Nginx configuration when a new container is up.

Create file docker-compose.proxy.yml (Run once on server):

YAML
services:
  nginx-proxy:
    image: nginxproxy/nginx-proxy
    ports:
      - "80:80"
      - "443:443"
    volumes:
      - /var/run/docker.sock:/tmp/docker.sock:ro
    networks:
      - web-network

networks:
  web-network:
    external: true
->
Run with: docker network create web-network && docker compose -f docker-compose.proxy.yml up -d

2. Refactor Application Configuration Files

Now, let's look at the docker-compose.yml file for our Golang API. There are some mandatory adjustments from the standard configuration.

YAML
# docker-compose.yml (For Applications)
services:
  api:
    image: my-registry.com/pos-api:latest
    # ANTI-PATTERN DOCKER ROLLOUT:
    # container_name: pos-api-1 (DELETE THIS)
    # ports: (DELETE THIS)# - "8080:8080"
    expose:
      - "8080"
    environment:
      - VIRTUAL_HOST=api.notopos.id
      - VIRTUAL_PORT=8080
      - DB_HOST=postgres
    networks:
      - web-network
      - db-network
    # MANDATORY: Healthcheck
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8080/health"]
      interval: 5s
      timeout: 3s
      retries: 3
      start_period: 10s

  postgres:
    image: postgres:15-alpine
    environment:
      - POSTGRES_PASSWORD=secret
    networks:
      - db-network
    volumes:
      - pgdata:/var/lib/postgresql/data

networks:
  web-network:
    external: true
  db-network:

volumes:
  pgdata:
->
Senior Engineer Analysis of the above Configuration:

  1. Deletion container_name: Docker Rollout must run 2 containers at once during transition. If the name is hardcoded, a conflict will occur. Docker Rollout will automatically rename the containers to [project]-api-1 and [project]-api-2.
  2. Removal of ports and Use of expose: Just like names, host ports cannot be used simultaneously. We only exposeexpose port 8080 into internal networkDocker.
  3. Environment VIRTUAL_HOST: This is magic variable captured by nginx-proxy.
  4. Presence healthcheck: This is the "soul" of zero downtime. Docker Rollout will hold the execution of deleting old containers until endpoint /health consistently returns an HTTP status of 200.

3. Deployment Execution with Docker Rollout

When there is a new version of code that has been build to image [my-registry.com/pos-api:latest](https://my-registry.com/pos-api:latest), the deployment process which previously used docker compose up -d is now changed to:

Bash
# 1. Pull the latest image from the registry
docker compose pull api

# 2. Do a rolling update specifically for the "api" service
docker rollout api
->
Log Output you will see:

Plaintext
Scaling service api to 2 instances...
Waiting for container project-api-2 to become healthy...
Container project-api-2 is healthy!
Stopping old container project-api-1...
Removing old container project-api-1...
Rollout completed successfully!
->
At this point, your application has updated without a single HTTP request getting an error message from Nginx.

Stage 3: "How to Solve" - Solving Problems (Troubleshooting Playbook)

In practice in production, zero downtime does not always run smoothly with just the command above.As a Senior Engineer, you are required to understand how to mitigate various legacy problems. Here is a troubleshooting playbook for complex scenarios.

Problem 1: Request Disconnects Suddenly in the Middle of Transition (Active Connection Drop)

Symptoms:
Even though the new container is up, there are some users who requests (e.g. upload files or queryheavy report) suddenly failed when the transition occurred.

Root Cause:
Docker Rollout kills old containers as soon as new containers healthy. However, Reverse Proxy (Nginx) may take 1-2 seconds to notice routing changes, and legacy containers may still be processing hundreds of request (in-flight requests).

How to Solve: Implement Container Draining with Pre-stop Hook
We have to "cheat" the system. Before the old container is shut down, we have to intentionally change its state to unhealthy. The goal is for Nginx to stop sending NEW traffic to the old container, but let the ONGOING traffic finish processing.

Add instruction pre-stop-hook via label in docker-compose.yml:

YAML
 labels:
      docker-rollout.pre-stop-hook: "touch /tmp/drain && sleep 15"
->
Then, change the /health endpoint logic in your Golang application source code:

Go
http.HandleFunc("/health", func(w http.ResponseWriter, r *http.Request) {
    // If drain file exists, return HTTP 503 (Unhealthy)
    if _, err := os.Stat("/tmp/drain"); err == nil {
        w.WriteHeader(http.StatusServiceUnavailable)
        return
    }
    w.WriteHeader(http.StatusOK)
    w.Write([]byte("OK"))
})
->
How this works:
When docker rollout finishes setting up the new container, it executes pre-stop-hook on the old container.The file /tmp/drain is created. Docker's Healthcheck now detects old containers as unhealthy. Nginx automatically removes old containers from upstream routing. The sleep 15 command gives the old container 15 seconds to complete all in-flight requests before it is completely stopped (SIGTERM). This is absolute Zero Downtime.

Issue 2: Database Migration Locking

Symptoms:
The new version of the application uses a changed database table structure (schema). If the old and new containers run simultaneously during the rollout phase, the old container will crash or produce corrupt data because the database schema has changed, while the codebases don't support it yet.

Root Cause:
A tie between the lifecycle of the application and the lifecycle of the database. Running Auto-Migrate (such as GORM AutoMigrate or Prisma migrate deploy) when the boot application is an anti-pattern on the Zero architecture Downtime.

How to Solve: Backward-Compatible Migrations & Separation of Concerns

  1. Separate execution: Database migration should not be executed inside the application container (via entrypoint). Use a separate container or run it before rollout.

    Bash
    docker compose run --rm api migrate up
    docker rollout api
    
    ->
  2. Backward Compatibility Rules: Your database schema should always support 2 code versions simultaneously (N and N-1).

    • If you want to delete a column, release new code that stops reading/writing that column first (Deploy 1).
    • After the old container dies, then do a migration release that drop that column (Deploy 2).
    • Never rename a column directly.Create a new column, copy the data, read/write from 2 columns, then gradually delete the old column.

Issue 3: Rollout Stuck (Timeout / Infinite Wait)

Symptoms:
Your terminal is stuck at Waiting for container project-api-2 to become healthy... for tens of minutes, or fails with a Timeout error.

Root Cause:

  • Healthcheck URL is incorrect or requires authentication.
  • Application crashes when booting (e.g. due to loss of Environment New variable), so it never reaches the Healthy state.
How to Solve: Debugging & Custom Timeout

  1. New Container Investigation: Do not stop the rollout process yet. Open a new terminal and check the log of the container being booting:

    Bash
    docker logs project-api-2 -f
    
    ->
    Look for stack trace or errors that cause the API to fail to run.
  2. Set Parameters start_period: If your application is a giant Spring Boot Java-based application or a large PHP monolith that takes 30-40 seconds for initial initialization, increase the start_period in your docker-compose.yml file (e.g. start_period: 60s).
  3. Configure Rollout Timeout: By default, docker rollout waits 60 seconds. You can extend the time tolerance via the -t:
    parameter
    Bash
    docker rollout api -t 120
    
    ->

Problem 4: Port Conflict Due to Hardcoded Ports in Codebase

Symptoms:
Docker responds with Error starting userland proxy: listen tcp4 0.0.0.0:8080: bind: address already in use.

Root Cause:
Even though you have removed the ports declaration in the docker-compose.yml file, your containerized application aggressively tries to open conflicting connections at the network stack, or you are running the application in host network mode.

How to Solve:
Make sure network_mode: host is off. Docker Rollout requires network isolation (default bridge mode). Make sure Reverse proxy acts as a mediator bridging external port 80/443 with your application's internal port 8080.

Conclusion

Switching schema Hard Stop (docker-compose up -d) to Zero Downtime Deployment with docker rollout is a massive leap in infrastructure maturity. Adopting this technology not only increases the SLA (Service Level Agreement) of the system you build, but also provides peace of mind (peace of mind) to the Engineering team when they have to deploy new features during busy working hours.

The main key to successful use of Docker Rollout lies in 3 architectural foundations:

  1. Stateless Containers: Avoid storing asset files or sessions inside file system container (use S3 bucket or Redis).
  2. Dynamic Reverse Proxy: Traefik or Nginx-Proxy is a mandatory companion so that new container internal IPs can be detected automatically.
  3. Accurate Healthcheck: Don't just rely on checking ping.Make sure your application's endpoint /health actually checks for database connection readiness and cache before responding with the status 200 OK.
By mastering the tooling and troubleshooting described above, you are ready to bring deployment reliability to the enterprise without having to worry about complicated cluster management. Happy refactoring and enjoy deploy on a drama-free Friday!
Sigit Wasis Subekti

Sigit Wasis Subekti

Software Engineer & Tech Educator

Software Engineer and Tech Educator sharing insights on web development and software architecture.