Server running AI apps with Docker containers on a VPS

How to Run AI Apps with Docker on a VPS

There’s a special kind of pain only AI developers know: it works on your machine, you deploy it to the server, and it explodes. Wrong CUDA version. A Python package that compiled fine locally but refuses on the VPS. A model file that was supposed to be there and isn’t. You fix one thing, break two more, and by midnight you’re rebuilding an environment you built last week.

When you run AI apps with Docker on a VPS, that whole class of problem disappears. The app, its dependencies, its CUDA libraries, and its config all live inside a container that behaves identically on your laptop and your server. Build once, pull on the VPS, and it runs. This guide covers why Docker fits AI workloads, the VPS specs that matter, full setup with GPU passthrough, a real docker-compose stack, and the security and cost realities nobody mentions upfront.

Why Docker for AI Apps? (The Honest Pitch)

AI apps are dependency nightmares. A typical stack pulls in PyTorch or llama.cpp, a specific CUDA toolkit, tokenizers, model weights in the tens of gigabytes, and a dozen pinned Python packages. On a bare server, one careless pip install and two projects start fighting over CUDA libraries. Docker fixes this with boring, reliable isolation:

  • Reproducibility. The container that works on your laptop works on the VPS — same base image, libraries, env vars. “Works on my machine” stops being a joke.
  • Isolation without VMs. Each app gets its own filesystem, Python, and CUDA stack. No virtualenv juggling, no rotting conda environments.
  • GPU passthrough that works. NVIDIA’s container toolkit hands your GPU straight to the container with near-zero overhead — the container sees it like bare metal.
  • Disposable experiments. Try a new model server, hate it, delete it — the host stays pristine. I test several AI tools a month; my VPS has never needed a reinstall.
  • One-command deploys. docker compose up -d starts your whole stack — model server, web UI, database — in order, every time. Reboots and updates become non-events.

The tradeoff: a learning curve, plus a sharp edge or two in GPU setup (covered below). For AI apps, it isn’t close. New to servers? Our VPS hosting primer covers the fundamentals first.

VPS Specs That Matter for AI-in-Docker

Docker itself is light — your specs are driven entirely by what the AI app does inside the container. A chatbot UI sipping from an API needs almost nothing; a local 8B-parameter LLM needs serious RAM; fast inference needs a GPU. For CPU-only AI apps (API-backed chatbots, embedding pipelines, small models), RAM is king, then CPU cores, then disk:

TierWorkloadRAMCPUDiskApprox. cost
StarterChatbot UI, API proxies, n8n automations, small embedding jobs4 GB2 vCPU80 GB SSD$6–12/mo
Sweet spotOllama with 7–8B models, RAG pipelines, Open WebUI + vector DB16 GB4–6 vCPU200 GB NVMe$20–40/mo
Heavy CPU13B+ models on CPU, batch embedding, multiple AI services32–64 GB8–12 vCPU400+ GB NVMe$50–100/mo
GPUFast local inference, image generation, fine-tuning16–32 GB4–8 vCPU200+ GB NVMe$80–300+/mo
Infographic comparing VPS spec tiers for running AI apps in Docker containers, from starter to GPU tiers
Pick your tier by workload, not by hype — most AI apps in Docker run happily on 16 GB RAM.

Local LLMs are RAM-hungry in a way that surprises people: a 7B model in 4-bit quantization needs ~6–8 GB just for weights, plus server and OS overhead. That’s why the sweet spot starts at 16 GB. If serving models is your real goal, read our guide to running AI models on a VPS alongside this one.

Start CPU-only, measure whether inference speed is actually your bottleneck, then pay the GPU tax — our roundup of cheap VPS options that don’t sacrifice performance covers the CPU tiers above. And budget 2x the disk you think you need: AI images run 5–10 GB, model weights 5–50 GB, and Docker’s build cache quietly eats the rest.

Docker vs Bare Metal vs Virtualenv: Which Should You Pick?

ApproachBest forBiggest strengthBiggest weakness
DockerMost AI apps: chatbots, RAG, model servers, multi-service stacksReproducible, isolated, one-command deploysLearning curve; overkill for single tiny scripts
Bare metalMax-performance single workloads, squeezing every token/sec from one modelZero container overhead, simplest debuggingDependency conflicts; messy uninstalls
VirtualenvQuick experiments, single Python scriptsZero setup, familiarDoesn’t isolate system deps (CUDA!); not reproducible across machines

If the app has more than one moving part, or you ever want to run it on a second machine, use Docker. A single API-calling script? Virtualenv is fine. An Ollama server plus a web UI plus a vector database? That’s Docker’s home turf. Bare metal only wins when you’re chasing the last 2–3% of performance on one dedicated workload.

How to Run AI Apps with Docker on a VPS: Step-by-Step Setup

This assumes a fresh Ubuntu 22.04 or 24.04 VPS. If you haven’t hardened a fresh box yet, run through our new VPS setup checklist first — firewall, SSH keys, updates — then follow along as a non-root user with sudo.

Step 1 — Install Docker and the Compose plugin

Skip the distro’s outdated Docker package and use Docker’s official repository. Details change over time, so I’ll point you at Docker’s official Ubuntu install docs as the canonical reference — the short version:

sudo apt update && sudo apt install -y ca-certificates curl
sudo install -m 0755 -d /etc/apt/keyrings
sudo curl -fsSL https://download.docker.com/linux/ubuntu/gpg \
  -o /etc/apt/keyrings/docker.asc
sudo chmod a+r /etc/apt/keyrings/docker.asc

echo "deb [arch=$(dpkg --print-architecture) \
  signed-by=/etc/apt/keyrings/docker.asc] \
  https://download.docker.com/linux/ubuntu \
  $(. /etc/os-release && echo "$VERSION_CODENAME") stable" | \
  sudo tee /etc/apt/sources.list.d/docker.list > /dev/null

sudo apt update
sudo apt install -y docker-ce docker-ce-cli containerd.io \
  docker-buildx-plugin docker-compose-plugin

Then allow Docker without sudo (log out and back in after):

sudo usermod -aG docker $USER
docker --version
docker compose version   # verify both work

Step 2 — GPU support: the NVIDIA Container Toolkit (GPU servers only)

Skip this on a CPU-only VPS. With an NVIDIA GPU on the server, this is what makes it visible inside containers:

curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \
  sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg

curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
  sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
  sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update
sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

The moment of truth — verify the GPU is reachable from inside a container:

docker run --rm --gpus all nvidia/cuda:12.6.0-base-ubuntu24.04 nvidia-smi

If your GPU appears in the output, passthrough works. If not, check troubleshooting first — it’s almost always a driver/toolkit version mismatch, not a Docker problem.

Step 3 — Your first AI container: Ollama

Ollama is the perfect first container: a real AI workload in one command. This pulls the image, starts the server, and keeps it running:

docker run -d \
  --name ollama \
  --restart unless-stopped \
  -p 127.0.0.1:11434:11434 \
  -v ollama-data:/root/.ollama \
  ollama/ollama

Two deliberate choices: -p 127.0.0.1:11434:11434 binds the API to localhost only — never internet-exposed (see security) — and -v ollama-data:/root/.ollama keeps models in a named volume so they survive restarts. On a GPU server, add --gpus all and Ollama uses it automatically.

docker exec ollama ollama pull llama3.1:8b
docker exec ollama ollama run llama3.1:8b "Explain Docker containers in one paragraph."

That pull downloads ~4.7 GB — your first taste of why disk space matters. You’re now running a local LLM in a container on a VPS. Everything next makes this robust and multi-service.

A Real Stack: Ollama + Open WebUI with Docker Compose

Single containers are toys; real AI apps are stacks. The setup I run myself: Ollama serving models plus Open WebUI as a ChatGPT-style interface, wired together with Docker Compose. One YAML file, one command:

mkdir -p ~/ai-stack && cd ~/ai-stack
services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama
    restart: unless-stopped
    ports:
      - "127.0.0.1:11434:11434"
    volumes:
      - ollama-data:/root/.ollama
    # Uncomment the next 2 lines on a GPU server:
    # deploy:
    #   resources:
    #     reservations:
    #       devices:
    #         - driver: nvidia
    #           count: 1
    #           capabilities: [gpu]

  open-webui:
    image: ghcr.io/open-webui/open-webui:main
    container_name: open-webui
    restart: unless-stopped
    ports:
      - "127.0.0.1:3000:8080"
    environment:
      - OLLAMA_BASE_URL=http://ollama:11434
      - WEBUI_SECRET_KEY=${WEBUI_SECRET_KEY}
    volumes:
      - webui-data:/app/backend/data
    depends_on:
      - ollama

volumes:
  ollama-data:
  webui-data:

The services talk over Docker’s internal network via service names (http://ollama:11434) — no host networking hacks. Both ports bind to 127.0.0.1 only; reach the UI through an SSH tunnel or reverse proxy. The secret comes from an env var — create a .env file next to the YAML:

WEBUI_SECRET_KEY=put-a-long-random-string-here
docker compose up -d
docker compose ps        # both services "running"?
docker compose logs -f   # watch startup; Ctrl+C to detach

Open WebUI needs 30–60 seconds on first boot. Reach it via SSH tunnel — ssh -L 3000:localhost:3000 user@your-server-ip — then open http://localhost:3000. Private ChatGPT-style interface, your own models, your VPS. For a chatbot reachable 24/7, our guide to hosting an AI chatbot on a VPS 24/7 takes it further.

Docker Compose terminal output showing Ollama and Open WebUI containers starting together as a connected AI stack
One YAML file, one command — Compose starts your entire AI stack in the right order, every time.

Volumes, Env Files, and Secrets Hygiene

The unglamorous stuff that separates a stack that survives from one that dies on the first update:

  • Volumes for everything stateful. Models, databases, uploads — if losing it would hurt, it lives in a named volume or bind mount, never in the container’s writable layer. docker compose down && docker compose up -d should lose nothing. Back up via cron: docker run --rm -v ollama-data:/data -v ~/backups:/backup ubuntu tar czf /backup/ollama-$(date +%F).tar.gz /data.
  • .env for config, never hardcoded secrets. API keys and passwords go in .env (chmod 600), referenced as ${VAR}. Never commit .env to git.
  • One compose project per app. Don’t stuff your chatbot, n8n automations, and experiments into one giant file — separate directories, separate stacks, so one app’s restart never touches the others. (Our n8n automation server guide makes a natural second stack.)

Resource Limits: Stop One Container from Eating the Server

AI workloads are terrible neighbors. A runaway inference job will eat every byte of RAM, and then the Linux OOM killer starts murdering processes — sometimes your container, sometimes SSH. Set limits so one container can’t take down the host:

services:
  ollama:
    image: ollama/ollama:latest
    deploy:
      resources:
        limits:
          memory: 12G        # hard ceiling: container killed past this
        reservations:
          memory: 8G         # soft guarantee for the scheduler
    # ... rest of service config

Leave 2–4 GB of RAM for the host OS and Docker itself — on a 16 GB box, cap the model server at ~12 GB. A container killed at its limit is a restart; a dead server is a provider-console reboot. Also watch swap: inference on swap runs seconds-per-token slow. If docker stats shows memory pegged and swap climbing, your model is too big — quantize down or size up. No config flag fixes physics.

Lock It Down: Docker Security on a Public VPS

The uncomfortable truth: Docker’s default port publishing bypasses UFW — Docker inserts its own iptables rules ahead of UFW’s, so a perfect firewall can still leave a container port public. I’ve seen unauthenticated AI dashboards exposed for exactly this reason:

  • Bind to localhost by default. Always 127.0.0.1:port:port, never bare port:port. Every compose file here does this — the single highest-value habit.
  • Reverse-proxy anything browser-facing. Nginx or Caddy on the host, terminating HTTPS, proxying to 127.0.0.1:3000. Our Nginx on Ubuntu guide covers setup; add the app’s own login on top.
  • Avoid root inside containers where possible. Many AI images default to root; add user: "1000:1000" where supported. A container escape as root is a host compromise.
  • Keep the Docker socket away from containers. Mounting /var/run/docker.sock grants root-equivalent host control. Say no unless you fully understand the tool asking.
  • Update images regularly, pinned deliberately. docker compose pull && docker compose up -d monthly; pin versions (ollama/ollama:0.5.11, not :latest) so updates are intentional.
  • Firewall + fail2ban on the host, as always. Docker doesn’t replace host hardening — our VPS security guide is the full checklist.

Updating and Backing Up Without Drama

Docker updates are anticlimactic by design — data lives in volumes, so pull, recreate, done:

cd ~/ai-stack
docker compose pull          # fetch new images
docker compose up -d         # recreate containers with new images
docker image prune -f        # delete old image versions

Snapshot before major updates: back up volumes (tar command above) and take a provider snapshot. If the new version breaks things, restore the volume, pin the old image tag, up -d — back in minutes. One gotcha: after editing .env, confirm with docker compose ps that the container actually restarted.

Troubleshooting: The Errors You’ll Actually Hit

Three problems cause ~90% of “Docker AI on VPS” headaches:

“GPU not visible inside the container”

Run nvidia-smi on the host first. Host doesn’t see the GPU? Driver problem — reinstall the NVIDIA driver for your exact kernel. Host sees it but the container doesn’t? Re-run sudo nvidia-ctk runtime configure --runtime=docker && sudo systemctl restart docker and confirm --gpus all is passed. A CUDA 12.x toolkit with ancient drivers is the most common mismatch.

“Port already in use”

Usually a forgotten container: docker ps to find it, then stop it or remap the host side (e.g. 3001 instead of 3000). The left side is yours to choose; the right side is the container’s port — don’t change it blindly.

“No space left on device”

Diagnose with docker system df, then climb the ladder: docker image prune -a (unused images), docker builder prune (build cache, often gigabytes). Last resort docker system prune -a --volumes also deletes unused volumes — back up first.

Containers that won’t start after reboot are usually missing restart: unless-stopped — every service here has it. If one still fails, docker compose logs usually reveals a volume permission issue or a dead env var.

What It Really Costs

  • The VPS: $6–40/month for CPU workloads (tier table above). The main cost, and the only mandatory one.
  • GPU upgrade (if needed): +$80–250/month over a CPU box. Only when CPU inference speed is measurably your bottleneck.
  • Docker itself: $0. Free and open source.
  • Model/API costs: $0 for local open models (Ollama, llama.cpp); $5–50/month if containers call commercial APIs — the sneakiest line item, budget it separately.
  • Backups/snapshots: usually $1–5/month. Cheap insurance.

Realistic total: $10–45/month for a serious CPU-based AI stack, or $100–300/month with a GPU. Managed AI hosting charges per-seat or per-token markups — for 24/7 workloads the VPS pays for itself fast.

When NOT to Use Docker for AI Apps

  • Chasing maximum inference speed on one model. Container overhead is 1–3%, but benchmarkers serving production traffic may want bare metal to remove all doubt.
  • A single tiny script. A 50-line API-calling script on a cron job doesn’t need a container. Virtualenv is simpler, and that’s fine.
  • Exotic kernel needs. Some HPC-style workloads need host kernel modules that don’t containerize cleanly. Rare, but real.
  • A 1–2 GB RAM VPS. Docker’s daemon plus one chunky AI image on a tiny box is just sad. Upgrade first; containerize second.

For the mainstream — chatbots, RAG pipelines, model servers, automation stacks — Docker on a VPS is my default recommendation.

FAQ

Can you run AI apps with Docker on a CPU-only VPS?

Yes — most AI apps don’t need a GPU. Chatbot UIs, API gateways, RAG pipelines, embedding jobs, and quantized 7–8B local models all run fine on CPU. A GPU matters only when inference speed or image generation becomes the bottleneck.

How much RAM do I need to run Ollama in Docker?

Budget the model’s quantized size plus ~2 GB overhead. A 7–8B model needs ~8–10 GB total, so 16 GB is the comfortable minimum. A 13B model wants 16+ GB for itself — that’s 32 GB territory.

Is Docker slower than running AI apps directly on the VPS?

Overhead is typically 1–3%, often unmeasurable, for CPU/memory-bound work; GPU passthrough is near-native. The operational wins — reproducibility, easy updates, isolation — dwarf the cost for nearly everyone.

How do I expose my AI app’s web UI safely?

Bind container ports to localhost (127.0.0.1:3000:8080), then reverse-proxy with Nginx or Caddy over HTTPS. Never publish AI dashboards to 0.0.0.0 — many ship without auth, and bots scan for them constantly.

What’s the difference between docker run and docker compose for AI apps?

docker run starts one container from command-line flags — fine for quick tests. docker compose defines multi-container stacks in YAML, handling networking and startup order, making the setup reproducible. Use compose for anything you’ll run more than once.

Ship It: Your AI Stack in an Afternoon

Your first AI stack takes an afternoon — install Docker, wire up the NVIDIA toolkit with a GPU, get Ollama and Open WebUI talking through Compose. The second takes twenty minutes: just another YAML file.

Start today: spin up a 16 GB box, run the Ollama container from Step 3, pull a small model. Then graduate to the Compose stack — memory limits set, ports bound to localhost, volumes backed up — and you’ll have a private, always-on AI setup for less than a takeaway dinner a month. Happy shipping.

Leave a Comment

Your email address will not be published. Required fields are marked *