Here is a scenario I see all the time. Someone builds a genuinely good AI chatbot — trained on their business docs, wired into their website — and then hosts it on their laptop. It works beautifully. Until they close the lid. Or the Wi-Fi drops. Or Windows decides 3 AM is the perfect time for an update.
And just like that, the chatbot is down. Visitors get a dead widget. Leads go nowhere.
The fix is simple: put the chatbot on a VPS that never sleeps. A small virtual server, running 24/7, serving your AI to the world while you do literally anything else. This guide shows you exactly how to host an AI chatbot on a VPS 24/7 — the model, the setup, the auto-restart plumbing, the web interface, and the security bits most tutorials skip. No fluff, just the working path.
Table of Contents
What “24/7” Actually Means (Be Honest With Yourself)
Before we touch a terminal, let us define the target. “24/7” does not mean your chatbot literally never goes down — even Google has outages. It means three things:
- It survives reboots. When the host reboots your VPS for maintenance at 4 AM, everything comes back up on its own. No SSH-ing in half-asleep to restart things.
- It recovers from crashes. If the model process dies, something restarts it within seconds — automatically.
- You know when it is down. Silent downtime is the real killer. A free uptime monitor that pings you beats hoping for the best.
That is the bar. Everything in this guide exists to clear it. And the honest truth? Getting there takes about an hour, and most of that is waiting for downloads.
What You Will Need
Nothing exotic. Here is the full shopping list:
- A VPS — 4 vCPUs, 8 GB RAM, 80 GB NVMe SSD is the sweet spot for a single chatbot (about $6–12/month). If you expect real traffic or want a smarter 13B model, go 8 vCPUs / 16 GB. Our VPS specs guide for AI websites has the full sizing math.
- Ubuntu 22.04 or 24.04 — every command below assumes it. Do not use some obscure distro to feel clever; you want Stack Overflow answers to match your system.
- A domain name (optional but recommended) —
chat.yourdomain.comlooks infinitely more trustworthy than an IP address. Costs ~$10/year. - About an hour — mostly waiting on model downloads.
That is it. No GPU needed — a quantized 7B or 8B model runs fine on CPU for chat workloads, answering at 10–30 tokens per second. Fast enough that visitors never notice.
Step 1: Set Up Your VPS

SSH into your fresh server. First, the boring-but-critical basics — updates, a firewall, and a non-root user. I know, I know, you want to get to the AI part. But an AI chatbot with an open port and a weak root password will be crypto-mining for a stranger by Thursday. Our VPS security guide goes deeper; here is the minimum:
apt update && apt upgrade -y
apt install -y ufw fail2ban
ufw allow OpenSSH
ufw allow 80,443/tcp
ufw --force enable
That gives you a patched server with only SSH, HTTP, and HTTPS reachable. If you have not done the rest of the new-server checklist (SSH keys, automatic updates), do it now — future-you will be grateful.
Step 2: Install Ollama and Pull a Model
Ollama is the easiest way to run open models on a server. One install command, one command per model, and you have an API serving completions. If you want the full background on how self-hosting models works, our complete guide to running AI models on a VPS covers it in depth.
curl -fsSL https://ollama.com/install.sh | sh
ollama pull llama3.1:8b
That downloads Meta’s Llama 3.1 8B — smart enough for customer support, FAQs, and document Q&A, small enough (~5 GB) to live comfortably in 8 GB of RAM. Want something else? ollama pull qwen2.5:7b is excellent for multilingual chat; mistral:7b is fast and sharp in English.
Quick sanity check that it works:
ollama run llama3.1:8b "Reply in one sentence: why do servers need uptime monitoring?"
If it answers, your model is alive. Now let us make sure it stays alive.
Step 3: Make It Survive Reboots (systemd)
This is the step most tutorials skip — and it is the entire difference between “a chatbot” and “a chatbot that runs 24/7.” Ollama installs a systemd service by default on most setups, but let us verify and harden it, because silent failures here are how you get the 4 AM outage.
systemctl enable ollama
systemctl start ollama
systemctl is-enabled ollama # must print: enabled
enable is the magic word — it means “start this automatically on every boot.” Now add a safety net: if the process crashes, systemd should restart it instead of giving up. Create an override:
mkdir -p /etc/systemd/system/ollama.service.d
cat > /etc/systemd/system/ollama.service.d/restart.conf <<'EOF'
[Service]
Restart=always
RestartSec=5
EOF
systemctl daemon-reload
systemctl restart ollama
Five lines, and your model now restarts itself within 5 seconds of any crash, and comes back after every reboot. That is 90% of “24/7” right there.
Step 4: Give It a Web Interface (Open WebUI)
Ollama alone is an API — great for developers, useless for visitors. Open WebUI is a polished ChatGPT-style interface that talks to Ollama. Your visitors get a familiar chat window; you get user accounts, chat history, and admin controls.
Install Docker, then run it:
apt install -y docker.io docker-compose-plugin
docker run -d --name open-webui \
--restart unless-stopped \
-p 3000:8080 \
-e OLLAMA_BASE_URL=http://host.docker.internal:11434 \
--add-host=host.docker.internal:host-gateway \
-v open-webui:/app/backend/data \
ghcr.io/open-webui/open-webui:main
Two details worth noticing. --restart unless-stopped means the web UI also comes back after reboots — same 24/7 principle, applied to every layer. And the data volume (-v open-webui:/app/backend/data) means user accounts and chat history survive container updates. I have seen people skip the volume, update the container, and lose everything. Do not be those people.
Visit http://your-server-ip:3000 — you should see the Open WebUI signup page. Create your admin account (the first account becomes admin automatically). But do not stop here: right now it is HTTP on a raw IP. Let us fix that.
Step 5: Put It Behind a Domain With Free HTTPS
Nobody trusts a chatbot on http://203.0.113.42:3000. Point a DNS A-record for chat.yourdomain.com at your server’s IP, then install Caddy — a web server that gets free Let’s Encrypt certificates automatically, with roughly four lines of config:
apt install -y caddy
Then edit /etc/caddy/Caddyfile:
chat.yourdomain.com {
reverse_proxy localhost:3000
}
systemctl reload caddy
That is genuinely all of it. Caddy fetches the certificate, renews it forever, and proxies visitors to Open WebUI over HTTPS. No certbot cron jobs, no manual renewals, no expiry panic. If you prefer Nginx, our Nginx on Ubuntu guide covers the manual route — but for a chatbot, Caddy’s simplicity wins.
Step 6: Monitoring and Auto-Recovery

You now have a chatbot that restarts itself and serves over HTTPS. The last piece of 24/7 is knowing — because “it restarted itself” is only comforting if you can verify it.
1. Free external uptime monitoring. Sign up for UptimeRobot’s free plan (50 monitors, plenty) and add https://chat.yourdomain.com. It pings every 5 minutes and emails you the moment the site stops responding. This is non-negotiable — it is the difference between finding out from a monitor in 5 minutes and finding out from an angry customer in 5 days.
2. Watch disk space. Chat logs and Docker images grow quietly. A full disk kills everything at once. Add a weekly check to your routine — or better, a one-line cron that warns you:
# Warn if disk usage passes 80%
0 9 * * * df / | awk 'NR==2 && $5+0 > 80 {print "Disk warning: "$5" used"}' | mail -s "VPS disk alert" you@example.com
3. Keep an eye on memory. If your model plus Open WebUI plus the OS creep past physical RAM, Linux starts swapping and response times fall off a cliff. Run free -h occasionally; if swap is anything above zero for more than a few minutes, you need a bigger plan — not a tweak. The specs guide tells you exactly which tier to move to.
4. Update on a schedule. Pick one quiet day a month: apt upgrade, pull a fresh Open WebUI image, restart, verify the chat loads. Boring, regular maintenance beats heroic firefighting every time.
How Much Does This Actually Cost?
Let us be concrete, because “cheap” means nothing without numbers:
| Component | Typical cost |
|---|---|
| VPS — 4 vCPU, 8 GB RAM, 80 GB NVMe | $6–12/month |
| VPS — 8 vCPU, 16 GB RAM (busier sites) | $24–48/month |
| Domain name | ~$10/year |
| Ollama + Open WebUI + Caddy | $0 (all open source) |
| UptimeRobot monitoring | $0 (free plan) |
| HTTPS certificates | $0 (Let’s Encrypt via Caddy) |
So a real, production, 24/7 AI chatbot: roughly $7–13/month all-in. Compare that to hosted chatbot platforms charging $50–500/month, and the VPS route pays for itself before the first invoice. If you are hunting for the lowest viable price, our cheapest VPS guide shows how to cut costs without cutting reliability.
Security: The Part People Skip (Do Not)
An AI chatbot is an attractive target — it burns CPU, which attackers love to borrow. Minimum viable security:
- Disable Ollama’s public exposure. Ollama listens on localhost by default — keep it that way. Only Caddy (port 443) should face the internet. Never expose port 11434 publicly; an open inference API will be found within days and abused.
- Require login. In Open WebUI’s admin settings, disable public signup after creating your accounts. An open-registration chatbot is a free GPU farm for strangers.
- SSH keys only. Disable password auth entirely (
PasswordAuthentication noin sshd_config). Fail2ban is already installed from Step 1 — verify it is running. - Set rate limits. If your chatbot is public, add basic rate limiting in Caddy or Cloudflare so one user cannot monopolize your CPU.
None of this is paranoia — it is the baseline. A chatbot that answers visitors at 3 AM is wonderful; a chatbot that mines crypto for someone else at 3 AM is a very expensive lesson.
Troubleshooting: When Something Breaks
“The chat page won’t load.” Check the layers bottom-up: systemctl status ollama → docker ps (is open-webui running?) → systemctl status caddy. Nine times out of ten, one of the three is down, and the status output tells you which.
“It was working, now it’s slow.” Run free -h. If swap is in use, you have outgrown your RAM — upgrade the plan. If load average exceeds your vCPU count persistently, same answer for CPU.
“Caddy says certificate error.” Your DNS A-record is probably wrong or still propagating. dig chat.yourdomain.com should return your server IP. DNS changes can take a few hours — this one is almost always just patience.
“The model gives garbage answers.” That is a model/prompt problem, not a hosting problem. Try a different model (ollama pull qwen2.5:7b) or tune the system prompt in Open WebUI’s settings. Hosting can only serve the brain you give it.
FAQs
Can I host the chatbot and my website on the same VPS?
Yes, if the specs allow it. A chatbot (4 vCPU / 8 GB) plus a small WordPress site fits comfortably on an 8 vCPU / 16 GB server. Just remember they share RAM — size for the sum, not each in isolation.
Do I need a GPU for a 24/7 chatbot?
No. CPU inference on a quantized 7B–8B model handles typical chat traffic fine. You would only add a GPU for image generation or very high concurrency — and that multiplies your hosting bill roughly 5–10×.
What happens when my VPS provider does maintenance?
The server reboots, and everything we set to auto-start (Ollama via systemd, Open WebUI via --restart unless-stopped, Caddy via systemd) comes back on its own. Your uptime monitor might catch a 2–3 minute gap. That is normal and fine.
How do I update the model later?
ollama pull llama3.1:8b again pulls the latest version, then switch to it in Open WebUI’s model picker. Old versions stay cached until you run ollama rm — handy for instant rollback if a new version misbehaves.
Conclusion: How to Host an AI Chatbot on a VPS 24/7
Hosting an AI chatbot on a VPS 24/7 comes down to six moves: a right-sized server, Ollama serving a quantized model, systemd keeping it alive across reboots and crashes, Open WebUI giving visitors a real chat interface, Caddy wrapping it in free HTTPS, and a free monitor telling you the moment something breaks. Total cost: about the price of two coffees a month. Total maintenance: maybe an hour a month.
The alternative — a laptop that sleeps, a free tier that throttles, a hosted platform that bills per message — always costs more in the end, in money or in 3 AM surprises. Build it on your own VPS, set the auto-restart plumbing once, and your chatbot will be answering visitors while you sleep. Which is, after all, the entire point. If you later want that chatbot to act on its own as a 24/7 AI agent, the same server is already the right foundation.


