Setting Up a Self-Hosted AI Chatbot on a Raspberry Pi: Step by Step

Dileep Solanki


I've been interested in self-hosted AI for a while, but there's something particularly appealing about running a chatbot on a computer that costs a fraction of a typical AI workstation.

That's where the Raspberry Pi comes in.

A Raspberry Pi isn't going to compete with a high-end NVIDIA GPU for running large language models. That's not really what I'm trying to do here.

The interesting part is that a Raspberry Pi can act as a small, private AI server for lightweight models. With a local model runner such as Ollama and a web interface such as Open WebUI, you can turn the Pi into a chatbot that you access from another computer or phone on your network.

No monthly chatbot subscription is required for the local model itself.

And, importantly, your prompts can stay on your own hardware.

In this guide, I'll walk through the setup from a fresh Raspberry Pi OS installation to a working browser-based chatbot.


What We're Building

The final setup looks like this:

Your phone or laptop → Open WebUI → Ollama → Local AI model

Open WebUI provides the interface.

Ollama manages and runs the language model.

The Raspberry Pi provides the hardware.

The model itself runs locally rather than sending every conversation to a cloud AI provider.

Open WebUI officially supports ARM64 systems, including Raspberry Pi, and provides Docker-based installation instructions.


What You'll Need

For this project, I'd recommend using a reasonably recent Raspberry Pi rather than an older model.

You'll need:

  • Raspberry Pi 4 or newer
  • Preferably 8GB RAM if available
  • A microSD card or other suitable storage
  • Power supply
  • Network connection
  • Raspberry Pi OS 64-bit
  • Docker
  • Ollama
  • Open WebUI

The 64-bit version is important here.

Current Raspberry Pi OS releases provide official 64-bit builds for several Raspberry Pi models, including the Pi 4 and Pi 5. Raspberry Pi's official documentation also provides Raspberry Pi Imager for installing the operating system.

I'd also recommend Ethernet if you have it.

Wi-Fi works, but if the Pi is going to act as a little AI server that you access regularly, a wired connection is generally the nicer option.


Step 1: Install Raspberry Pi OS

The easiest way to get started is with Raspberry Pi Imager.

Install Raspberry Pi Imager on your computer and insert the microSD card.

Select:

Raspberry Pi → Raspberry Pi OS → 64-bit

For a server-like setup, Raspberry Pi OS Lite is worth considering because it doesn't include the desktop environment and has a smaller footprint. Raspberry Pi currently provides an official 64-bit Lite release.

Before writing the image, configure your username, password and network settings if you're using the Imager's customization options.

Then write the operating system to the card.

Put the card into the Raspberry Pi and boot it.


Step 2: Update the Pi

Once you're connected to the Raspberry Pi, update the operating system.

Run:

sudo apt update
sudo apt full-upgrade -y

Then reboot:

sudo reboot

After the Pi comes back online, reconnect to it.

Keeping the operating system updated is particularly important when you're turning a small computer into a network service.


Step 3: Check That You're Running 64-bit Linux

Before installing the AI stack, check the architecture:

uname -m

Ideally, you'll see:

aarch64

That's the 64-bit ARM architecture.

If you see armv7l instead, you're running a 32-bit operating system.

That's not what I'd use for this project.

Current Docker documentation specifically recommends the 64-bit ARM route for Raspberry Pi OS, while support for 32-bit Raspberry Pi OS is being phased down in newer Docker releases.


Step 4: Install Docker

Open WebUI's official documentation recommends Docker as the easiest installation method for most users.

Docker's current documentation supports ARM64 installations on Debian-based systems.

For a test or personal project, Docker's installation method can be used to get Docker running on the Pi.

After installation, verify it:

sudo docker run hello-world

If everything is working, Docker should download a small test image and display a confirmation message.

That's our first checkpoint.


Step 5: Install Ollama

Now we need something that can actually run the language model.

That's where Ollama comes in.

Ollama provides a local API and command-line interface for running language models.

Once Ollama is installed, check that it's available:

ollama --version

The exact installation command can change over time, so I'd use the current Linux installation instructions from Ollama rather than copying an old tutorial's installer command.

The important thing is that we want Ollama running as the local model server.


Step 6: Start With a Small Model

This is where I would resist the temptation to download the biggest model I can find.

A Raspberry Pi is a small ARM computer.

It doesn't have the memory bandwidth or GPU acceleration of a desktop AI workstation.

So start small.

For example:

ollama run llama3.2:1b

Ollama currently lists Llama 3.2 in 1B and 3B text-only versions. The 1B model is listed at about 1.3 GB in its standard tag, while the 3B version is about 2.0 GB.

You can also use a specific quantized variant if you want to experiment with a smaller model footprint.

The first download can take a while depending on your internet connection.

After it's downloaded, Ollama should open an interactive prompt.

Try:

Hello. Explain what a Raspberry Pi is in simple terms.

If you get a response, congratulations.

You already have a local AI model running.


Step 7: Check Ollama's API

Ollama also exposes an API.

You can test it locally with:

curl http://localhost:11434/api/chat \
  -d '{
    "model": "llama3.2:1b",
    "messages": [
      {
        "role": "user",
        "content": "Say hello in one sentence."
      }
    ]
  }'

Ollama documents this API format for its models.

If the Pi returns a response, the backend is working.

This is important because Open WebUI will communicate with Ollama through this API.


Step 8: Install Open WebUI

Now we can make the setup much nicer.

Instead of interacting with the model through a terminal, we'll use a browser.

Open WebUI provides a self-hosted interface for AI models and officially supports ARM64 systems, including Raspberry Pi.

Before launching it, generate a secret key:

openssl rand -hex 32

Copy the result somewhere safe.

Then start Open WebUI:

docker run -d \
  -p 3000:8080 \
  --add-host=host.docker.internal:host-gateway \
  -v open-webui:/app/backend/data \
  -e WEBUI_SECRET_KEY=YOUR_SECRET_KEY \
  --name open-webui \
  --restart always \
  ghcr.io/open-webui/open-webui:main

Replace:

YOUR_SECRET_KEY

with the value generated by openssl.

Open WebUI's current documentation uses this general Docker configuration and recommends keeping its persistent data volume.


Step 9: Open Your Chatbot

Find your Raspberry Pi's IP address:

hostname -I

You might see something like:

192.168.1.50

Now open a browser on another computer connected to the same network.

Go to:

http://192.168.1.50:3000

Obviously, replace the IP address with the address of your Pi.

You should see the Open WebUI interface.

Create the administrator account when prompted.


Step 10: Connect Open WebUI to Ollama

If Ollama is running directly on the Raspberry Pi and Open WebUI is running inside Docker, Open WebUI needs to know where Ollama is.

The Docker command above includes:

--add-host=host.docker.internal:host-gateway

This allows the container to reach services running on the host.

Open WebUI's documentation explains that its Docker setup can connect to Ollama through:

http://host.docker.internal:11434

depending on the deployment configuration.

In Open WebUI, go to:

Settings → Admin → Connections

Then check the Ollama connection.

If necessary, set the Ollama URL to:

http://host.docker.internal:11434

Once the connection works, your installed Ollama models should become available in Open WebUI.


Step 11: Start Your First Chat

Create a new chat.

Choose:

llama3.2:1b

Then ask something simple:

Explain how DNS works to a beginner.

If the model responds, the complete stack is working.

You now have:

Raspberry Pi

↓

Ollama

↓

Llama 3.2

↓

Open WebUI

↓

Browser

And all of it is running on your own hardware.


What Can You Actually Use This For?

This is where expectations matter.

I wouldn't expect a Raspberry Pi to replace ChatGPT, Claude or Gemini for every task.

That's not the goal.

A small local model can still be useful for things like:

Private notes

You can experiment with asking a local model questions about information you don't want to send to a cloud service.

Simple writing

Short rewrites, summaries and brainstorming can work well with smaller models.

Learning

It's an excellent way to learn how local LLM systems actually work.

Instead of just typing prompts into a website, you can see the components involved.

Small internal tools

You can build simple applications that communicate with Ollama's local API.

Experimentation

You can swap models and compare their behaviour without changing the basic interface.


Where the Raspberry Pi Starts to Struggle

This is the part I wouldn't hide.

The Raspberry Pi is not a high-performance AI workstation.

Large language models require significant memory and compute.

A small 1B model is a very different workload from a large 30B or 70B model.

As models get larger, response times can become frustrating on CPU-only hardware.

The Pi also has limited RAM compared with modern AI desktops.

So I would think of the Raspberry Pi as a small local AI server, not an AI powerhouse.

That's actually what makes it interesting.

You're trading raw performance for:

  • Low power consumption
  • Small size
  • Local control
  • Privacy
  • Always-on availability
  • Low operating cost


Don't Expose It Directly to the Internet

This is important.

If you're just experimenting, keep the chatbot accessible only from your local network.

Don't simply forward port 3000 from your router to the Raspberry Pi and assume the login page is enough security.

If you eventually want remote access, use a properly secured approach such as a VPN or another authenticated access layer.

Also keep the Raspberry Pi and Docker environment updated.

A self-hosted service becomes your responsibility once you put it on a network.


One More Thing: Storage

AI models aren't tiny.

Even relatively small models can take hundreds of megabytes or several gigabytes depending on their size and quantization.

If you download multiple models, your storage requirements grow quickly.

That's why I wouldn't buy the smallest possible storage option if I already know I'm going to experiment with several models.

An external SSD can also make sense for a more serious setup.


Is It Worth Doing?

For me, the interesting part isn't whether the Raspberry Pi is the fastest machine for AI.

It isn't.

The interesting part is that you can take a tiny computer, install an open operating system, run a local model and access your own AI chatbot from a browser.

That changes the relationship you have with the technology.

Instead of simply being an AI user, you're running part of the AI stack yourself.

You can see the server.

You can see the API.

You can change the model.

You can modify the interface.

And you can build your own applications on top of it.

That's a much more interesting learning experience than simply opening another AI website.


Final Setup

At the end, your setup should look roughly like this:

                 Your laptop / phone
                         │
                         ▼
                  Open WebUI
                    :3000
                         │
                         ▼
                      Ollama
                   :11434 API
                         │
                         ▼
                Local AI Model
                 Llama 3.2 1B
                         │
                         ▼
                  Raspberry Pi

The whole thing fits on a small computer sitting quietly on your desk.

And that's probably the best part.

You don't need a massive server to start experimenting with local AI.

You just need realistic expectations about what the hardware can handle.

Start with a small model.

Get the basic stack working.

Learn how Ollama and Open WebUI communicate.

Then experiment.

Because once you've built your own local chatbot, the next question becomes much more interesting:

What else can you build on top of it?

3/related/default