Ollama Installation: Local LLM Running Guide
Ollama installation, running local LLM models (Llama 3, Mistral, Gemma), API usage, model management, GPU acceleration and Docker deployment.
NVIDIA driver installation, CUDA toolkit, nvidia-container-toolkit, model serving with vLLM, GPU monitoring and AI model deployment with Docker guide.
GPU servers are essential for running AI models in production environments. This guide covers everything from NVIDIA driver installation to CUDA configuration, Docker GPU integration, and high-performance model serving with vLLM.
| Component | Minimum | Recommended |
|---|---|---|
| GPU | NVIDIA T4 (16 GB) | NVIDIA A100 (40/80 GB) |
| RAM | 32 GB | 64 GB+ |
| Disk | 100 GB SSD | 500 GB NVMe |
| OS | Ubuntu 22.04 LTS | Ubuntu 22.04/24.04 LTS |
| CUDA | 12.0+ | 12.4+ |
First, clean existing drivers and install the new one:
# Remove existing NVIDIA drivers
sudo apt-get purge nvidia-* -y
sudo apt-get autoremove -y
# System update
sudo apt-get update && sudo apt-get upgrade -y
# Add NVIDIA driver repository
sudo apt-get install -y linux-headers-$(uname -r)
sudo apt-get install -y software-properties-common
sudo add-apt-repository ppa:graphics-drivers/ppa -y
sudo apt-get update
# Check recommended driver
ubuntu-drivers devices
# Install NVIDIA driver
sudo apt-get install -y nvidia-driver-550
sudo reboot
Verify after installation:
# Driver version and GPU info
nvidia-smi
Output:
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 550.54.15 Driver Version: 550.54.15 CUDA Version: 12.4 |
|-------------------------------+----------------------+----------------------+
| GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |
|===============================+======================+======================|
| 0 NVIDIA A100-SXM4-80GB On| 00000000:00:04.0 Off | 0 |
| N/A 32C P0 45W / 400W| 0MiB / 81920MiB | 0% Default |
+-------------------------------+----------------------+----------------------+
# Add CUDA keyring
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt-get update
# Install CUDA Toolkit
sudo apt-get install -y cuda-toolkit-12-4
# PATH configuration
echo 'export PATH=/usr/local/cuda-12.4/bin:$PATH' >> ~/.bashrc
echo 'export LD_LIBRARY_PATH=/usr/local/cuda-12.4/lib64:$LD_LIBRARY_PATH' >> ~/.bashrc
source ~/.bashrc
# Verify
nvcc --version
# cuDNN installation (for CUDA 12.x)
sudo apt-get install -y libcudnn8 libcudnn8-dev
# Verify
cat /usr/include/cudnn_version.h | grep CUDNN_MAJOR -A 2
The nvidia-container-toolkit is required to use GPUs in Docker containers:
# NVIDIA Container Toolkit repository
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \
sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update
# Installation
sudo apt-get install -y nvidia-container-toolkit
# Docker runtime configuration
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
# Test
sudo docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
vLLM is a high-throughput LLM inference engine. It optimizes memory usage with PagedAttention technology.
docker run -d \
--name vllm \
--gpus all \
-p 8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env HUGGING_FACE_HUB_TOKEN=$HF_TOKEN \
vllm/vllm-openai:latest \
--model meta-llama/Meta-Llama-3.1-8B-Instruct \
--max-model-len 4096 \
--gpu-memory-utilization 0.9
version: '3.8'
services:
vllm:
image: vllm/vllm-openai:latest
container_name: vllm-server
ports:
- "8000:8000"
volumes:
- huggingface_cache:/root/.cache/huggingface
environment:
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN}
command: >
--model meta-llama/Meta-Llama-3.1-8B-Instruct
--max-model-len 4096
--gpu-memory-utilization 0.9
--tensor-parallel-size 1
--dtype auto
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
restart: unless-stopped
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
volumes:
huggingface_cache:
vLLM provides an OpenAI-compatible API:
# Chat completion
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Meta-Llama-3.1-8B-Instruct",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is Docker?"}
],
"max_tokens": 512,
"temperature": 0.7
}'
# Model list
curl http://localhost:8000/v1/models
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="not-needed"
)
response = client.chat.completions.create(
model="meta-llama/Meta-Llama-3.1-8B-Instruct",
messages=[
{"role": "system", "content": "You are a Linux expert."},
{"role": "user", "content": "How to configure RAID?"}
],
max_tokens=1024,
temperature=0.7
)
print(response.choices[0].message.content)
# Tensor parallelism with multiple GPUs
docker run -d \
--name vllm-multi \
--gpus all \
-p 8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:latest \
--model meta-llama/Meta-Llama-3.1-70B-Instruct \
--tensor-parallel-size 4 \
--max-model-len 4096 \
--gpu-memory-utilization 0.9
# Current status
nvidia-smi
# Continuous monitoring (every 1 second)
watch -n 1 nvidia-smi
# GPU usage and memory only
nvidia-smi --query-gpu=index,name,utilization.gpu,memory.used,memory.total,temperature.gpu --format=csv -l 1
# docker-compose.monitoring.yml
version: '3.8'
services:
dcgm-exporter:
image: nvcr.io/nvidia/k8s/dcgm-exporter:3.3.5-3.4.0-ubuntu22.04
container_name: dcgm-exporter
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
ports:
- "9400:9400"
restart: unless-stopped
prometheus:
image: prom/prometheus:latest
container_name: prometheus
ports:
- "9090:9090"
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
restart: unless-stopped
grafana:
image: grafana/grafana:latest
container_name: grafana
ports:
- "3000:3000"
environment:
- GF_SECURITY_ADMIN_PASSWORD=admin
restart: unless-stopped
# Enable persistence mode (lower latency)
sudo nvidia-smi -pm 1
# Lock GPU clock speed
sudo nvidia-smi -lgc 1410,1410
# Set power limit
sudo nvidia-smi -pl 300
server {
listen 443 ssl;
server_name ai-api.example.com;
ssl_certificate /etc/letsencrypt/live/ai-api.example.com/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/ai-api.example.com/privkey.pem;
location / {
proxy_pass http://127.0.0.1:8000;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_read_timeout 300s;
proxy_send_timeout 300s;
}
}
Run your AI models with high performance on REXE GPU servers with NVIDIA A100/H100 GPUs. Contact us for professional GPU server solutions.
For 7B-13B parameter models, NVIDIA T4 (16 GB) or RTX 4090 (24 GB) is sufficient. For 70B models, A100 (40/80 GB) or H100 (80 GB) is recommended. VRAM capacity is the most critical factor.
Ollama is ideal for easy setup and development environments. vLLM provides high throughput and low latency in production environments. vLLM offers much more efficient memory usage with continuous batching and PagedAttention.
Docker containers cannot access GPUs by default. nvidia-container-toolkit adds GPU passthrough support to the Docker runtime. This allows GPU usage from within containers using the --gpus all parameter.
Yes, multiple models can run on the same GPU as long as VRAM is sufficient. In vLLM, you can set memory limits for each model with the --gpu-memory-utilization parameter. With MIG (Multi-Instance GPU), you can physically partition an A100 GPU.
Yes, GPU access is interrupted during driver updates and running containers are affected. Stop containers before updating, update the driver, restart the system, and restart the containers.
Reduce model size (use quantized version), lower --max-model-len value, set --gpu-memory-utilization between 0.85-0.9, or use a GPU with more VRAM.
Ollama installation, running local LLM models (Llama 3, Mistral, Gemma), API usage, model management, GPU acceleration and Docker deployment.
Open WebUI installation with Docker, Ollama integration, user management, RAG (document querying), model management, custom prompts and multi-user support.
Stable Diffusion installation, AUTOMATIC1111 and ComfyUI interfaces, model downloading, LoRA training, ControlNet usage and Docker self-hosted image generation guide.