GPU compatibility: built and tuned around a 16GB card, but the underlying mechanics (Steps 1–5: Proxmox passthrough, Secure Boot handling, Docker, NVIDIA Container Toolkit) work unchanged on any Turing-or-newer RTX GPU — RTX 20/30/40/50-series. CUDA 13.x (installed in Step 1g) dropped support for anything older than Turing, so GTX 10-series and earlier need a different, older CUDA branch not covered here. On different VRAM, the model choices in Steps 6–8 need adjusting, not the setup steps themselves:
- 8GB — use the 7B coding model as primary (not fallback), FLUX.1 klein instead of FLUX.1-dev
- 12GB — this guide's picks mostly fit, but with less headroom than on 16GB
- 16GB (this guide) — Qwen2.5-Coder 14B + FLUX.1-dev + Wan 2.2 14B GGUF Q4, as written
- 24GB+ — everything here fits with room to spare; a 30B-class coding model becomes realistic, and the "switch sessions" VRAM juggling in this guide may not be necessary at all
Root install folder: /usr/src/ai
Service account: skynetai — a normal, non-root user with sudo rights. Every command below runs as this user; sudo is used only for the specific steps that genuinely need elevation. Nothing in this stack runs logged in as root, and no daemon here runs as root either — see the "who runs as what" table at the end.
Target network: 192.168.1.0/24 — used below for firewall scoping.
Ports (all above 20000, all listening on 0.0.0.0):
| Service | Port |
|---|---|
| Ollama API | 21434 |
| ComfyUI | 21188 |
| SearXNG | 21081 |
| Open WebUI | 21080 |
Stack:
- Ollama — serves your coding LLM (Qwen2.5-Coder 14B)
- NVIDIA Container Toolkit — lets Docker containers use your GPU
- ComfyUI — image (FLUX.1-dev) and video (Wan 2.2) generation engine
- SearXNG — self-hosted, private web search backend
- Open WebUI — unified chat interface: talks to Ollama, auto-decides when to search the web via SearXNG, and generates images via ComfyUI directly from chat
- Aider — autonomous terminal coding agent, git-aware, uses the same Ollama model (client only, doesn't listen on any port)
Install order, and why: Step 1 provisions the VM and OS itself, including GPU passthrough and CUDA. Steps 2–3 set up the service account and folder structure. Ollama, ComfyUI, and SearXNG (Steps 6–9) are each installed and verified working on their own. Open WebUI (Step 10) goes in last, since it's the piece that connects to all three — by the time you set up its integrations, every backend it needs is already running.
A note on step numbering: this guide previously started from "assume Ubuntu Server + CUDA are already installed." Step 1 below now actually builds that assumption from scratch — provisioning the VM in Proxmox, installing Ubuntu 26.04, and finishing with GPU passthrough and CUDA — which necessarily comes before creating the skynetai user (you can't create a Linux user before the Linux install exists). Every step from the old "Step 0" onward has shifted down by two to make room for it.
VRAM rule of thumb: run the coding stack (Ollama) OR the creative stack (ComfyUI) loaded at a time, not both — see "Switching between sessions" near the end. The Open WebUI container itself (:cuda image tag, Step 10) also keeps a small embedding model on the GPU for RAG/search — a few hundred MB, running continuously in the background regardless of which session you're in.
Integration map (what talks to what):
Open WebUI (21080) ──chat/completions──> Ollama (21434)
Open WebUI (21080) ──web search────────> SearXNG (21081)
Open WebUI (21080) ──image generation──> ComfyUI (21188)
ComfyUI (21188) ──video generation─── Wan 2.2 (used directly in ComfyUI's own UI — see note in Step 8b)
Aider ──chat/completions──> Ollama (21434) [separate CLI, not routed through Open WebUI]
Step 1 — Provision Ubuntu Server 26.04 in Proxmox, then finish OS setup
Everything in this step happens before skynetai exists — you'll be working as whatever admin account already has access to your Proxmox host, and later as the account you create during the Ubuntu installer.
Order matters here, deliberately: create the VM → install the OS → update the system → disable the default video driver → then GPU passthrough → then CUDA, last. Passthrough is placed right before CUDA rather than at the very start, since CUDA is the one piece that actually needs the GPU physically present and working — everything before it (VM creation, OS install, updates, driver blacklisting) can happen without the GPU attached at all.
1a — Download the Ubuntu Server 26.04 LTS ISO
Ubuntu 26.04 LTS, codenamed "Resolute Raccoon," released April 23, 2026, supported until April 2031. Download the Server ISO from ubuntu.com/download/server.
On the Proxmox host, upload it via the web UI: Datacenter → [your node] → local (storage) → ISO Images → Upload. Or, if you're on the Proxmox host shell directly:
cd /var/lib/vz/template/iso
wget https://releases.ubuntu.com/26.04/ubuntu-26.04-live-server-amd64.iso
(Check releases.ubuntu.com/26.04/ yourself for the exact current filename if this one has been superseded by a point release, e.g. 26.04.1.)
1b — Create the VM in Proxmox (no GPU attached yet)
In the Proxmox web UI: Datacenter → [your node] → Create VM
- General: VM ID (any free number), Name (e.g.
ai-server) - OS: select the ISO you uploaded, Guest OS Type = Linux, Version = 6.x – 2.6 Kernel
- System: BIOS = OVMF (UEFI), Machine = q35 — both required later for clean PCIe passthrough. Add an EFI disk when prompted — leave "Pre-Enroll keys" unchecked. This is the important part: if left checked, Proxmox pre-enrolls Microsoft's Secure Boot keys and the VM boots with Secure Boot enabled, which will later reject the NVIDIA kernel module in Step 1g as unsigned (
module verification failed: signature and/or required key missing, orKey was rejected by serviceindmesg). Leaving it unchecked means Secure Boot starts disabled and CUDA's driver just loads — the right tradeoff for a private compute VM where Secure Boot isn't protecting against anything meaningful. SCSI Controller = VirtIO SCSI single. Check Qemu Agent. - Disks: 300GB minimum — Ubuntu itself is small, but FLUX.1-dev, Wan 2.2, and your Ollama models together run well over 100GB, and Docker images add more on top. Size up if you plan to keep multiple model variants side by side. Use your fastest storage (SSD/NVMe-backed) if available, and enable Discard if it is.
- CPU: Type = host (needed for passthrough compatibility and best performance). Allocate most of your cores, leaving 1–2 for Proxmox itself.
- Memory: with 96GB total, allocate around 80GB to the VM — leaves comfortable headroom for the Proxmox host. Turn Ballooning off for a more stable passthrough experience.
- Network: VirtIO, bridged to whichever interface maps to your
192.168.1.0/24network. - Confirm and create — don't start it yet.
1c — Install Ubuntu Server 26.04
Select the VM → Console (noVNC) → Start. Boot from the ISO and run through the installer:
- Language, keyboard layout
- Network: DHCP is fine, or assign a static IP on
192.168.1.0/24now if you'd rather not hunt for the DHCP-assigned address later - Storage: guided, use the entire disk. Uncheck "Set up this disk as an LVM group" if offered — LVM's guided default frequently under-allocates the root logical volume, leaving a large chunk of the disk sitting unused as free space inside the volume group rather than mounted anywhere, and it's easy not to notice until you're wondering where the space went. A plain ext4 partition using the whole disk avoids this entirely for a single-purpose VM like this one. If you do want LVM (e.g. for snapshotting), check the installer's partition summary screen before confirming and make sure the root LV is sized to consume the full volume group, not a fraction of it.
- Profile setup: create your initial admin account here — this is the account you'll use to bootstrap everything, distinct from
skynetaiwhich gets created explicitly in Step 2 - Enable OpenSSH server when offered — you'll want this for everything downstream in this guide
- Skip the optional server snaps (Docker, etc.) — this guide installs those explicitly later, with specific versions and configuration
- Let it finish, reboot, and remove the ISO from the VM's CD drive (Hardware → CD/DVD Drive → do not use any media) so it doesn't try to boot the installer again
1d — Update the system
SSH into the VM (or use the Proxmox console) as the admin account you just created:
sudo apt update && sudo apt full-upgrade -y
sudo reboot
1e — Disable the default video driver (Nouveau)
Ubuntu loads Nouveau, the open-source community NVIDIA driver, automatically whenever it detects NVIDIA hardware — even before the GPU is actually passed through, and it will conflict with the proprietary driver CUDA installs later.
sudo tee /etc/modprobe.d/blacklist-nouveau.conf << EOF
blacklist nouveau
options nouveau modeset=0
EOF
sudo update-initramfs -u
sudo reboot
Confirm it's gone:
lsmod | grep nouveau
No output is correct. Also check for any stray Ubuntu-repo NVIDIA driver package that might have been auto-installed:
apt list --installed | grep nvidia
If anything shows up here before you've installed CUDA yourself, remove it (sudo apt purge <package-name>) to avoid a conflict later.
1f — GPU passthrough (Proxmox host side, done now — last before CUDA)
This is the one part of this step that happens outside the VM, on the Proxmox host itself.
Shut down the VM first (sudo shutdown now inside the guest, or Shutdown from the Proxmox UI) — passthrough changes require it to be off.
On the Proxmox host shell:
- Confirm IOMMU is enabled. Check which bootloader Proxmox is using:
ls /etc/kernel/cmdline 2>/dev/null && echo "systemd-boot" || echo "GRUB"
- If systemd-boot (common on ZFS/UEFI installs), edit
/etc/kernel/cmdlineand addintel_iommu=on iommu=pt(oramd_iommu=on iommu=pton AMD) to the existing line, then:
sudo proxmox-boot-tool refresh
- If GRUB, edit
/etc/default/grub, add the same parameters toGRUB_CMDLINE_LINUX_DEFAULT, then:
sudo update-grub
Reboot the Proxmox host afterward.
- Identify your GPU's PCI address and vendor:device ID:
lspci -nn | grep -i nvidia
Note the address (e.g. 01:00.0) and the ID pair in brackets (e.g. 10de:XXXX) — you'll use both below. Use the actual values shown on your system, not a guessed or copied ID.
- Blacklist the GPU drivers on the Proxmox host (separate from the blacklist you already did inside the guest in Step 1e — this one stops Proxmox itself from grabbing the card):
sudo tee /etc/modprobe.d/blacklist-gpu.conf << EOF
blacklist nouveau
blacklist nvidia
blacklist nvidiafb
blacklist nvidia_drm
EOF
- Bind the GPU to
vfio-pciinstead, using the vendor:device ID from step 2:
sudo tee /etc/modprobe.d/vfio.conf << EOF
options vfio-pci ids=<your-vendor:device-id-from-step-2>
EOF
sudo update-initramfs -u -k all
sudo reboot
- Verify the bind took effect:
lspci -k -s <your-pci-address-from-step-2>
Look for Kernel driver in use: vfio-pci — not nvidia or nouveau.
Attach the GPU to the VM (Proxmox web UI): select the VM → Hardware → Add → PCI Device → choose your NVIDIA GPU → check All Functions and PCI-Express → leave Primary GPU unchecked (this is a headless compute box, not a display passthrough — you'll keep using the Proxmox console for access, and the GPU is purely for compute). Confirm the VM's Machine type is still q35 and BIOS is still OVMF from Step 1b.
Start the VM back up, SSH in, and confirm the GPU is now visible inside the guest:
lspci -nn | grep -i nvidia
1g — Install CUDA (last, now that the GPU is actually present)
Ubuntu 26.04 changed things here: NVIDIA's own CUDA APT repository now officially lists Ubuntu 26.04 as supported, and it installs a matched driver + toolkit pair together — the same "one installer handles both" property you get from the older .run file approach, just via apt instead:
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2604/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt install -y cuda
sudo reboot
Verify:
nvidia-smi
nvcc --version
nvidia-smi should show your RTX 4060 Ti; nvcc --version should report a CUDA 13.x release line.
If apt install cuda doesn't find the package or the keyring filename above has moved on (NVIDIA occasionally revises the exact keyring version number and package names between CUDA releases): go to developer.nvidia.com/cuda-downloads, select Linux → x86_64 → Ubuntu → 26.04 → deb (network), and use the exact commands shown there instead — the underlying method is the same, just with whatever the current filenames are.
If nvidia-smi fails after this, or dmesg | grep -i nvidia shows module verification failed: signature and/or required key missing or Key was rejected by service: this is Secure Boot, not a CUDA problem — the VM's OVMF firmware has it enabled and is refusing to load the unsigned NVIDIA kernel module. This can happen even if you left "Pre-Enroll keys" unchecked in Step 1b, if Proxmox's OVMF defaults changed since this guide was written, or if the VM was created differently than described. Two ways to fix it:
- Disable Secure Boot (simplest, recommended here): from the Proxmox console, reboot the VM and press
Escduring boot to enter the OVMF firmware setup menu → Device Manager → Secure Boot Configuration → uncheck Attempt Secure Boot → save and exit. Reboot, then re-trigger the driver load:
sudo modprobe nvidia
nvidia-smi
If that still doesn't pick it up, reinstalling the package rebuilds and loads it cleanly: sudo apt install --reinstall cuda.
- Keep Secure Boot on (more setup, only worth it if you specifically need it here): enroll a MOK (Machine Owner Key) so DKMS-signed builds are trusted instead of disabling Secure Boot. This involves generating a signing key, importing it with
mokutil --import, confirming the import via a blue on-screen prompt on next boot (works through the Proxmox console), and configuring DKMS to sign with it automatically going forward. This is meaningfully more involved — Proxmox's own wiki page "Secure Boot Setup" walks through it if you need this route.
For a private, single-purpose AI compute VM like this one, disabling Secure Boot is the pragmatic choice — it's not protecting against a threat model that applies here, and it removes an entire class of driver-signing failures for every future kernel/driver update too.
One more thing worth knowing: Ubuntu 26.04 also now ships CUDA directly through its own archive (sudo apt install cuda-toolkit with no NVIDIA repo needed at all) as a simpler, Canonical-maintained alternative. It's typically a version or two behind NVIDIA's own repo, which is why this guide uses NVIDIA's repo above — but it's a legitimate fallback if you'd rather not add an external repository at all.
This is the point where the rest of the guide's original assumption — "Ubuntu Server + CUDA already installed" — is now actually true.
Step 2 — Create the service user (skynetai)
You're already logged in as the admin account you created during Ubuntu installation (Step 1c) — this step creates skynetai from there, without ever logging in as root.
sudo adduser skynetai
# follow the prompts to set a password
# Grant sudo rights (manual authorization — skynetai will be prompted for
# their own password on every privileged command, never silently elevated)
sudo usermod -aG sudo skynetai
Switch into that account for the rest of the guide:
su - skynetai
(or log out and back in as skynetai over SSH)
Confirm:
whoami
# should print: skynetai
Also confirm CUDA's PATH setup carried over to this account — NVIDIA's installer (Step 1g) drops a script into /etc/profile.d/ that applies to every login shell system-wide, so this should just work without any extra config:
nvcc --version
If that resolves normally, you're done — nothing further needed for skynetai specifically. If it says command not found (this would only happen if you reached this shell without a login shell, e.g. skipped the su - above in favor of plain su or a non-login SSH session), add it manually just for this account:
echo 'export PATH=/usr/local/cuda/bin:$PATH' >> ~/.bashrc
echo 'export LD_LIBRARY_PATH=/usr/local/cuda/lib64:$LD_LIBRARY_PATH' >> ~/.bashrc
source ~/.bashrc
nvcc --version
Step 3 — Create the root install folder
sudo mkdir -p /usr/src/ai
sudo chown -R skynetai:skynetai /usr/src/ai
cd /usr/src/ai
Everything below lives under here: /usr/src/ai/ollama, /usr/src/ai/comfyui, /usr/src/ai/searxng, /usr/src/ai/open-webui.
Step 4 — Install Docker
sudo apt update
sudo apt install -y ca-certificates curl gnupg
sudo install -m 0755 -d /etc/apt/keyrings
curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo gpg --dearmor -o /etc/apt/keyrings/docker.gpg
sudo chmod a+r /etc/apt/keyrings/docker.gpg
echo \
"deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.gpg] https://download.docker.com/linux/ubuntu \
$(. /etc/os-release && echo "$VERSION_CODENAME") stable" | \
sudo tee /etc/apt/sources.list.d/docker.list > /dev/null
sudo apt update
sudo apt install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin
# Let skynetai run docker commands WITHOUT sudo and without being root.
# (Note: docker-group membership is effectively root-equivalent on this host,
# same tradeoff as any Docker install — worth knowing, not worth avoiding Docker over.)
sudo usermod -aG docker skynetai
newgrp docker
docker run hello-world
Group membership takes effect on your next login, not just in the current shell — you only need newgrp docker again if you stay in this exact original shell for later steps; a fresh terminal already has it.
Step 5 — Install NVIDIA Container Toolkit
Required so Docker containers (Open WebUI's GPU-accelerated features) can see your GPU.
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \
sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update
sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
If your RTX 4060 Ti shows up, GPU passthrough into Docker is confirmed.
Step 6 — Install Ollama, with models stored under /usr/src/ai, listening on port 21434
curl -fsSL https://ollama.com/install.sh | sh
ollama --version
Redirect model storage, and move Ollama off its default port (11434) onto 21434, listening on all interfaces:
sudo mkdir -p /usr/src/ai/ollama/models
sudo chown -R ollama:ollama /usr/src/ai/ollama
sudo mkdir -p /etc/systemd/system/ollama.service.d
sudo tee /etc/systemd/system/ollama.service.d/override.conf << 'EOF'
[Service]
Environment="OLLAMA_MODELS=/usr/src/ai/ollama/models"
Environment="OLLAMA_HOST=0.0.0.0:21434"
EOF
sudo systemctl daemon-reload
sudo systemctl restart ollama
curl -s http://127.0.0.1:21434
curl should return Ollama is running.
Confirm it's reachable from outside the host too:
curl -s http://<your-vm-ip>:21434
Security note: Ollama has no built-in authentication — see "Network exposure" near the end.
Step 7 — Pull your coding model
The ollama CLI is a separate client from the server — it doesn't know about the port change from Step 6 unless you tell it. Without this, ollama pull/ollama run fail with could not connect to ollama server, since the CLI defaults to the old port (11434) with nothing listening there anymore:
echo 'export OLLAMA_HOST=127.0.0.1:21434' >> ~/.bashrc
source ~/.bashrc
ollama pull qwen2.5-coder:14b
ollama pull qwen2.5-coder:7b # optional lighter/faster fallback
ollama run qwen2.5-coder:14b "Write a Python function that reverses a linked list."
Confirm GPU usage in another terminal:
watch -n 1 nvidia-smi
~9GB VRAM used, ollama process listed.
Verify models landed in the custom path:
sudo du -sh /usr/src/ai/ollama/models
Ollama backend for Open WebUI is now ready — Step 10 will point at your VM's real IP on this port (not 127.0.0.1 — see the note in Step 10 for why).
Step 8 — Install ComfyUI (image + video generation), listening on port 21188
Installed ahead of Open WebUI so its image-generation integration in Step 10 has something to connect to immediately.
Install prerequisites first — git isn't guaranteed present on a minimal Ubuntu Server install, python3 -m venv needs the separate python3-venv package or it fails at creation, and ffmpeg is needed by Wan 2.2's video-output nodes:
sudo apt update
sudo apt install -y git python3-venv python3-pip ffmpeg
cd /usr/src/ai
git clone https://github.com/comfyanonymous/ComfyUI.git comfyui
cd comfyui
python3 -m venv venv
source venv/bin/activate
pip install --upgrade pip
# Install a CUDA-matched torch build. As of Step 1g's CUDA 13.3, PyTorch's
# confirmed current wheel index is cu130 — you don't need an exact minor-version
# match, since PyTorch's pip wheels bundle their own CUDA runtime libraries and
# only require your driver to be at or above the wheel's target (13.3 > 13.0, so
# this is fine). This tag will drift over time as PyTorch adds newer CUDA support,
# so if this specific one 404s later, check https://pytorch.org/get-started/locally/
# in an actual browser (its install command is generated by JS, so it won't show
# correctly if fetched by a tool instead of loaded normally) and use whatever it
# shows for CUDA instead.
nvcc --version # confirm your CUDA version first
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt
python3 -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"
Expected: True NVIDIA GeForce RTX 4060 Ti
Install ComfyUI-Manager:
As of ComfyUI-Manager v4.0, it's a proper pip package (comfyui-manager on PyPI, currently 4.2.x) rather than something you clone into custom_nodes — the older git-clone method installs a now-legacy version that current ComfyUI actively flags as outdated on load, demanding an upgrade:
pip install comfyui-manager
(still inside the venv activated earlier in this step — if you opened a new shell since then, source /usr/src/ai/comfyui/venv/bin/activate first)
Important, and easy to miss silently: ComfyUI-Manager is disabled by default even when installed — it requires an explicit opt-in flag. Installing it via pip above is not enough on its own; every launch command needs --enable-manager explicitly, or Manager will be installed but inactive, with no obvious error — just a missing Manager button in the UI. Both launch commands below already include it.
A second flag matters too: the new Manager UI isn't yet a full feature match for the older one — model downloading specifically (what Steps 8a/8b below need) is one of the features still reported as living only in the legacy interface. Both launch commands below also add --enable-manager-legacy-ui so the Model Manager feature is guaranteed present rather than something you discover is missing mid-step.
Test-launch manually first, on port 21188:
cd /usr/src/ai/comfyui
source venv/bin/activate
python3 main.py --listen 0.0.0.0 --port 21188 --enable-manager --enable-manager-legacy-ui
Open http://<your-vm-ip>:21188, confirm it loads and that Settings → Model Manager is actually present (not just that the page loads) — see the note in Step 8a if it's not where you expect. Ctrl+C to stop.
Run as a systemd service under skynetai (not root):
sudo tee /etc/systemd/system/comfyui.service << 'EOF'
[Unit]
Description=ComfyUI
After=network.target
[Service]
Type=simple
User=skynetai
Group=skynetai
WorkingDirectory=/usr/src/ai/comfyui
ExecStart=/usr/src/ai/comfyui/venv/bin/python3 main.py --listen 0.0.0.0 --port 21188 --enable-manager --enable-manager-legacy-ui
Restart=always
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl daemon-reload
sudo systemctl enable --now comfyui
8a — Install the image model (FLUX.1-dev)
Manager's install flow has proven unreliable across ComfyUI versions during this kind of setup — its batch-install code path can crash with an unhandled server error depending on exact version combinations, silently, with no useful progress indicator when it does. Downloading the file directly is more reliable and just as fast:
mkdir -p /usr/src/ai/comfyui/models/checkpoints
cd /usr/src/ai/comfyui/models/checkpoints
wget -c https://huggingface.co/Comfy-Org/flux1-dev/resolve/main/flux1-dev-fp8.safetensors
-c makes it resumable if the connection drops mid-download (it's ~17GB). This single file already bundles both of FLUX's text encoders, so no separate CLIP/VAE files are needed — it's immediately usable via ComfyUI's Load Checkpoint node once the download finishes (~12GB VRAM at inference).
If wget returns 401/403 instead of downloading: this model is gated. Accept the license once at huggingface.co/Comfy-Org/flux1-dev in a browser, then authenticate the download:
pip install --user huggingface_hub
huggingface-cli login
(paste a token from huggingface.co/settings/tokens, generated after accepting the license)
If you'd rather try Manager first anyway: current versions tuck it inside Settings → Model Manager — not a standalone top-level menu, and not the Models sidebar tab (that's for browsing what's already installed). This has moved before and may move again, so check the sidebar Models tab and Help → Extension Management as fallback locations. If it crashes or hangs with no progress, fall back to the manual download above rather than troubleshooting Manager further.
8b — Install the video model (Wan 2.2)
Manager path: same Settings → Model Manager → search "Wan2.2 14B GGUF" → install the Q4 version (~10–12GB VRAM). If this hits the same kind of failure as 8a did, the direct-download approach applies here too — ask for the specific current Hugging Face file path if Manager's install proves unreliable for this one as well, since GGUF-quantized model repos change more often than FLUX's does. Load the example workflow from the Workflows tab to confirm.
Limitation, stated plainly: Open WebUI's image integration (Step 10b) is stills-only. Wan 2.2 video generation happens directly in ComfyUI's own interface, not inside Open WebUI chat.
8c — Export the API-format workflow (needed for Step 10b)
Build a basic text-to-image FLUX workflow → Settings → Enable Dev Mode → Save (API Format) → note the node IDs (checkpoint loader, positive/negative prompt, image size, sampler) → keep the JSON handy for Step 10b.
ComfyUI backend for Open WebUI is now ready.
Step 9 — Install SearXNG (self-hosted web search), published on port 21081
mkdir -p /usr/src/ai/searxng/searxng-data
cd /usr/src/ai/searxng
cat > docker-compose.yml << 'EOF'
services:
searxng:
image: searxng/searxng:latest
container_name: searxng
ports:
- "0.0.0.0:21081:8080"
volumes:
- ./searxng-data:/etc/searxng
environment:
- SEARXNG_BASE_URL=http://127.0.0.1:21081/
restart: unless-stopped
EOF
docker compose up -d
The container's internal process still listens on port 8080 inside its own namespace (fixed by the upstream image, not configurable) — Docker republishes it on host port 21081, which is what's actually reachable on your network. Security note: SearXNG has no login by default — see "Network exposure."
Enable JSON output:
nano /usr/src/ai/searxng/searxng-data/settings.yml
Under search::
search:
formats:
- html
- json
cd /usr/src/ai/searxng
docker compose restart
Verify:
curl -s "http://127.0.0.1:21081/search?q=test&format=json" | head -c 300
curl -s "http://<your-vm-ip>:21081/search?q=test&format=json" | head -c 300
SearXNG backend for Open WebUI is now ready.
Step 10 — Install Open WebUI (port 21080) and connect everything
By now Ollama, ComfyUI, and SearXNG are all installed and verified.
Important — use your VM's real 192.168.1.0/24 address for every connection below, not 127.0.0.1. This looks backwards at first (Open WebUI runs with --network=host, so loopback should work identically to the real IP), but in practice it doesn't: Open WebUI applies SSRF protection to outbound fetches, requiring an exact origin match between the configured Base URL and whatever address a service actually reports back in its own responses. 127.0.0.1 and the VM's real IP aren't treated as interchangeable for this check even though they resolve to the same host — using the real IP throughout is what actually works, confirmed directly against this exact stack.
mkdir -p /usr/src/ai/open-webui/data
docker run -d \
--network=host \
--gpus all \
-v /usr/src/ai/open-webui/data:/app/backend/data \
-e OLLAMA_BASE_URL=http://<your-vm-ip>:21434 \
-e PORT=21080 \
--name open-webui \
--restart always \
ghcr.io/open-webui/open-webui:cuda
Open http://<your-vm-ip>:21080, create your admin account (first account created becomes admin automatically — this is a web-app account, unrelated to Linux users).
Sanity-check the container actually stayed up on the new port, rather than assuming a clean page load means everything's fine:
docker ps --filter name=open-webui
Should show Up X minutes/seconds, not Restarting. PORT + --network=host is the correct, current way to rebind this image's port — but if you ever do see it stuck restarting despite the page loading once, check docker logs open-webui for the actual reason before assuming the port change is the cause.
10a — Connect web search (SearXNG)
Admin Panel → Settings → Web Search → enable → Search Engine: searxng → SearXNG Query URL: http://<your-vm-ip>:21081/search?q=<query>
10b — Connect image generation (ComfyUI)
Admin Panel → Settings → Images → Image Generation Engine: ComfyUI → ComfyUI Base URL: http://<your-vm-ip>:21188 → paste the workflow JSON from Step 8c → Model field: flux1-dev-fp8.safetensors → leave API Key blank (no auth is configured on ComfyUI)
Mapping the workflow's node IDs — a "ComfyUI Workflow Nodes" section asks for specific node IDs from your exported JSON. Using the standard 7-node FLUX workflow from Step 8 (checkpoint loader → 2× CLIP Text Encode → Empty Latent Image → KSampler → VAE Decode → Save Image), the mapping is:
| Field | Node ID | Why |
|---|---|---|
| prompt | 2 |
The CLIPTextEncode node wired into KSampler's positive input |
| model | 1 |
CheckpointLoaderSimple |
| width | 4 |
EmptyLatentImage |
| height | 4 |
Same node — it has both fields |
| steps | 5 |
KSampler |
| seed | 5 |
Same node — it has both fields |
Enter just the bare number in each field. If your own export has different node IDs (they're assigned in build order and can vary), open the JSON and match by class_type instead of assuming these exact numbers — CheckpointLoaderSimple is always the model node, KSampler is always steps/seed, EmptyLatentImage is always width/height, and whichever CLIPTextEncode node feeds KSampler's positive input (not negative) is the prompt node.
10c — Fix tool-calling before testing either integration
Before testing 10a or 10b, one more setting matters more than either of them — without it, the model will produce fake-looking JSON text instead of actually calling either tool, regardless of how correctly the URLs above are configured.
Settings → Models → click the pencil icon next to your model (e.g. qwen2.5-coder:7b) → Advanced Params → Function Calling → switch from Default to Legacy, and save. Do this for every model you plan to use with web search or image generation.
This matters because current Open WebUI defaults to Native tool-calling, which small local Ollama models are frequently unable to execute correctly — they emit a tool-call-shaped text response without it actually being parsed and run, which looks like a malformed hallucination but is really a mismatch between what the model can reliably produce and what Native mode expects. Legacy mode uses an older, prompt-based mechanism instead, which is far more forgiving for smaller local models. Note the naming trap: in the current UI, Default is not the same as Legacy — despite the name, Default behaves like Native. Legacy is the option you actually want.
Now test both:
- Web search: ask a normal question needing current information, e.g. "what time is it right now in New York?"
- Image generation: use the message action toolbar's image icon (🖼️) on a response, or your version's equivalent trigger, with a prompt like "a red fox in snow"
Step 11 — Install Aider (autonomous terminal coding agent)
sudo apt install -y python3-pip pipx
Ubuntu 26.04 ships Python 3.14 by default, and aider-chat will very likely fail to install on it — one of its older transitive dependencies (numpy==1.24.3, a 2023-era release) has no prebuilt wheel for a Python this new, and its source-build tooling can't compile against 3.14 either, failing with BackendUnavailable: Cannot import 'setuptools.build_meta'. Sidestep this by installing Aider into an isolated environment on an older Python instead, via the Deadsnakes PPA — note it publishes 3.11 for 26.04, not 3.12 (source-build only for that one), so use 3.11:
sudo apt install -y software-properties-common
sudo add-apt-repository ppa:deadsnakes/ppa
sudo apt update
sudo apt install -y python3.11 python3.11-venv
pipx install aider-chat --python python3.11
pipx ensurepath
source ~/.bashrc
Deadsnakes is a well-established, widely-used community PPA — but its own maintainers note it doesn't carry the same security-update guarantees as Ubuntu's own repos. Fine for a single isolated tool like this; worth knowing if you'd lean on it more broadly.
Verify it's actually reachable:
aider --version
Projects Aider works on live under /usr/src/ai/projects, consistent with everything else in this stack — not scattered in your home directory:
mkdir -p /usr/src/ai/projects
If you already have a project, move or clone it in there, e.g.:
cd /usr/src/ai/projects
git clone <your-repo-url> my-app
cd my-app
If you just want to confirm Aider works before pointing it at anything real, create a throwaway repo in the same location:
mkdir -p /usr/src/ai/projects/aider-test
cd /usr/src/ai/projects/aider-test
git init
Make Aider's connection to Ollama permanent, so you don't have to re-export it every new terminal session:
echo 'export OLLAMA_API_BASE=http://127.0.0.1:21434' >> ~/.bashrc
source ~/.bashrc
This stays on 127.0.0.1, not the real IP — unlike Open WebUI (Step 10), Aider is a plain CLI process on the same machine, not a Docker container with SSRF-style origin checks, so loopback is both correct and preferable here (faster, and independent of the VM's IP or network state).
Then, from inside your project's folder:
cd /usr/src/ai/projects/<your-project-folder>
aider --model ollama/qwen2.5-coder:14b
Aider needs the folder to be a git repo — it relies on git to track and commit its changes — so git init is required if you haven't already run one for that project.
If it doesn't connect on the first try, double-check OLLAMA_API_BASE is still the correct env var name for your installed Aider version — check --help or the docs if it's changed.
Switching between sessions (16GB VRAM constraint)
ollama ps
ollama stop <model-name-shown-above>
Before going back to coding, make sure ComfyUI isn't still holding a model (sudo systemctl stop comfyui if needed).
Who runs as what (privilege summary)
| Process | Runs as | Why |
|---|---|---|
| Your shell / all commands from Step 2 onward | skynetai |
normal user, sudo only when a step needs it |
| Steps inside Step 1 (before skynetai exists) | your Ubuntu installer admin account | bootstraps the system before the dedicated service account is created |
| Docker daemon | root (standard for Docker) |
inherent to how Docker works on Linux; mitigated by not logging in as root yourself |
| Open WebUI, SearXNG containers | inside their own containers, launched by skynetai via the docker group |
container process isolation, not host root |
| Ollama service | dedicated ollama system account |
created automatically by Ollama's installer |
| ComfyUI service | skynetai (explicit in the systemd unit) |
pinned deliberately, not left to a template default |
| Aider | skynetai |
plain CLI tool, no elevation required |
Nothing in this stack requires you to log in as root at any point after your initial Ubuntu install (Step 1c).
Network exposure — read this now that everything is on 0.0.0.0
Every service listens on all interfaces, reachable from anywhere on 192.168.1.0/24. Ollama and SearXNG have no authentication at all; Open WebUI requires its own login; ComfyUI has none. If untrusted devices can join that subnet, scope access with a firewall:
sudo apt install -y ufw
sudo ufw allow from 192.168.1.0/24 to any port 21080 proto tcp
sudo ufw allow from 192.168.1.0/24 to any port 21188 proto tcp
sudo ufw allow from 192.168.1.0/24 to any port 21434 proto tcp
sudo ufw allow from 192.168.1.0/24 to any port 21081 proto tcp
sudo ufw allow OpenSSH
sudo ufw enable
Quick reference
| Service | Address | Notes |
|---|---|---|
| Open WebUI | http://<vm-ip>:21080 |
login required |
| ComfyUI | http://<vm-ip>:21188 |
no login |
| Ollama API | http://<vm-ip>:21434 |
no authentication |
| SearXNG | http://<vm-ip>:21081 |
no authentication |
| Path | Contents | Owner |
|---|---|---|
/usr/src/ai/ollama/models |
LLM weights | ollama |
/usr/src/ai/comfyui |
ComfyUI + models | skynetai |
/usr/src/ai/searxng |
SearXNG config | skynetai |
/usr/src/ai/open-webui/data |
Open WebUI data | skynetai (via container) |
/usr/src/ai/projects |
Aider project repos | skynetai |