I wanted a local assistant that could answer questions about embedded work from my own reference material, without sending anything to a hosted service. That meant three things running at home: a model server, a chat front end, and a retrieval (RAG) pipeline over a pile of documentation.

The stack

Everything runs in a Debian VM on Proxmox, with the GPU passed through via VFIO and the services in Docker:

  • Ollama as the main model server, with llama.cpp as a second OpenAI-compatible backend for models that need tuning Ollama doesn’t expose.
  • Open WebUI as the front end, talking to both.
  • BGE-M3 for embeddings and bge-reranker-v2-m3 for hybrid search reranking.
  • SearXNG for web search when local documents don’t have the answer.

Stage one: the cards I had

GTX 1660 Ti (6 GB)

The first version ran on a GTX 1660 Ti. Six gigabytes of VRAM is the whole story here: small models fit, and that’s about it. 7B-class models at 4-bit quantization run, but there’s little room left for context, and the embedding model and reranker compete with the chat model for the same memory.

It was enough to prove the stack worked end to end. It wasn’t enough to make it pleasant.

GTX 1060 via PCIe passthrough

Next I moved the work into a proper Proxmox VM (Debian 12) and passed through a GTX 1060. This was more of an architecture change than an upgrade: the 1060 is an older card with the same 6 GB ceiling. What it bought me was isolation. The inference VM became its own thing, with its own driver stack, that I could snapshot, rebuild, or move without touching the host.

The passthrough itself needed the usual care: q35 machine type, cpu: host, VirtIO for disk and network, ballooning disabled, and hugepages sized so the VM didn’t OOM the host. GPU passthrough pins all of a VM’s RAM, so the VM’s memory size is a host-level decision, not just a guest one.

Stage two: RTX 5060 Ti 16 GB

The real limit was VRAM, so the upgrade was about memory more than speed. The RTX 5060 Ti 16 GB is a modest card with enough memory to change what the setup could do, and it went into a Debian 13 VM.

Getting it to boot

The install did not go smoothly. With the new card in the chipset-fed second x16 slot, the host failed to boot. Resizable BAR asked for a large MMIO window, and the firmware ran out of address space before it got to the SATA controllers, which came up with unassigned BARs and no disks.

The fix had two parts: move the GPU to the CPU-fed primary x16 slot, and update the BIOS to a newer AGESA. The machine is headless, so I flashed it remotely through a PiKVM’s virtual mass storage.

The BIOS update brought its own small surprises: the Ethernet interface was renamed (fixed in /etc/network/interfaces), and SVM (AMD virtualization) was reset to disabled, which is not a great default on a hypervisor. Inside the guest, Blackwell needed a current driver with NVIDIA’s open kernel modules.

One gotcha for anyone running the NVIDIA container toolkit: the CDI spec has to be regenerated after any GPU or driver change, or containers silently lose the GPU.

nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml

What 16 GB changed

  • Bigger models, fully on the GPU. A 14B coding model now runs with a 32K-token context in about 15 GB of VRAM. On 6 GB that model didn’t fit at all.
  • Mixture-of-experts models became practical. Gemma 4 26B-A4B runs at around 35 tokens/s with every language-model layer resident on the GPU, at a stable 16K context.
  • Larger MoE models with CPU offload. A 35B-A3B Qwen model runs through llama.cpp with some experts offloaded to the CPU, fitting a 32K context in just under 15 GB of VRAM.

The MoE offload had one surprise worth sharing. Generation started at about 16 tokens/s. Turning off mmap (no-mmap = true) took it to about 28.5 tokens/s. With mmap on, the expert weights left in system RAM were being faulted in from disk as they were needed, so the “CPU offload” was really disk offload.

Things I learned along the way

  • Bigger isn’t always better for RAG. For retrieval-backed questions, qwen2.5-coder:7b gave better answers than qwen3:4b. Qwen3’s thinking mode spent a lot of tokens reasoning about simple lookups.
  • Chunking matters. Chunks of roughly 2,000 to 2,500 characters with 300 characters of overlap worked better for technical documentation than smaller defaults.
  • Find the real context limit empirically. Model cards and calculators are a starting point. The 14B model’s practical ceiling on this card was 32K; it spilled out of VRAM around 43K.
  • When a model misbehaves, check the plumbing first. A vision model that seemed to hallucinate about images turned out to be hitting a known image ingestion bug in Ollama, and a model running mostly on the CPU turned out to be its vision projector being kept off the GPU.

Where it’s going

The VM has since moved to a newer host with DDR5 and more RAM, which helps the CPU-offloaded experts. The next project is a layered reference collection for RAG: a shared base of C standards, POSIX, Linux kernel docs, RTOS docs and compiler manuals, with smaller per-project collections on top. One early conclusion: portable, pre-built datasets for this are thin, partly because of licensing and partly because embeddings are tied to the model that made them.