Local AI inference on hardware I already had, and then a bit more
Running local LLMs and a RAG stack in a Proxmox VM: starting on a 6 GB GTX 1660 Ti, moving to a passed-through GTX 1060, and finally upgrading to an RTX 5060 Ti 16 GB, plus what changed at each step.