Gemma 4's quantized models finally made local AI practical in my homelab
I run Home Assistant in Proxmox on my mini PC, along with multiple other services. I use local LLMs running in Ollama, but they're very slow on my local hardware. I've always wanted to use a local LLM to make my local voice assistant smarter, but until now, it's always been far too slow or inaccurate to be of any use.
