Quantization: Making LLMs Smaller Without Making Them Stupid
What LLM quantization is, how GPTQ, AWQ, QLoRA, and GGUF work, and how to choose the right method for local inference, GPU serving, or fine-tuning.
8 posts
What LLM quantization is, how GPTQ, AWQ, QLoRA, and GGUF work, and how to choose the right method for local inference, GPU serving, or fine-tuning.
VS Code's Agents Window is not just a UI: it's a multi-engine harness hosting Copilot CLI, Claude, and cloud agents. How the IDE became an operating system for software engineering agents.
How Copilot CLI hosts OpenSpec and Spec Kit flows on local hardware: propose, apply, archive. Why bounded inference cost makes specification-driven iteration practical.
How custom instructions, AGENTS.md, and skills give your local Copilot the context it needs to stop making stupid mistakes.
How MCP servers turn a local LLM into a real agent: test runners, linters, semantic search, and why tool access compensates for weaker reasoning.
The complete setup guide: vLLM and llama.cpp configs, environment variables, the dual-mode trick, debugging /v1/models, every failure mode, and why owning the runtime changes the equation.
Text becomes tokens, tokens become vectors, vectors consume memory, and memory pressure adds latency. A complete breakdown of the LLM inference pipeline.
Running out of tokens? Will you be part of this new form of slavery or will join the local ai rebellion? Your AI. Your rules.