Quantization: Making LLMs Smaller Without Making Them Stupid
What LLM quantization is, how GPTQ, AWQ, QLoRA, and GGUF work, and how to choose the right method for local inference, GPU serving, or fine-tuning.
1 post
What LLM quantization is, how GPTQ, AWQ, QLoRA, and GGUF work, and how to choose the right method for local inference, GPU serving, or fine-tuning.