All Courses

Practical Quantization for Large Language Models

Chapter 1: Foundations of Model Quantization

Introduction to Model Compression

Why Quantize Large Language Models?

Representing Numbers: Floating-Point vs. Fixed-Point

Integer Data Types in Quantization

Quantization Schemes: Symmetric vs. Asymmetric

Quantization Granularity Options

Measuring Quantization Error

Overview of Quantization Techniques

Quiz for Chapter 1

Chapter 2: Post-Training Quantization (PTQ)

Principles of Post-Training Quantization

Calibration: Selecting Representative Data

Static vs. Dynamic Quantization

Common PTQ Algorithms

Handling Outliers in PTQ

Applying PTQ to LLM Layers

Limitations of Basic PTQ

Hands-on Practical: Applying Static PTQ

Quiz for Chapter 2

Chapter 3: Advanced PTQ Techniques

Introduction to GPTQ

Understanding GPTQ Algorithm Mechanics

AWQ: Activation-aware Weight Quantization

SmoothQuant: Mitigating Activation Outliers

Comparing Advanced PTQ Methods

Implementation Considerations for Advanced PTQ

Hands-on Practical: Quantizing with GPTQ

Quiz for Chapter 3

Chapter 4: Quantization-Aware Training (QAT)

Need for Quantization-Aware Training

Simulating Quantization Effects During Training

Straight-Through Estimator (STE)

Implementing QAT with Deep Learning Frameworks

Fine-tuning Models with Quantization Nodes

Benefits and Drawbacks of QAT vs. PTQ

Practical Considerations for QAT Execution

Hands-on Practical: Setting up a Simple QAT Run

Quiz for Chapter 4

Chapter 5: Quantization Formats and Tooling

Overview of Common Quantized Model Formats

GGUF: Structure and Usage

GPTQ Format: Library Support and Implementation

AWQ Format Details

Working with Hugging Face Transformers and Optimum

Using bitsandbytes for Quantization

Tools for Model Conversion and Loading

Practice: Converting and Loading Quantized Formats

Quiz for Chapter 5

Chapter 6: Evaluating and Deploying Quantized LLMs

Metrics for Evaluating Quantized Models

Benchmarking Inference Speed and Memory Usage

Hardware Considerations for Quantized Inference

Deployment Strategies for Quantized LLMs

Troubleshooting Common Quantization Issues

Analyzing Accuracy vs. Performance Trade-offs

Practice: Benchmarking a Quantized LLM

Quiz for Chapter 6

Tools for Model Conversion and Loading

Was this section helpful?

References

llama.cpp: Inference of LLaMA model in pure C/C++, Georgi Gerganov and the llama.cpp community, 2023 - The open-source project that introduced the GGUF format and provides tools for model conversion and inference.
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh, 2022 ICLR 2023 DOI: 10.48550/arXiv.2210.17323 - Introduces the GPTQ algorithm, a method for post-training quantization of large language models without performance degradation.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration, Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, Song Han, 2023 MLSys 2024 DOI: 10.48550/arXiv.2306.00978 - Presents the AWQ method for weight quantization, focusing on retaining model accuracy through activation-aware adjustments.
Hugging Face Transformers Documentation, Hugging Face team, 2023 - Official documentation for the transformers library, a standard for loading and using various models, including quantized ones.
TimDettmers/bitsandbytes: 8-bit and 4-bit quantization for PyTorch, Tim Dettmers and bitsandbytes contributors, 2023 - The GitHub repository for bitsandbytes, a library enabling efficient 8-bit and 4-bit quantization for PyTorch models, integrated with transformers.

© 2025 ApX Machine LearningEngineered with