Track
LLM Internals
How modern LLMs actually work
Chapters
- GPT Pretraining Explained: From Raw Text to Next-Token Prediction
- KV Cache in LLMs Explained: How Transformers Make Inference Fast
- LoRA in LLMs: Beginner's Guide to Fine-Tuning
- Mixture of Experts (MoE) in LLMs: A Beginner's Guide
- Quantization in LLMs Explained: INT8, INT4, GPTQ, AWQ, NF4
- RoPE Explained: Rotary Position Embedding in Transformers
- Speculative Decoding in LLMs: A Beginner's Guide
No pages match that filter.