Chi Wang, Peiheng Hu | Hands-On LLM Serving and Optimization. Hosting LLMs at Scale (2026) [PDF, EPUB]
Автор: Chi Wang, Peiheng Hu
Издательство: O’Reilly
ISBN: 979-8-341-62149-7, 979-8341621466
Жанр: Software Design Tools, Natural Language Processing, Computer Science
Язык: Английский
Формат: PDF, EPUB
Качество: Изначально электронное (ebook)
Иллюстрации: Цветные и черно-белые
Описание:Large language models (LLMs) are the reasoning engines of modern AI. Today, a major inflection point has arrived: as the world races to deploy AI at scale, model inference has moved to the center of the stack. Welcome to the inference era.
Without proper optimization, however, LLMs can be expensive and slow to serve. Hands-On LLM Serving and Optimization is a comprehensive guide to the complexities of deploying and optimizing LLMs at scale.
In this hands-on, engineering-focused book, authors Chi Wang and Peiheng Hu combine practical examples, code, and strategies for building robust, performant, and cost-efficient AI token factories. Whether you're building the LLM inference infrastructure or the applications that consume it, a deep understanding of LLM serving will make you a more effective, future-ready engineer as AI transforms how we work and build.
Learn the foundations of model serving with core concepts, design paradigms, and industry best practices
Understand the common challenges of hosting LLMs at scale
Balance latency and throughput to meet the demands of AI applications and business requirements
Host LLMs cost-effectively with practical, code-backed techniques
Preface ix
Chapter 1. Introduction to Model Serving and Optimization 1
Chapter 2. Large Language Model Serving 31
Chapter 3. Model Serving System Design: A Deep Dive 65
Chapter 4. Model Serving Best Practices 107
Chapter 5. Challenges When Serving LLMs 157
Chapter 6. Essential LLM Optimization Techniques 189
Chapter 7. Advanced LLM Optimization Techniques 233
Chapter 8. LLM Serving Frameworks 271
Chapter 9. LLM Optimization in Practice 293
Chapter 10. Advancements in LLM Serving 315
Index 333
Скриншоты:
Время раздачи: с 10 до 20 (минимум до появления первых 3-5 скачавших)