Holden Karau, Adi Polak, Rachel Warren | High Performance Spark. Best Practices for Scaling and Optimizing Apache Spark. 2nd Edition (2026) [PDF, EPUB]
Автор: Holden Karau, Adi Polak, Rachel Warren
Издательство: O'Reilly Media
ISBN: 978-1098145859
Жанр: Data Mining, Data Processing, Database Storage & Design
Язык: Английский
Формат: PDF, EPUB
Качество: Изначально электронное (ebook)
Иллюстрации: Цветные и черно-белые
Описание:Apache Spark is amazing when everything clicks. But if you haven't seen the performance improvements you expected or still don't feel confident enough to use Spark in production, this practical book is for you. Authors Holden Karau, Adi Polak, and Rachel Warren walk you through the secrets of the Spark code base and demonstrate performance optimizations that will help your data pipelines run faster, scale to larger datasets, and avoid costly antipatterns.
Ideal for data engineers, software engineers, data scientists, and system administrators, the second edition of High Performance Spark presents new use cases, code examples, and best practices for Spark 4.x and beyond. This book gives you a fresh perspective on this continually evolving framework and shows you how to work around bumps on your Spark and PySpark journey.
With this book, you'll learn how to:
Accelerate your ML workflows with integrations including PyTorch
Handle key skew and take advantage of Spark's new dynamic partitioning
Make your code reliable with scalable testing and validation techniques
Make Spark high performance
Deploy Spark on Kubernetes and similar environments
Take advantage of GPU acceleration with RAPIDS and resource profiles
Get your Spark jobs to run faster
Use Spark to productionize exploratory data science projects
Handle even larger datasets with Spark
Gain faster insights by reducing pipeline running times
Preface
Chapter 1. Introduction to High Performance Spark
Chapter 2. How Spark Works
Chapter 3. Upgrading Spark
Chapter 4. What’s New in Spark 4.2 Since 2.4
Chapter 5. DataFrames, Datasets, and Spark SQL
Chapter 6. Joins (SQL and Core)
Chapter 7. Effective Transformations
Chapter 8. Working with Key/Value Data
Chapter 9. Going Beyond Scala
Chapter 10. Spark: A Thoughtful Step into Generative AI
Chapter 11. Testing, Validation, and Side-By-Side Runs
Chapter 12. Spark Components and Packages
Appendix A. Just Enough Iceberg and Friends
Appendix B. Spark Connect
Appendix C. When Not to Use Spark
Appendix D. Advanced Task Scheduling: Gangs and Resource Profiles
Appendix E. Spark Streaming
Appendix F. The Spark Web UI: Debugging and Optimizing Your Jobs
Index
Скриншоты:
Время раздачи: с 10 до 20 (минимум до появления первых 3-5 скачавших)