Napačna izbira? Nič za to! Izdelke lahko vrnete do 30 dni
Z darilnim bonom ne morete zgrešiti. Obdarovanec lahko v zameno za darilni bon izbere karkoli iz naše ponudbe.
Do 30 dni za vračilo
Learn how to build practical big data pipelines with Hadoop, Apache Spark, and PySpark.
As datasets grow beyond the practical limits of a single computer, organizations need reliable ways to store, process, clean, transform, and analyze information at scale. Engineering Big Data Pipelines with Hadoop and Spark provides a practical introduction to the tools and workflows used to solve these problems.
This beginner friendly guide takes you from the foundations of big data and distributed computing to complete processing pipelines using Hadoop, Spark, and Python. You will learn not only what the technologies do, but how their individual components fit together in a repeatable data engineering workflow.
Inside this book, you will learn how to:
The book emphasizes a practical workflow rather than memorizing commands. You will learn to move data through a complete pipeline: store it, inspect it, clean it, transform it, analyze it, optimize the processing job, and save the results.
Practical examples use familiar business data including sales transactions, customer orders, product records, website logs, incoming events, and customer behavior data. The final project combines the major skills developed throughout the book by moving raw transaction data into HDFS, processing it with PySpark, analyzing it with Spark SQL, optimizing the job, and saving the results as Parquet.
You do not need previous experience with Hadoop, Spark, distributed systems, or data engineering. Basic computer skills are enough to begin, while some familiarity with Python and SQL can be helpful as you progress into PySpark and Spark SQL.
Whether you are a student learning data processing, a Python user working with larger datasets, an analyst exploring distributed computing, a software developer moving into data engineering, or a technical professional preparing to work with Hadoop and Spark, this book provides a structured path from fundamentals to practical implementation.
Build the skills to design, process, optimize, and understand large scale data pipelines with Hadoop, Spark, and PySpark.