Summary
A Reddit user is starting to learn PySpark for ETL purposes and will use AWS Glue to run and orchestrate his data pipelines. He plans to share daily updates on his progress and questions to maintain discipline.
Deepen your knowledge
Knowledge Base
ETL Explained — Extract, Transform, Load in plain language
What is ETL? Learn how Extract, Transform, and Load works, the difference with ELT, and which tools to use. Clearly expl...
Knowledge BaseData Lakehouse Explained — The best of both worlds
What is a data lakehouse and why does it combine the best of data warehouses and data lakes? Architecture, comparison, a...
Knowledge BaseWhat is Power BI? Everything you need to know
Discover what Microsoft Power BI is, how it works, what it costs, and why it's the world's most popular BI tool. Complet...