Categories
Fabric

Materialising DataFrames in Spark

We recently dived into Spark and saw how DataFrames use lazy evaluation to defer execution until needed. This is useful for defining segments of logic to be reused, but if this causes large, complex calculations to execute multiple times, there may be a better way. Lazy evaluation Let’s take an example with some sales data: The display triggers execution […]

Categories
Fabric

Thinking in SQL, Working in Spark

I’ve spent years shaping data with SQL Server, however after pulling at the threads of Fabric I’m opening notebooks and finding PySpark. At first glance the difference is stark, but it’s not quite the dramatic shift it appears. If you’re not familiar, let’s look at what’s very similar, and where the true differences are. Same data, different […]