Part 4 - Bad records path

In part 4, the final part of this beginner’s mini-series of how to handle bad data, we will look at how we can retain flexibility to capture bad data and proceed uninterrupted.

Jun 2021

Part 3 - Permissive

In the 3rd instalment of this 4-part mini-series, we will look at how we can handle bad data using PERMISSIVE mode. It is the default mode when reading data using the DataFrameReader but there’s a bit more to it than simply replacing bad data with NULLs.

May 2021

Part 2 - Dropmalformed

In the second part, we’ll continue to focus on the DataFrameReader class and look at the option, **DROPMALFORMED** to **remove** bad data.

May 2021

Part 1 - Failfast

Receiving bad data is often a case of “when” rather than “if”, so the ability to handle bad data is critical in maintaining the robustness of data pipelines.

May 2021