Data pipeline loaded erroneous $4.5 million order without raising errors
A developer conducted an experiment by loading a sabotaged data batch with 12 planted defects into a test data warehouse. The load succeeded without errors, and one erroneous order valued at $4.5 million was included, causing reported revenue to spike from about $395,751 to over $4.9 million. Standard and custom data quality tests in the dbt pipeline correctly identified all 12 defects, including the outlier amount, after the load. However, the pipeline's core loading process, focused only on moving valid data formats, did not inherently block the erroneous row from being ingested. The experiment highlights a distinction between data movement and quality validation, where checks run after loading.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in