SShortSingh.
Back to feed

Automate Data Quality Checks on Pull Requests Using Great Expectations and GitHub Actions

0
·1 views

A developer tutorial published on DEV Community outlines a method to catch bad data at the pull request stage, before any code is merged. The approach combines a YAML-based data contract, a Python validation script using Great Expectations 1.x, and a GitHub Actions workflow that runs on every pull request. Instead of checking production data, the setup tests what the pipeline code does to a known sample input, keeping checks fast and predictable. The YAML contract defines column-level rules such as required fields, allowed values, and value ranges, making rule changes visible as reviewable diffs. Together, these three components cause a pull request to fail automatically in GitHub if the code would produce data that violates the defined contract.

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

How Mojo Uses SIMD to Speed Up Text Search Across Large Documents

SIMD, or Single Instruction Multiple Data, allows a processor to compare multiple bytes simultaneously rather than one at a time, making it far more efficient for large-scale text searches. The Mojo programming language exposes this CPU-level capability through a typed SIMD construct where both the data type and number of lanes are part of the value's type signature. A tutorial on DEV Community walks through building a Mojo application called ResearchLens, which uses SIMD to quickly identify candidate word matches in research documents. The approach pairs a fast SIMD scan for first-byte matches with a scalar verifier to confirm full word matches and handle edge cases like partial vectors. No GPU is required, as the SIMD hardware used is already present in modern CPUs.

0
ProgrammingDEV Community ·

Why HPC Clusters Use InfiniBand: Low Latency and RDMA Explained

High-performance computing (HPC) clusters often deploy InfiniBand alongside standard Ethernet because tightly coupled workloads demand more than just high bandwidth. Applications running across dozens or hundreds of nodes continuously exchange data, making both latency and throughput critical to overall performance. InfiniBand's key advantage lies in Remote Direct Memory Access (RDMA), which enables data transfers directly between nodes' memory with minimal CPU and operating system involvement, reducing communication overhead. However, InfiniBand is not the only option — high-speed Ethernet and RoCE, which brings RDMA capabilities to Ethernet, are also used depending on workload requirements. Ultimately, the network is as vital to HPC performance as the CPUs themselves, since powerful processors are ineffective if nodes spend excessive time waiting for data.

0
ProgrammingDEV Community ·

WebDecoy Plugin Lets Fastify APIs Detect Scrapers and Bots in Monitor Mode

Developers can add request detection to Fastify APIs using the open-source WebDecoy plugin, which identifies whether traffic comes from frontends, integrations, or scrapers hitting the same endpoint. The plugin operates in a non-blocking monitor mode, meaning suspicious requests are flagged and logged but not blocked, leaving API responses unchanged. A tripwire rule can be configured on sensitive paths such as /.env to trigger a DENY verdict whenever those routes are probed. Detection results are logged locally by default, and connecting an API key unlocks a cloud dashboard where flagged requests appear as labeled records. The walkthrough targets Node.js 22.12 or later and uses Fastify 5 alongside the @webdecoy/fastify and @webdecoy/node packages.

0
ProgrammingDEV Community ·

Cloudflare's AI Bot Rules Silently Block Googlebot, Bypassing robots.txt

A September 2026 Cloudflare policy change altered default settings for AI-related crawlers, causing Googlebot, Bingbot, and Applebot to receive 403 errors on ad-supported pages across new and free-tier accounts. Because Cloudflare classifies Googlebot as a mixed-use crawler serving both search and AI training, it falls under the more restrictive AI training block when that setting is enabled. The problem goes undetected because robots.txt still shows an open allow directive, while the actual block is enforced at the CDN edge layer before the origin server is ever reached. Cloudflare's own data shows Googlebot's crawl share dropped to 27.49% in Q2 2026, down sharply from 57.20% a year prior. Site owners can diagnose the issue by comparing HTTP response codes for browser versus Googlebot user-agent requests, or by using Google Search Console's URL Inspection tool for verified confirmation.

Automate Data Quality Checks on Pull Requests Using Great Expectations and GitHub Actions · ShortSingh