SShortSingh.
Back to feed

Postgres Creator: LLMs Achieve 0% Accuracy on Real Enterprise Database Queries

0
·44 views

Turing Award winner and Postgres creator Mike Stonebraker has claimed that large language models score 0% accuracy when tested on real-world production data warehouse queries, far below the 80–85% figures reported on popular benchmarks like Spider and Bird. Speaking on the Data Renegades podcast, Stonebraker explained that standard benchmarks use clean, simple datasets that bear little resemblance to the complex, idiosyncratic schemas found in actual enterprise systems. He identified four core reasons for the gap: enterprise data is absent from LLM training sets, real queries are far more complex, production schemas are messy and inconsistently named, and companies use domain-specific terminology that models have never encountered. To support his findings, Stonebraker released BEAVER, an anonymized benchmark based on four real data warehouses, urging AI researchers to test their models against realistic data. He drew a parallel to past tech hype cycles, warning that AI vendors may be shipping demos of capabilities that do not yet function reliably in production environments.

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

Tessvia AI testing tool developed from spec to beta in 72 hours

A developer documented the creation of Tessvia, a new software testing tool. The process began on June 10, 2026, by handing a complete MVP specification to an AI coding environment named Claude Code. Within the first day, a foundational codebase with 11 packages, a Chrome extension, and a functional CLI was established. Over the next two days, the team set up end-to-end testing and selected the product's name, completing a working beta version within a 72-hour timeframe.

0
ProgrammingDEV Community ·

Security analysis finds Polygon Bridge carries moderate-to-high risk due to design

A senior DeFi security researcher analyzed the Polygon Bridge on October 2, 2026. The bridge, which holds over $2.8 billion in assets, facilitates transfers between Ethereum and Polygon. Its optimistic security model uses a seven-day challenge period to reduce costs but creates a significant attack surface. The analysis rated the overall risk as moderate-to-high due to past exploits and the system's complexity. Key vulnerabilities include potential validator collusion and exploitation of the challenge window.

0
ProgrammingDEV Community ·

Critical vm2 sandbox flaw allows host system command execution

A critical vulnerability designated CVE-2026-92957 was published on October 1, 2026, affecting the vm2 sandbox package. The flaw exists in versions 3.11.6 and earlier of vm2's NodeVM component. It allows sandboxed code to bypass security policies and load restricted host modules like child_process. This bypass occurs due to improper handling of the 'node:' prefix in policy configuration, enabling remote code execution on the host system. The issue is fixed in version 3.11.7 of the vm2 package.

0
ProgrammingDEV Community ·

Developer details pitfalls of using hard-coded values instead of configurable settings

A project has adopted a rule that any value someone might reasonably want to change without altering code should be a configurable setting, not a literal value. One failure mode is creating a setting that no code reads, making the feature silently non-functional. The more common error is keeping the old hard-coded value alongside a new setting, causing changes to have no effect. The article notes that properly removing the original literal after registering a setting is critical for the change to be real.

Postgres Creator: LLMs Achieve 0% Accuracy on Real Enterprise Database Queries · ShortSingh