SShortSingh.
Back to feed

Ghost database from cancelled CI job silently faked a passing migration test

0
·5 views

A database migration at an unnamed engineering team appeared to pass all pipeline checks in July but failed immediately in staging because it relied on a Postgres extension that was never explicitly installed. The root cause was traced to long-lived CI runner virtual machines where Docker Compose reused containers from previous jobs, including one from a different branch that had been cancelled mid-run three days earlier. Because the teardown step only executed on success, the cancelled job left its database — complete with the installed extension — running on the shared host. An audit of the four machines uncovered 61 orphaned containers and 140 volumes, some dating back to February. The team resolved the issue by switching to ephemeral runners destroyed after each job, scoping Compose project names to the job ID, running teardown unconditionally, and adding a preflight check to ensure no pre-existing containers are present before a job starts.

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

Wazuh 4.14.7 silently drops events in two of three overload scenarios

Security researchers tested Wazuh 4.14.7 under heavy log ingestion load, simulating one agent reading logs from 1,600 firewalls, and found three distinct ways events can be lost. Only one scenario — when the agent's client buffer is enabled and overwhelmed — triggers an alert, specifically rule 203, while the other two fail silently. Disabling the client buffer caused over 840,000 log lines to vanish with only a single warning written to the agent's local ossec.log file. A third scenario involving an overloaded manager dropped tens of thousands of events, leaving traces only in the manager's ossec.log and an analysisd state file. Researchers warn that if those log files are not actively monitored, busy agents or overloaded managers can lose data with no dashboard alert or rule firing.

0
ProgrammingDEV Community ·

NETO: Open-Source P2P Chat Tool Lets Dev Teams Message Without Cloud Servers

NETO is a peer-to-peer chat application built for development teams who want to communicate without relying on third-party cloud platforms. The tool operates entirely within a local network, meaning no messages leave the office and no user accounts are required. It uses mDNS (Multicast DNS) to automatically discover other instances on the same network, eliminating the need for manual IP configuration or a dedicated discovery server. The setup is designed to be straightforward — users simply install and open the application to begin chatting with teammates. NETO positions itself as a privacy-focused alternative to tools like Slack, where sensitive information such as credentials or architecture discussions may be stored on external servers.

0
ProgrammingDEV Community ·

GitHub Codespaces Bills for 32GB Allocated Disk, Not Actual File Usage

Computer Science teacher Jim McCormick raised concerns in a GitHub Community discussion after his students repeatedly hit free-tier storage limits despite their codebases being under 2GB. The root cause, clarified by community experts, is that GitHub Codespaces provisions a 32GB block storage disk by default, and billing is based on this allocated capacity rather than actual files stored. Storage charges continue to accumulate even when a Codespace is stopped, as the provisioned disk persists until the environment is explicitly deleted. Because the free tier offers only 15 GB-months, a student with just two stopped Codespaces can exhaust the entire monthly quota within roughly a week. The issue highlights a broader cloud resource management challenge, where the gap between perceived and billed usage can lead to unexpected costs for educators, developers, and organizations alike.

0
ProgrammingDEV Community ·

Developer spends $402 on one AI ticket, builds fix to track per-turn costs accurately

A software team lead at Oizom discovered that a single ticket on his AI-powered board had accumulated $402 across nine agent runs, with no cost visibility in the app. The issue stemmed from how Claude Code reports total_cost_usd as a cumulative session total rather than a per-run figure, causing naive readings to double-count earlier runs. To fix this, the open-source orchestrator was updated to subtract the previously saved session total from the new one at the end of each turn, isolating the actual spend per run. Edge cases such as new sessions, resumed sessions, and runs that exit without a result line were all accounted for in the revised logic. The solution was implemented in the project's ticket-tracker orchestrator, which manages headless Claude Code sessions for AI agents acting as board members.