SShortSingh.
Back to feed

OpsMind AI Agent Uses Two-Stage Memory System to Learn from Past SRE Incidents

0
·3 views

Engineers behind OpsMind, an AI site reliability engineering agent, designed a workflow that separates incident memory into two distinct operations: recalling historical context before diagnosis and retaining new outcomes after resolution. The system uses a tool called Hindsight as a persistent memory layer, querying past incidents based on current symptoms, service behavior, error patterns, and resource utilization rather than matching exact incident IDs. To avoid faulty automation, current telemetry and logs are treated as the primary evidence source, with historical memory serving only as supporting context. This prevents the agent from blindly reapplying old fixes to new incidents that may share surface-level symptoms but have different root causes. The design positions memory between evidence collection and AI reasoning, ensuring historical experience informs but does not override live operational data.

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

Developer Builds Hook to Feed AI Past Architecture Decisions at Session Start

A developer working on a Claude Code plugin in May 2026 noticed that the AI recommended a previously rejected library, TipTap, despite an existing Architecture Decision Record (ADR) ruling it out in favour of BlockNote. The root cause was that AI sessions do not automatically inherit prior design context, leaving documented decisions effectively invisible to the model. To address this, the developer wrote a hook that scans a project's ADR folder at session start and injects a concise index of accepted or amended decisions into the context. The implementation was refined over several months, reducing startup time from a median of 3,842ms to 94.7ms and cutting repeated context injections to a lightweight pointer on session compaction. The author notes the hook only ensures the index is delivered, not that the model will consult or follow the recorded decisions.

0
ProgrammingDEV Community ·

Transcription and image-to-3D tools lead AI website growth in August 2026

An analysis of 2,916 AI websites tracked by Anjin Radar found combined monthly visits rose 4.6% from July to August 2026, reaching an estimated 393 million. Transcription and meeting-note tools drove a 35.3% surge in the text generation category, while 3D modeling sites grew 34.6%, making them the two fastest-expanding segments. Standout performers included iamtypist.dev, which climbed from roughly 5,600 visits in June to 6.61 million in August largely through organic search, and hi3d.ai, which posted three consecutive months of growth. By contrast, image generation — the most crowded AI category — saw visits dip slightly by 1.5%. The data, based on third-party traffic estimates rather than self-reported figures, suggests high-intent free tools and niche creator-focused utilities are outpacing more saturated AI product categories.

0
ProgrammingDEV Community ·

How a Wrong rm -rf Command Wiped GitLab's Production Database in 2017

On January 31, 2017, a GitLab engineer accidentally deleted the PostgreSQL data directory on the primary production server while intending to resync a lagging replica, destroying approximately 300 GB of data within seconds. When the team attempted recovery, all five backup and replication mechanisms in place were found to be non-functional or improperly configured, including nightly pg_dump jobs that had been failing silently due to a PostgreSQL version mismatch. GitLab.com remained offline for roughly 18 hours, with recovery relying on a manual snapshot taken six hours earlier for an unrelated test. The incident resulted in the permanent loss of around 5,000 projects, 5,000 comments, and 700 new user accounts, though Git repositories and wikis were unaffected. GitLab published a detailed postmortem on February 10, 2017, and CEO Sid Sijbrandij issued a public apology, making the incident one of the most widely cited database failure case studies in the industry.

0
ProgrammingDEV Community ·

Seven AI Workflow Failures Health Insurers Must Fix Before Automating Claims

AI adoption in health insurance is outpacing its evidence base, with 84% of large US insurers using AI or machine learning in operations, yet a September 2026 systematic review found only 16 real-world studies on AI in prior authorization and coverage decisions. Common workflow failures include relying on incomplete clinical data, applying outdated coverage rules, and providing inadequate human oversight that lacks meaningful review authority. Regulatory pressure is also mounting, as CMS now mandates faster prior-authorization decisions with specific denial reasons, and API compliance requirements are set to take effect in 2027. Experts recommend that insurers version-control all policy rules with effective dates, build data-quality checks before adjudication, and establish risk-based thresholds that route low-confidence cases to human reviewers. The core challenge for payers has shifted from whether to automate to whether every automated decision can be proven accurate, explainable, compliant, and reversible.