SShortSingh.
Back to feed

Retrying entire AI workflows on Gemini 500 errors causes duplicate side effects

0
·2 views

Developers using Google's Gemini or Vertex AI APIs often respond to intermittent 500, 502, 503, or 429 errors by adding broad retry logic around their entire automation workflow. However, retrying the full workflow replays every side effect, leading to duplicate CRM updates, Slack alerts, Jira tickets, and inconsistent database states. The root cause is misplaced retry boundaries: the retry logic should wrap only the isolated model call, not the steps before or after it. Google's Gemini Python SDK already handles transient failures with exponential backoff by default, meaning raw HTTP clients or full-workflow retries are the more likely culprits. The recommended fix is to isolate the LLM call into a dedicated sub-workflow or worker, so that a transient model failure triggers only that step to retry while downstream writes execute exactly once.

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

How to Measure AI Visibility Reproducibly: Fixed Question Bank and Explicit Denominators

Brasil GEO, a Brazilian AI visibility monitoring firm, has published a reproducible protocol for measuring brand mentions across AI search engines including ChatGPT, Claude, Gemini, and Perplexity. The methodology treats AI visibility as a distribution variable, reporting mention rates as the share of executions in which a brand appears, always paired with sample size, time window, variance, and collection coverage. The protocol requires a frozen bank of 30 to 40 real-customer questions, a minimum of 5 executions per question for continuous monitoring and 30 for before-and-after comparisons, with all parameters fixed before the first data collection round. In a sample run from August 29, 2026, covering 37 questions across four engines, the combined brand mention rate was 21.1 percent, but per-engine rates ranged from 16.2 percent on ChatGPT and Claude to 35.1 percent on Perplexity, underscoring why aggregated averages can obscure meaningful platform-level differences. The approach draws on findings from a January 2026 SparkToro and Gumshoe.ai study of nearly 3,000 prompts, which found that even category-leading brands appeared in only 55 to 77 percent of responses.

0
ProgrammingDEV Community ·

GEONMI-MEMS Architecture Claims 40% Weight Cut for VLEO Satellites via Plasma Harvesting

A developer has proposed the GEONMI-MEMS VLEO Architecture, a software-defined system aimed at extending the operational life of satellites flying at approximately 250 km altitude, where atmospheric drag rapidly depletes conventional fuel. Rather than countering drag, the design harvests ambient ionospheric plasma at orbital velocities of around 7.5 km/s to assist micro-propulsion, while safely redirecting electrostatic discharge away from onboard batteries. The system comprises two subsystems — GEONMI-MEMS VLEO Engine 2 for autonomous propulsion modeling and AeroCore-3 for real-time energy harvesting — both governed by deterministic C++17 software running at a fixed 100Hz control loop. The developer claims the architecture reduces satellite wet mass by 40%, potentially stretching VLEO satellite lifespans from months to years and allowing more revenue-generating payloads per launch. The project is described as proprietary intellectual property intended for civil and commercial space applications, with source code published on GitHub.

0
ProgrammingDEV Community ·

Why a Three-Model AI Jury Can Create False Confidence Instead of Safety

Using three AI models to reach a consensus decision — a design pattern called a three-model jury — is gaining traction in agentic AI systems as a form of cross-validation. However, models trained on overlapping data can share the same biases, meaning unanimous agreement may reflect a shared blind spot rather than correctness. Running three models in parallel also significantly increases token usage and latency, while non-deterministic outputs mean quorum outcomes can shift unpredictably with model temperature settings. Engineers are advised to log not just the majority decision but also dissenting rationales and abstentions, as minority reports often expose ambiguities the majority ignored. Without instrumentation tracking dissent, abstentions, and correlated failures, the author argues the setup functions as three copies of the same guess rather than a genuine safety mechanism.

0
ProgrammingDEV Community ·

Developer rebuilds 28-year-old Turkish IRC network in Go and WebAssembly

A solo developer has completely rewritten Yudum.NET, a Turkish IRC network originally launched in 1998, replacing its legacy C daemon and scripted bots with a modern platform built entirely in Go. The new system runs as a single binary on one VPS, combining an IRC daemon, user services, a website, and a webchat into one unified process backed by a SQLite database. Notably, the webchat and interactive site components are compiled to WebAssembly using TinyGo, eliminating the need for hand-written JavaScript and reducing a class of cross-site scripting vulnerabilities by design. The platform retains classic IRC connectivity for existing clients while adding social feeds, table games, voice rooms, and media discovery features accessible directly within the chat window. The rebuild was motivated by the failure of traditional IRC interfaces to retain casual users who gravitated toward modern apps like Discord and WhatsApp.