SShortSingh.
Back to feed

Kong AI Gateway Used to Add Per-Tool Access Controls to Meta's Muse Code Agent

0
·2 views

A developer tested whether Meta's terminal coding agent, Muse Code, could autonomously handle the first stage of an incident investigation by connecting it to operational tools via remote MCP servers. A key concern was that Muse Code's MCP tools run outside the client's sandbox and approval mechanisms, meaning the agent could execute high-risk actions like rolling back production deployments without any client-side prompt. To address this, the developer routed the agent through Kong AI Gateway 2.0, which converts a REST API into MCP tools and enforces per-tool access control lists based on caller identity. Two identities were configured — an investigator and an operator — where the investigator could open incidents but was entirely hidden from the rollback tool, not merely blocked from using it. The full setup, including gateway configuration, a mock ops API, and real agent transcripts, has been published on GitHub.

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

TeamTavern uses shared PureScript types to unify HTTP client and server routing

TeamTavern, an online teammate-finding platform, is built entirely in PureScript with both its browser app and API server sharing a single codebase compiled by one build tool. Every HTTP endpoint is defined as a single type-level route, from which the server handler and client call are both automatically derived. The routing library used is Jarilo, which follows a similar philosophy to Haskell's Servant, encoding path, request body, and all possible responses directly into the type. The compiler enforces correctness end-to-end: handler argument types, return types, and response variants are all inferred from the shared route type, making mismatches a compile-time error. This architecture eliminates an entire class of client-server contract bugs by ensuring neither side can reference a response or parameter that the route does not declare.

0
ProgrammingDEV Community ·

Study of 24 AI Agent Payment Lanes Finds Zero Verified Payouts

A research snapshot by T3rnel Market Pulse evaluated 24 lanes where AI agents are claimed to earn money, finding not a single verified payout as of 29 September 2026. Across 115 job-board applications submitted between July and September 2026, agents received no responses whatsoever — no rejections, no offers, nothing. On the Algora platform, 14 bounty pull requests were successfully merged, generating $740 owed to agents, yet none of that money was received. Combined with other pending claims, the ledger shows $1,115 owed or pending against $0 paid out. The study notes that 17 of the 24 lanes are live and accessible, meaning absent payout evidence is not proof those lanes are fraudulent — but it does show that claims of agent earnings remain unverified measurements rather than established facts.

0
ProgrammingHacker News ·

Magnitude launches self-optimizing local AI inference engine, claims 2x llama.cpp speed

YC S25 startup Magnitude, founded by engineers Anders and Tom, has launched an open-source inference engine designed specifically for running AI agents on local hardware. Unlike existing solutions such as llama.cpp or vLLM, Magnitude uses on-device kernel compilation and tuning to maximize performance on the user's specific hardware across Mac, Linux, and Windows. The engine employs dynamic memory allocation and hybrid paged attention to support multiple concurrent agent sessions without monopolizing system resources. Benchmarks against llama.cpp using Qwen 3.6 35B show up to 92% faster decode speeds on Apple M4 Pro and 23% faster prefill on CUDA hardware, alongside roughly 27-28% lower per-agent memory usage. Built in Rust and licensed under Apache 2.0, Magnitude ships as a desktop app and integrates with popular agent tools, with future plans including loading oversized models from RAM or disk just-in-time.

0
ProgrammingDEV Community ·

Developer builds offline natural-language search for GitHub starred repos using on-device AI

A developer has built a fully offline semantic search system that lets users query their GitHub starred repositories using plain English sentences. The system uses Google's EmbeddingGemma 300M model — a 308-million-parameter embedding model built on Gemma 3 — to convert search queries into 768-dimensional vectors entirely on the user's device. Stored repository vectors are ranked against the query vector using PGlite and pgvector, returning the top 50 matches without any data leaving the machine. The search interface is built with TanStack Start, featuring a 600ms debounce and URL-based query state, while the Elysia backend handles input validation and routes requests to the local Deno process. The project builds on earlier components covering GitHub authentication, a background embedding worker, and a real-time UI update bus.