SShortSingh.
Back to feed

Jeff: Lightweight 0.8B Decision Models Trained Locally with 30ms Inference

0
·2 views

A developer has released Jeff, an open-source project featuring small 0.8-billion-parameter decision-focused language models. The models are designed to be compatible with the Jev specification and can be trained on consumer hardware at home. Inference runs at approximately 30 milliseconds, making the models suitable for fast, lightweight decision-making tasks. The project is publicly available on GitHub under the repository firelex/jeff. It has garnered early attention on Hacker News, though community discussion is still limited.

Read the full story at Hacker News

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

Why Incoming Service Links Can't Be Trusted From Your Own Manifest

Software engineer Anton, working on decomposing a PHP monolith into Go microservices, highlights a critical architectural blind spot: a service's outgoing dependencies can be self-declared, but its incoming callers cannot. A single configuration change can introduce a live production edge without touching code, tests, or diagrams, leaving architecture maps silently wrong. The danger is not the drift itself but the uncertainty — a diagram that might be current is more dangerous than one known to be outdated. Anton illustrates this with a real audit finding: a cache-invalidation consumer was written and subscribed but never registered as a daemon, meaning invalidation silently never ran in production. Because the code compiled and unit tests passed, no alert was raised, demonstrating that declared dependency maps can show edges that do not actually exist at runtime.

0
ProgrammingDEV Community ·

Developer Builds AI Invoice Agent That Learns from Past Human Approval Decisions

A developer created VendorSense, an accounts payable automation agent that uses a memory layer called Hindsight to retain context from previous human invoice decisions. Unlike stateless AI models that treat each invoice independently, VendorSense retrieves a vendor's historical approval patterns, payment terms, and purchase-order formats before the reasoning model evaluates a new invoice. The system is built with Python, Streamlit, pypdf, and Groq, and makes its memory and pipeline activity visible through a dedicated dashboard. A key design principle is that historical memory serves as evidence to inform decisions, not as authoritative truth, keeping humans in the loop to generate the learning signal. The architecture prioritizes loading vendor history before analysis, rather than consulting it as an optional afterthought.

0
ProgrammingDEV Community ·

SupportMind AI Hackathon Project Gives Customer Support Agents Long-Term Memory

A team built SupportMind, an AI-powered customer support assistant, as a hackathon project to address a common frustration: chatbots forgetting past interactions with returning customers. The system assigns each customer a separate long-term memory bank, which is searched before any response is generated, giving the AI relevant context from prior conversations. After each interaction, the new conversation is stored back into that customer's memory bank, making the assistant progressively more informed over time. Built using Flask, Hindsight, Groq, and a large language model, SupportMind keeps customer histories strictly isolated so one user's data never influences another's responses. The project also aims to help human support agents by generating concise briefings drawn from a customer's stored interaction history.

0
ProgrammingDEV Community ·

Google Gemini 3.8 Live Avatar Now Generally Available for Enterprise AI Agents

Google has made Gemini 3.8 Live with Live Avatar generally available within Gemini Enterprise, bringing a near real-time on-screen avatar to production conversational agents. The avatar synchronizes visually with Gemini's spoken responses and works alongside live voice and video understanding for more natural customer-facing interactions. Organizations can customize avatars using a reference photo and audio samples, then deploy them via Gemini Studio and the Live API. The feature is available on Gemini Enterprise endpoints in the US and EU, with enterprise-grade throughput, compliance, and data governance provisions. Intended use cases include live customer support, branded marketing experiences, and sales assistants, while the related Extended Thinking capability remains in preview.