Cactus fine-tunes Gemma 4 with a confidence probe to cut cloud AI costs
Startup Cactus has post-trained Google's Gemma 4 E2B model with a lightweight 68,000-parameter probe layer that predicts how likely the model is to be wrong on any given response. Each output is assigned a confidence score between 0 and 1, allowing developers to process high-confidence queries entirely on-device and route only uncertain ones to a larger cloud model. By sending just 15–55% of queries to Gemini Flash-Lite depending on the task, the hybrid system matches the cloud model's benchmark performance across text, vision, and audio tasks. The probe achieves an average AUROC of 0.814 across 12 hold-out benchmarks, significantly outperforming token entropy heuristics which scored 0.549. Cactus has released the model weights on HuggingFace and the code under an MIT license, with support for Transformers, MLX, and Llama.cpp available now.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in