Harness AI Assistant Runs Perception Stack On-Device to Keep Costs Near Zero
Harness, a screen-aware AI assistant, processes video frames, text, and audio entirely in the user's browser using on-device models, meaning compute costs are borne by the user's own hardware rather than the developer. The tool relies on several lightweight models — including CLIP for image embedding, PaddleOCR for text extraction, and a quantized 2.6B language model — stored in a roughly 1.7GB download. For complex reasoning tasks that smaller models cannot handle reliably, Harness routes requests through Surplus Intelligence, a marketplace reselling provider quota below standard list prices, reportedly achieving savings of up to 66% on certain models. Revenue is generated from the spread between discounted surplus pricing and the list price charged to users, a margin the developer describes as disclosed bridge revenue rather than the core business. The architecture is designed so that high-frequency, perception-based tasks scale freely on user hardware, while lower-frequency frontier model calls incur real but reduced costs.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in