Developer finds local LLMs handle only half of Claude Code's individual steps

A developer tested whether a local LLM running on an RTX 4070 with Qwen 3.5 4B could replace Claude Code, finding that 96 out of 100 full requests required the frontier model due to complex planning demands. However, when breaking sessions into 200 individual steps, nearly half proved manageable locally, particularly simple extraction tasks like WebFetch. This led the developer to build a step-level router that directs tasks to local or cloud models based on rule matching rather than routing entire requests. A simple rule — sending WebFetch calls to the local model while pinning planning-heavy steps like sub-agents to the frontier — outperformed a dedicated 9B probability-judge model. The resulting open config file defaults to frontier for anything rules cannot classify, with an optional probability-based fallback for edge cases.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in