Latent-GRPO Trains AI to Reason in Continuous Vectors, Not Words

Researchers have proposed Latent-GRPO, a reinforcement learning framework that replaces token-based chain-of-thought reasoning with continuous recurrent thought vectors in embedding space. Current AI reasoning models, including DeepSeek-R1 and OpenAI o1, require neural networks to generate lengthy internal monologues as discrete text tokens, consuming 80–90% of generation capacity on prose no human reads. In multi-step engineering tasks, this verbose reasoning exhausts context windows and causes 19–34% of model rollouts to be cut off mid-generation, receiving zero reward and corrupting training signals. Latent-GRPO addresses this by allowing models to cycle through continuous hidden-state vectors internally without emitting vocabulary tokens, removing the vocabulary projection bottleneck. The approach also introduces two-pass gradient replay to enable exact backpropagation through these continuous thought steps during RL post-training.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in