OpenAI adds long-context support to GPT-5.6 Fast mode, doubling costs above 272K tokens
OpenAI updated its API on August 5 to allow GPT-5.6 models — Sol, Terra, and Luna — to process prompts exceeding 272,000 tokens in Fast mode, which was previously called Priority processing. While Fast mode can deliver responses up to 2.5 times quicker, it carries an additional per-token premium on top of already elevated long-context rates. Prompts surpassing the 272K token threshold are priced at twice the standard input rate and 1.5 times the standard output rate, and enabling Fast mode compounds that cost further. For example, a 300K-token request using GPT-5.6 Terra costs $1.38 in Standard mode but doubles to $2.76 in Fast mode. Developers are advised to benchmark both modes on fixed test cases, use feature flags for gradual rollout, and only enable Fast mode when the latency gains justify the added expense.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in