OpenAI Narrows Codex Context Window to 272K Tokens in Surprise Update
Codex Context Window Shrinks by 100,000 Tokens
OpenAI has quietly rolled back the context window for its Codex AI coding assistant, reducing it from 372,000 tokens to 272,000 tokens. The change was introduced through a backported update to version 0.144 of the Codex software, merged on July 18, 2026, via a pull request titled "Backport refreshed bundled model metadata to 0.144."
The pull request, authored by OpenAI engineer sayan-oai, shows a net change of +64 and -54 lines of code. However, the exact rationale for the reduction remains unexplained in the commit message. The update appears to be a hotfix, suggesting the previous 372K context size may have been a bug or an unintended configuration.
Context Window Realities and Trade-offs
Large context windows have been a double-edged sword for AI models. While they allow agents to process entire codebases in a single session, performance degrades sharply as the context fills. A recent hands-on report from XDA detailed how a 284-billion-parameter local model dropped from 3,000 tokens per second to under 10 tokens per second after processing 100,000 tokens of code.
"By the time I'd fed it well over 100,000 tokens of codebase, prefill had dropped to around 200 tokens per second and generation was down under 10," the report notes. "Every new token now has to attend over a much bigger cache, and each checkpoint the engine wrote out to disk had grown past a gigabyte." This scaling issue makes massive context windows impractical for real-time coding loops.
OpenAI's reduction to 272K tokens aligns with a broader industry trend toward optimizing for cost and performance rather than chasing raw context size. The move also comes as the company faces intense competition in the AI pricing war, with Meta and Elon Musk's xAI slashing model costs.
Codex Micro and Hardware Ambitions
The context window reduction coincides with OpenAI's entry into the hardware market. The company recently unveiled the Codex Micro, a $230 specialized keyboard designed for power users. Produced in collaboration with Work Louder, the device features mechanical keys in "clicky" and "silent" variants.
Business Insider reported that the Codex Micro is part of a larger push to integrate Codex deeper into developer workflows. OpenAI has also merged Codex into the ChatGPT desktop app, and Codex along with ChatGPT Work now boast 8 million active users, according to Codex engineering lead Thibault Sottiaux.
The hardware launch signals that OpenAI is betting big on Codex as a platform, even as it makes technical trade-offs like the context window reduction. The keyboard is designed to shorten "side quests" and keep developers locked into the coding flow.
Pricing War Context
The context window change must be viewed against the backdrop of the ongoing AI pricing war. OpenAI CEO Sam Altman recently told CNBC, "Every enterprise now is thinking about spend and the value they're getting in exchange for AI." The company's latest flagship model, GPT-5.6, is designed to use significantly fewer tokens, making it more cost-efficient for customers.
Meta CEO Mark Zuckerberg has been equally aggressive on pricing, stating, "The pricing from some of the other labs is very extreme and has very high margins. We think that there's a real ability to be able to offer frontier or very high-level intelligence at a much more affordable cost."
By reducing the context window, OpenAI may be prioritizing inference speed and cost savings over the ability to process massive codebases in a single pass. This aligns with the company's strategy to make Codex more accessible and affordable for enterprises.
Technical Implications for Developers
For developers using Codex, the reduction from 372K to 272K tokens means they can now fit approximately 200,000 words of code or documentation into a single prompt, down from about 280,000. While still substantial, the change may impact workflows that rely on loading entire large repositories into context.
However, the trade-off may be acceptable if the reduced context leads to faster response times and lower latency. The XDA report highlighted that even with aggressive quantization (down to 2 bits per weight), performance degrades significantly with large contexts. OpenAI's move suggests the company believes 272K tokens is the sweet spot for its current model architecture.
Developers who need larger context windows may need to adopt chunking strategies or use external retrieval-augmented generation (RAG) systems to compensate.
Industry Reactions and Outlook
The quiet nature of the update has raised eyebrows in the developer community. Some speculate that the 372K context size was a bug introduced in a previous release, while others believe it was an experimental feature that proved too costly to maintain.
OpenAI has not issued a public statement explaining the change. The pull request was merged without discussion, and the commit message offers no details beyond "Backport refreshed bundled model metadata."
As the AI coding assistant market heats up, with competitors like GitHub Copilot and Anthropic's Claude gaining ground, OpenAI's decision to prioritize efficiency over raw capability may prove strategic. The company is clearly betting that a faster, cheaper, and more reliable Codex will win over enterprise customers, even if it means a smaller context window.
Related News

Airport Simulator: A Minimalist Take on Air Traffic Control Gaming

AI Mania Is Crippling Enterprise Decision-Making, Experts Warn

NYC Mayor Mandates AI Disclosure in Rental Listings

GPT-5.6 Solves 30-Year Convex Optimization Gap with Single Prompt

Kaiser Nurses: AI Surveillance Is Harming Patient Care and Jobs

