OpenAI announces GPT-5.3-Codex-Spark will retire next week, ending the first non-NVIDIA inference experiment in under 8 months
Codex lead Tibo announced on September 12 that GPT-5.3-Codex-Spark, the February Cerebras WSE-3 inference experiment, will be retired next week and replaced by GPT-5.6 Sol Ultrafast.
On September 12, 2026, OpenAI Codex lead Thibault Sottiaux announced that GPT-5.3-Codex-Spark will be formally retired next week. The experimental coding model, released in research preview on February 12, had a life cycle of less than eight months and is the first OpenAI model deployed on non-NVIDIA inference silicon to be retired. GPT-5.3-Codex-Spark was positioned as a smaller, faster model for real-time interactive coding, suitable for code editing, local modifications, and targeted testing. Its defining feature was the underlying hardware: it ran on Cerebras WSE-3 wafer-scale processors, the first time OpenAI deployed production inference on anything other than NVIDIA GPUs. Cerebras integrates compute and on-chip SRAM on a single wafer, reducing the latency of weight transfers between compute chips and external memory typical in GPU systems. The experiment came against the backdrop of OpenAI's broader effort to diversify compute suppliers. In January 2026 OpenAI announced a multi-year deal with Cerebras to deploy up to 750 MW of Cerebras inference capacity, valued at more than $20 billion, phased through 2028. Speed, however, did not translate into retention. Spark was a lightweight real-time branch available only to ChatGPT Pro users, in Codex clients, the CLI, and the VS Code extension. Developer feedback was split: some praised its speed on mechanical refactors while others called it "a fun model and an absolutely terrible one that simply does not work". In August, OpenAI shipped GPT-5.6 Sol Ultrafast, also running on Cerebras hardware, peaking at 750 tokens per second. Unlike Spark, Ultrafast runs the full GPT-5.6 Sol rather than a speed-compressed sub-model. When Spark shipped, OpenAI's engineering capacity only allowed trading model quality for response time on real-time coding. With GPT-5.6 Sol Ultrafast, the company is trying to get both at once.