OpenAI is previewing a new Ultrafast inference mode designed for latency-sensitive AI applications.
The company says GPT-5.4 Mini can reach up to 1,000 tokens per second in the new mode, pointing toward faster coding, voice and interactive AI products where waiting time directly shapes the user experience.