OpenAI previews Ultrafast mode for real-time AI workloads
OpenAI is previewing a new Ultrafast inference mode designed for latency-sensitive AI applications.
The company says GPT-5.4 Mini can reach up to 1,000 tokens per second in the new mode, pointing toward faster coding, voice and interactive AI products where waiting time directly shapes the user experience.
Source
Published: