When to Choose Each Option
Clear guidance based on your specific situation and needs.
Our Recommendation
Neither is the successor to the other. The Realtime API wins whenever a human is waiting mid-sentence: sub-second turn-taking, barge-in, and native SIP telephony that REST cannot reach at all. REST wins everything a production team gets paged about — retryability, stateless recovery, per-request cost attribution, and access to gpt-5.5 and gpt-5.5-pro, which the realtime model line simply does not offer. The cost story is more interesting than the transport story: fresh realtime audio input runs $32.00 per 1M tokens against $5.00 for gpt-5.5 text, but cached realtime audio input is $0.40 per 1M — slightly cheaper than gpt-5.5's own cached text at $0.50. Caching, not transport, is the lever. The pattern most production teams land on is a hybrid: keep the realtime session as a thin conversational interaction layer, and delegate hard reasoning to a REST call against a frontier model behind it. That is what OpenAI itself did with GPT-Live, which hands reasoning to GPT-5.5 in the background. Choose Realtime for the ear. Choose REST for the brain.
- Choose OpenAI Realtime API when...
- A human is waiting mid-sentence and sub-second turn-taking decides whether the product feels alive.
- Your audio starts or ends on a phone line and you want native SIP rather than a self-built telephony bridge.
- Barge-in and interruption handling are product requirements, not nice-to-haves.
- A voice-first interaction layer is the product, and the heavy reasoning can be delegated behind it.
- Choose OpenAI REST API when...
- The task needs gpt-5.5 or gpt-5.5-pro reasoning, which the realtime model line does not offer.
- You need per-request cost attribution and billing that scales with work done, not with connection time.
- Retries, queue replay and stateless recovery are load-bearing parts of your reliability story.
- You need every turn to be a discrete, loggable, replayable artifact for debugging or audit.