OpenAI

OpenAI Product Day Sept 10: Agents API, GPT-Live-1, ChatGPT for Financial Services — The Agent Backend Becomes a Product

OpenAI Product Day Sept 10: Agents API, GPT-Live-1, ChatGPT for Financial Services — The Agent Backend Becomes a Product

TL;DR: On September 10, 2026, OpenAI bundled its agent backend into three products: The Agents API delivers the Codex harness as a managed service, GPT-Live-1 replaces the STT–LLM–TTS chain with a single full-duplex voice model, and ChatGPT for Financial Services brings licensed market data with citable references directly into the model. For builders, this means: calculate each layer individually — what do I pull first-party, what do I wire up myself?

The Day in a Table

ProductLayerStatusKey Takeaway
Agents APIOrchestrationPublic BetaCodex harness as an API: Context Compaction, Tool Search, Subagents
GPT-Live-1VoiceAvailable in APIFull duplex in one model, delegation to backend models
ChatGPT for Financial ServicesDataLaunchedPremium data (Daloopa, PitchBook, LSEG, etc.) indexed on OpenAI infrastructure

Agents API: The Codex Harness Becomes a Product

Useful agents need more than a strong model: they need context management, efficient tool usage, and subagent coordination. Exactly this harness, which powers Codex and ChatGPT for Work, is now available as the Agents API in public beta — with a single API call, defined by task, model, tools, and environment.

The most important building blocks:

  • Selectable Environment: OpenAI-hosted sandbox, custom infrastructure, or partners like Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop, Vercel, and Blaxel.
  • Automatic Compaction: As sessions approach the context limit, the API automatically summarizes earlier context. Workflows across multiple context windows no longer need custom compaction logic.
  • Tool Search: Loads only relevant tool definitions — this lowers token costs and preserves the model cache.
  • Programmatic Tool Calling: Agents call tools in parallel, chain operations, and filter results in code, instead of pulling everything back into the context.
  • Multi-Agent Support: Complex tasks are broken down into independent parts and delegated to parallel subagents; each maintains its own context.
  • Open Foundation: The API runs on the open-source Codex harness and supports MCP, custom functions, and built-in tools like web search.

The minimal call follows a simple pattern:

json
{
  "task": "Analyze the quarterly figures and create a comparison table",
  "model": "gpt-5.6",
  "tools": ["web_search", "code_interpreter"],
  "environment": "openai_hosted_sandbox"
}

GPT-Live-1: Full Duplex Instead of a Relay Race

Symbolic image: listening and speaking at the same time – one model instead of an STT–LLM–TTS chain.
Symbolic image: listening and speaking at the same time – one model instead of an STT–LLM–TTS chain.
Classic voice agents chain together speech recognition, reasoning model, and speech synthesis. Every handoff costs latency and disrupts rhythm. GPT-Live-1 hears and speaks simultaneously in one model — and delegates deeper reasoning and tool tasks to a backend model like GPT-6 Astra or a third-party model.

The measurable results from OpenAI's evaluations:

  • Full Duplex Bench: +30 percentage points over GPT-Realtime-2.1, especially in turn-taking latency and interaction behavior.
  • Tau3 (Intelligence of voice agents in end-to-end tasks): 1st place, paired with GPT-6 Astra for medium reasoning effort.
  • Real-world Value at Speak: Learners got more thinking time, interruptions dropped by almost 80 percent compared to turn-based systems.

Further strengths: Tone, pacing, and style can be controlled via the system prompt; background noise and pauses are processed without disruption; telephony is natively supported, from restaurant reservations to support IVR menus. Natively, the model outputs ASR transcripts and response text, and supports keyword biasing — with explicit turn detection, without being turn-based itself.

ChatGPT for Financial Services: The Data Layer as a Moat

Symbolic image: tracing numbers back to their source – the data layer as the real lever.
Symbolic image: tracing numbers back to their source – the data layer as the real lever.
On the same Thursday, ChatGPT for Financial Services launched, developed with Morgan Stanley and Evercore as design partners. The product runs on GPT-6 Astra and covers the daily life of an analyst: valuations, LBO models, buyer screening, earnings analyses, and pitchbooks with proprietary templates.

The real leverage is the data layer: Premium data from Daloopa, PitchBook, and LSEG News (including Reuters) is directly built-in, alongside integrations for S&P Capital IQ, MSCI, Moody's, Dow Jones Factiva, Preqin, Intapp, Datasite, and Box — without separate contracts or self-configured connectors. The data is indexed on OpenAI infrastructure, which is intended to improve retrieval, latency, and the new, more granular citations: numbers can be traced right back to the source.

As proof, OpenAI cites the OfficeQA Pro benchmark (U.S. Treasury bulletins with tables, charts, and footnotes): GPT-6 Astra achieves 69.9 percent, while its predecessor GPT-5.6 Sol scored 60.2 percent. Read honestly, this means: a roughly 30 percent error rate remains — so there is still room for improvement regarding numerical accuracy. A starting price was not published.

Sidenotes from the Same Day

  • Data Agent in ChatGPT Work: Structured data access directly within the work client.
  • Government Access: Government agencies receive a connection to the model line.

Checklist: Which Layer Do I Use First-Party?

  1. Orchestration: Do you need compaction, tool search, and subagents yourself? Then the Agents API is the shortest path — the hosted harness saves you maintenance.
  2. Voice: Is your STT–LLM–TTS stack stable but sluggish? GPT-Live-1 consolidates the chain into one model; weigh latency against cost per minute.
  3. Data: Does your company already use Daloopa, PitchBook, or LSEG? Check whether the first-party citations actually reduce your error rate before maintaining double connectors.
  4. Escalate: Only build your own models if late-night infrastructure or data residency requirements are not met by the API version.

FAQ

What is the difference between the Agents API and the existing Responses API?

The Agents API is a managed setup on the same open-source Codex harness that also powers Codex and ChatGPT for Work. You define the task, model, tools, and environment in one call, while the API handles context compaction, tool discovery, and subagent coordination. The Responses API remains the base primitive for direct model calls; the Agents API adds the orchestration layer on top.

Why is full duplex in voice agents an advantage over turn-based systems?

In turn-based chains, the system waits for a completed speech segment and then replies — pauses and interruptions cost latency every time. GPT-Live-1 hears and speaks simultaneously, reacts to interruptions and affirmations in real-time, and avoids the stiff handoffs of the STT–LLM–TTS chain. In practice, this means: shorter wait times, more natural dialogues, and almost 80 percent fewer disruptions in the learning test at Speak.

How do I check if ChatGPT for Financial Services replaces my data setup?

Compare the built-in sources with your existing licenses: Daloopa for financial statements, PitchBook for private company data, LSEG including Reuters for news. If your team already uses these services, you save on connectors and get more granular citations from an indexed source. Afterward, test on two or three real models whether the citation quality lowers your error rate — at 69.9 percent on OfficeQA Pro, there is still enough manual review required.

Is the Agents API worth it for teams with their own infrastructure?

Yes, because the environment is freely selectable: OpenAI-hosted sandbox for a quick start, VPC, or partner sandboxes like Cloudflare, E2B, or Modal for controlled setups. The main benefit lies in maintenance — OpenAI delivers compaction logic, tool search, and multi-agent coordination versioned with every model release. Only those mapping special cases beyond this harness are more flexible with a custom-wired stack.

What exactly does delegation mean for GPT-Live-1?

The voice model itself handles hearing and speaking, but passes reasoning and tool calls to a backend model, such as GPT-6 Astra for complex cases, or a cheaper model for high-volume tasks like scheduling. This allows you to separate speech latency from depth of thought and match costs per task class. The architecture thus remains modular without you having to orchestrate the chain yourself.

Sources

Relevant for your team? Let's talk for 30 minutes.

We sort out what of this actually works in your company — concrete, no slide marathon.

No commitment · 30 minutes · Proposal within 48 h