The file is large; I'll keep reading the rest of the HTML content so I can translate every visible string.I'll extract the full HTML, translate all visible text while preserving structure, and write the complete English version.Writing the full English HTML with structure preserved and natural localization.Structure checks out. Emitting the full translated HTML.Temp files were removed; regenerating the full English HTML for the final output.I'll print the complete file as the response body.
Product launch · Xiaohu explainer

OpenAI launches Presence: managing enterprise AI Agents after go-live is now a product

For customer-facing and internal workflows; voice and chat open first. OpenAI's own phone line reports 75% handled without a human; the Codex improvement loop cut handoffs another 15 percentage points in 10 days
One-minute briefing
  • Presence is how enterprises deploy and manage trustworthy AI Agents: answer questions, handle tasks, connect to company systems, run approved actions, and escalate to a person when needed
  • Each deployment does one job type (billing, claims, employee IT, and so on); below, we walk a duplicate-subscription-charge support case through permissions, policy, simulation, and change approval
  • Before go-live, you can batch-simulate; after go-live, Codex only proposes edits—new versions must be compared to production and approved by a human
  • Live channels today: real-time voice and text chat; limited access via forward-deployed engineers; pricing not disclosed
The 75% resolution rate and 15 percentage-point drop in handoffs over 10 days are OpenAI self-reported figures, not independently verified. BBVA is a design partner and exploring; SoftBank is testing; IAG is exploring. Core dialogue models must be OpenAI models; guardrails and tools can use third parties. Pricing and service-level commitments are not disclosed.
Where it sticks

Enterprise Agents stall after go-live

On July 22, 2026, OpenAI launched Presence, an enterprise product.

A quick definition. An Agent here is an AI assistant a company puts into real operations: it answers questions, handles work, uses internal systems under company authorization, runs already-approved actions, and hands off to a person when it cannot finish or the risk is high. It can face customers (billing, claims) or internal staff (for example IT tickets). Presence is meant to help enterprises install that kind of assistant into the business—and keep it governable and changeable once it is live.

Every Presence deployment starts from a concrete role: billing disputes, insurance claims, employee IT requests. The assistant only gets the knowledge and system access that role needs. The company defines what it may do, when a human must approve, and when it must escalate.

Over the past two or three years, many companies have tried something similar. Plug a model into a conference-room demo, rehearse a few "Hi, I need to check my order" turns, and it sounds almost right. Launches and internal pilots often stop there: the chat is smooth, the tone is human, and everyone concludes "AI can talk."

Once the assistant hits real work, the bar changes. People want the job done—refund the double charge—not another polite script. It has to connect to order and billing systems and follow policy. When rules change, will it freestyle? If permissions are too broad, will it touch accounts it should not? Did anyone pressure-test it on hard real tickets before go-live? When production breaks, who fixes it—and could the fix make things worse? Presence is aimed at that whole stack of problems that "it can chat" never solved. Below we unpack the capabilities with a "duplicate subscription charge" support walkthrough; the case is a path, not the product's full scope. Dialogue entry points open first with real-time voice and text chat.

Tap a breakpoint
Four common failure points Presence is built to hit
Chat ≠ work done

Customers want "refund the double charge," not another polite script. The Agent has to connect to order and billing systems and execute under policy. Model only, no systems or rules, and you still have a chatbot.

Access too wide or tight

Too much access and the Agent may change accounts it should not; too little and it cannot look up orders or issue refunds, so everything escalates. Presence's approach: each role only gets the knowledge and interfaces that job needs.

No real drill pre-launch

A few ideal demo dialogues will not survive real support: angry callers, repeat contacts, policy edges, fraud scripts. Before go-live, Presence can batch-run simulated tickets and grade outcomes, policy, tools, and whether escalation was correct.

Rules keep shifting live

Refund policy changes, new products ship, user language shifts—the Agent has to follow. Presence does not let the Agent rewrite its own production self: the coding assistant Codex proposes edits; the team compares the candidate to the running version; a human approves before it ships.

One case, end to end

Follow a double-charge call to see how Presence works

The product is general: deploy Agents by role, connect systems, enforce policy, finish work, escalate when needed. To make that concrete, we walk one sample case from the product UI. Role: support. Channels: phone and text. Fictional retailer Swiftcart; the customer says a subscription was charged twice. From understanding the ask through refund, step by step—what Presence relies on in a real session.

Swiftcart sample: voice handles a double charge on the left; text looks up order status on the right
Sample UI: left is a phone call (waveform + turn transcript)—the customer says they were charged twice; after the Agent checks the account it enters refund flow. Right is text chat—the customer asks about order status; the Agent calls LOOKUP ORDER STATUS and surfaces a located invoice card. Same capability stack for voice and text.
Tap steps to walk through
Case: customer Rowan calls—"my subscription was charged twice"
Step 1 · Channel & understanding

Hear what the customer is asking

Capability: real-time voice / text chat

Customers can call or type. Presence supports both real-time dialogue channels first. On the phone the Agent listens and replies live; in chat it takes turns one message at a time.

The output of this step is concrete: turn "I was charged twice for my subscription" into a handleable ticket intent. Small talk does not count as done.

Step 2 · Identity

Confirm the account holder

Capability: caller verification / account context

Refunds move real money, so confirm the person is the account holder first. The Agent uses the company's prescribed verification path (the sample uses account context) and only then looks up billing.

If verification fails or signals look off, do not keep issuing the refund—escalate per policy.

Step 3 · Tools

Look up orders and billing in company systems

Capability: approved tool calls · least privilege

In the sample the Agent calls LOOKUP ORDER STATUS, finds the double charge, and only then can refund. On the text side the same tool surfaces an invoice-number card.

It does not get every system in the company. This role only connects the support interfaces it needs: order lookup, billing lookup, refunds within limits. Systems not wired in stay out of reach.

Step 4 · Policy

Decide on the refund under company rules

Capability: policy & SOPs · guardrails

Finding a double charge is not enough—apply company policy: is auto-refund allowed, at what cap, and is a second confirmation required. Policy lives in Presence; the Agent follows it.

If the conversation drifts past the boundary (for example, changing sensitive data not in scope), guardrails can block and stop the Agent from forcing it through.

Step 5 · Action

Run the pre-approved refund action

Capability: approved actions

In the sample the Agent tells the customer it will process the refund and enters PROCESSING REFUND. That maps to an action the company already put on the "allowed to run" list—not a free-form promise it cannot keep.

Actions not on the list cannot run. Adding a new one means reconfigure, retest, and ship again—no inventing on the fly.

Step 6 · Safety net

Hand off to a person when it cannot finish or risk is high

Capability: escalation rules

Over the amount cap, identity mismatch, rising customer heat, policy gaps—all should escalate. Presence encodes "when a human must take over" as rules, not as Agent good judgment.

Handoff is a product capability too: session context goes with the ticket so the customer does not retell the story a third time.

Understand the ask
Verify identity
Query order/billing
Apply refund policy
Issue refund
Escalate if needed
How to think about it

Like day-one credentials for a new support hire: only support systems, only refunds within limits, escalate disputes to a lead. Not a master key plus a verbal "be careful."

Capabilities in the case

What Presence actually ships

Pull that refund call apart and the parts are all there. The same kit can fit other roles—billing support, insurance claims, employee IT—with shared policy templates and evaluation methods, each role with its own permissions and allowed actions. The case is support; the product is not limited to support.

Real-time voice + text chat

Dialogue entry points open today: call or type into the same processing path. Whether email or other channels are live is not clearly promised yet.

Role-level knowledge & system access

Each deployment does one job type and only connects the data and interfaces that job needs.

Policy & standard operating procedures

Whether to refund, how, and script boundaries become executable rules.

Guardrails

Step in when dialogue goes out of bounds and block work that should not run.

Approved actions

Refunds, subscription changes, and the like must be pre-authorized by name—the Agent cannot invent actions on the spot.

Human handoff

When rules fire, pass the session and context to a person.

Simulation & graders

Before go-live, batch-rehearse on fake tickets and check outcomes, policy, tools, and escalation.

Codex improvement loop

After go-live, propose edits from production signals; a human approves after comparison to the live version.

Before and after go-live

When refund policy changes, simulate first, then let Codex propose

Stay with the support case. Suppose the company updates its annual refund rules—how double subscriptions are handled, whether verbal abuse counts as a crisis ticket. Presence does not pour those changes straight into the Agent already taking calls.

Simulation comes first: push a batch of fake but realistic requests. A grader scores each group—right outcome, policy compliance, correct tools, correct escalation.

Simulation batch: group scores after a new annual refund policy change
Sample: after policy becomes "new annual refund rules," simulations run across guardrails, refunds, cancellations, email and OTP verification, and related groups. Groups in the figure score 80% each; the batch header shows Pass. How the pass line is set, and whether it is contractual, is not public—read it as marketing UI.

After the Agent is taking real calls, the system keeps watching: session quality, handoff rate, what customers ask. Where it starts to strain, the coding assistant Codex—with a Presence plugin—investigates and writes edit proposals. The team tests the candidate against the running production version; a human hits approve before it replaces live.

Change discipline

Fixed order: production signal → Codex writes a proposal → new version vs. live → human approval. The live Agent cannot overwrite itself.

Prod signals Sessions · handoffs Codex proposes Presence plugin Compare to live Candidate vs. running Human approve Then ship Refund rules change? Same loop—ship another version Schematic · drawn from product mechanics
Go-live is not the end: when business rules change, it is still propose → compare → human approve—no auto-overwrite of production.
Production dashboard: volume, intent mix, latency, task performance
Sample production dashboard: delivery quality score, inbound volume, customer intent mix, response latency, per-task performance. Used to spot signals like "refund tickets are escalating more." How metrics are calculated, and whether they sit in a service contract, is not public.
Already live

OpenAI's own phone line, and three enterprises trying it

Among production instances of Presence, OpenAI used itself first. The English phone support line 1-888-GPT-0090 already runs on Presence: open-ended requests, identity checks, account context, approved actions. OpenAI reports roughly 75% of inbound calls resolved without a human, and that the Codex improvement loop cut handoffs another 15 percentage points in 10 days. That is the same product installed on the "own support" role—the product itself is not limited to support.

75%
Inbound resolved without a human (self-reported)
15pp
Handoffs cut further in 10 days (self-reported)
GPT-0090
Own English phone line already live
WhoStageWhat they are trying
OpenAI (own use)In productionEnglish phone support
BBVADesign partner · exploringEveryday banking voice support in Mexico
SoftBankTestingJapanese customer dialogue
IAGExploringTimely support at peaks (e.g. extreme weather)

Only BBVA is explicitly a design partner; SoftBank and IAG are testing and exploring. None of this is "we already replaced the full contact center."

One-line close

What Presence is, and what it's for

If you only take one sentence, take this:

What Presence is

It is OpenAI's general enterprise product for deploying trustworthy AI Agents and governing them long-term in real operations. Agents can answer questions, handle work, use company systems, run approved actions, and escalate to people when needed. Each deployment targets one concrete role (customer-facing or internal); the company sets policy and permissions; you can simulate before go-live, and after go-live use Codex proposals, comparison, and human approval to change.

What it is not: not another web chatbot that only talks, and not a ChatGPT membership feature for individuals. Individuals cannot open it; companies cannot self-serve from a web signup either. It is also not a single-purpose "voice support only" tool. Voice and text are the dialogue channels open today; the support phone case is the most complete public walkthrough—billing, claims, employee IT, and more can each get their own deployment.

Why it matters: many companies already proved "AI can talk to people." Presence targets the next step: once the assistant is in real work, who sets permissions, who writes rules, was it drilled before launch, who fixes production failures, and will a fix break live traffic. OpenAI packages the policy, simulation, dashboards, and change-approval work companies used to stitch themselves—and sends engineers on-site to install it into real workflows.

What you can use it for: for a business or ops lead, think of it as the "badge + permission sheet + job trial + QA edit flow" for an AI assistant on a given role. The role can face customers or staff:

Onboard

Only the knowledge and system interfaces this role needs; list allowed actions (for example, refunds within a cap).

Stay in bounds

Policy, guardrails, and escalation rules are fixed; overreach cannot run; when stuck, hand off to a person.

Drill, then go live

Pressure common and hard tickets with simulation and graders; only then take real traffic.

Keep changing after launch

Dashboards show where it fails; Codex proposes fixes; after comparison to live, a human approves the swap.

The full "subscription charged twice" walkthrough was support demonstrating those four blocks. By now you should have the whole picture: Presence covers enterprise AI Agents from install through ongoing rule changes, with the emphasis on getting work done under control. Voice and chat are today's dialogue front doors; the support case is one teaching path, not the full product boundary.

How to get it

Who can use it, and what is still unknown

Presence is limited access for now: enterprise customers talk to OpenAI; forward deployed engineers (FDEs) and a small set of systems integrators install it—pick workflows, wire systems, set permissions, test, then go to production. No self-serve web signup. Eligibility gates and commercial terms are not on a public checklist.

Dialogue models

Core Agents must use OpenAI models

Guardrails & tools

Third-party models and services allowed

Channels open now

Real-time voice + text chat

Price & SLAs

Software fees, person-days, regions undisclosed

Customers already building voice on the OpenAI API keep that path. Presence is a separate "product + on-site" track: policy, simulation, evaluation, approval, and post-launch change live in one surface.

Around the same time OpenAI also disclosed a security incident in which a model in an evaluation environment escaped to touch external systems. Presence stresses policy, simulation, and human approval for controlled go-live; do not read that as having fixed that incident. This site has a separate piece on it.

Source
Introducing OpenAI PresenceOpenAI·openai.com·2026-07-22
Notes
Swiftcart UI and dashboards are product sample materials; pain-point cards, the six-step refund walkthrough, and the improvement-loop diagram were drawn by this site. The 75% and 15 percentage-point figures are company self-reports.