In short
OpenAI is expanding ChatGPT into AI agents that can complete office tasks across email, calendars and workplace apps. The move could deepen adoption and revenue, but it also exposes the biggest barrier in agent AI: users may not trust software with that much access.
- ChatGPT Work is OpenAI’s attempt to turn ChatGPT into a task-completing workplace agent.
- The product targets non-engineers and knowledge workers, not just developers.
- OpenAI sees higher token usage and broader adoption as major business upside.
- Trust, permissions and confusing setup remain the biggest obstacles to mainstream use.
- Anthropic helped define the market first, but OpenAI is trying to catch up with more powerful models and broader integration.
OpenAI is pushing its software beyond chat and into task execution, betting that AI agents that can read email, work across apps and complete office jobs will become the next mainstream use of ChatGPT. The company’s new ChatGPT Work product, released last month, is designed to help workers hand routine digital tasks to models — a move that could expand OpenAI’s market far beyond programmers while raising fresh questions about trust, permissions and whether most people want to give an AI that much control.
The shift matters because it moves artificial intelligence from answering questions to taking action. For OpenAI, that could mean more valuable subscriptions, longer usage sessions and a route into white-collar workflows that have so far been harder to crack than software engineering.
OpenAI wants ChatGPT to do more than talk
OpenAI’s latest workplace push is built around a simple idea: the most useful AI is not just a conversational partner, but a digital operator. ChatGPT Work is intended to let non-engineers use AI agents for multi-step jobs such as compiling reports, building dashboards, drafting planning materials and moving information between the many tools that define modern office life.
The product arrived last month and is available on OpenAI’s lower-cost subscription tier at $20 per month, making it an unusually accessible attempt to bring agentic AI to mainstream users. OpenAI describes the goal as a future in which artificial intelligence helps people transform ideas into outcomes, not merely produce text on demand.
That ambition echoes the success of AI coding tools, which already allow software engineers to delegate large chunks of work to models. OpenAI now wants similar behavior for accountants, investors, doctors, operations teams and other workers whose jobs revolve around email, spreadsheets, calendars and specialized business apps.
Why the company is betting on agents now
OpenAI’s timing reflects both product maturity and commercial pressure. Agents that stay active longer and handle multi-step tasks tend to consume more tokens, which can increase revenue per user. Just as importantly, OpenAI and its rivals need to prove that AI can add value in professions beyond software development if they want to justify the enormous cost of training frontier models.
The company’s leadership has framed this as an extension of OpenAI’s core mission: if the technology is going to change how people work, it has to reach more than just engineers.
Thibault Sottiaux, who oversees core product work including Work, said the objective is to build a system that can complete complex tasks autonomously while still feeling useful and safe for everyday users.
In practice, that means OpenAI is trying to move from a narrow, developer-friendly world into the messy reality of office software, legacy systems and fragmented workflows that were never designed for AI to control them.
How ChatGPT Work is different from a normal chatbot
ChatGPT Work is not just another chat window with a few extra features. It is built on a software layer often described by engineers as a “harness,” meaning the set of tools, rules and permissions that tells the model what it can see, what it can touch and how it should carry out a task.
That distinction is crucial. A chatbot can answer a question. An agent can use apps, move data and perform steps across a workflow. The harness makes that possible by connecting the model to the outside tools people already use every day.
OpenAI’s internal experience helped shape the product. According to the company, many of its own employees use agentic coding tools extensively, but those same tools were too technical and code-centric to be helpful for staff outside engineering. The challenge, then, was to make something that preserved the power of agents while hiding the complexity.
Why the interface matters so much
OpenAI executives say the user experience has to be simple enough for mainstream workers who are not comfortable with command lines, code diffs or technical readouts. That is one reason the company has leaned on a more visual, button-driven interface for Work, even if some internal voices believe users should simply ask the model directly.
The argument inside OpenAI is not really about aesthetics. It is about adoption. Engineers may be happy with bare-bones tools, but mass-market users often need prompts, labels and visible actions before they understand what the system is doing.
OpenAI’s team has compared this to older interface design patterns that once helped people move from physical objects to digital ones. The point is not that the old designs were beautiful; it is that they made unfamiliar software feel less alien.
What the data says about adoption inside and outside OpenAI
The gulf between internal enthusiasm and outside adoption is one of the clearest signs of how difficult this market could be. OpenAI-backed research found that in June, 98% of employees were using Codex, the company’s agentic coding tool, but only 17% of organizational subscribers and fewer than 1% of individual subscribers were using it.
That gap captures the opportunity OpenAI is chasing. The company has proven that agents can be indispensable in a technical environment. The unanswered question is whether the same can be true for ordinary knowledge workers who do not live in code all day.
OpenAI did not disclose how many users had adopted Work compared with Codex, but it did say the combined app has reached roughly 20 million users. That is far below the scale of ChatGPT itself, which the company says has surpassed a billion users online.
| Product | Main use case | Target user | Current scale or adoption signal |
|---|---|---|---|
| ChatGPT Work | Automating office tasks across email, documents and SaaS tools | White-collar workers and teams | Part of an app family used by about 20 million people |
| Codex | Agentic coding and software development support | Developers and technical teams | Used by 98% of OpenAI employees in June; far lower use outside the company |
| ChatGPT online | General-purpose Q&A and content generation | Broad consumer and professional audience | More than 1 billion users, according to OpenAI |
| Claude Code / Claude-style tools | Interactive coding and agent workflows | Developers and early adopters | Previously led downloads before Codex recently gained ground |
What kinds of work are agents already doing?
ChatGPT Work is being pitched for tasks that are repetitive, information-heavy and easy to break into steps. In OpenAI’s telling, that includes weekly reporting, spreadsheet-based planning, data cleanup, document assembly and basic coordination work that currently eats into the day of analysts and managers.
The examples are intentionally mundane, because that is where the immediate productivity wins are easiest to prove. Agents do not need to replace a full professional role to be useful; they only need to remove the most frustrating, time-consuming pieces of it.
Some early users are already treating the tool that way. Investors have reportedly used AI agents to collect communications and company research into memos. Operations teams are spinning up dashboards and visualizations. One OpenAI engineer described using an agent to turn a Slack discussion about a technical problem into charts.
OpenAI leaders also say the system is showing promise for personal tasks, not just work assignments. In one example, the company said Sam Altman has used it to organize travel. Another user managed to move a school calendar from email into Google Calendar, which may sound small but illustrates the category of tedious digital labor OpenAI wants to own.
Examples of tasks OpenAI says agents can handle
- Weekly metrics reports
- Spreadsheet organization and planning
- Investment memo preparation
- Dashboard creation
- Data visualization
- Calendar transfers and scheduling
- Research aggregation from multiple sources
Why trust and permissions are the real hurdle
The biggest obstacle to agent adoption is not whether the model can generate plausible output. It is whether users are willing to let it touch sensitive systems and whether they can understand the access it needs.
To be useful, an agent often needs broad permissions across email, Slack, cloud storage, calendars and work apps. That creates obvious risks: a model could surface information from a private message, misread a document, or take an action the user did not intend.
Andrew Ambrosino, the lead engineer for OpenAI’s desktop app, said he is comfortable granting that access for the sake of testing the product, even if it means occasional exposure to private information during development.
His stance reflects a broader truth about AI products: the people building them are often more willing than average users to accept imperfect behavior in exchange for learning what the tools can do.
For everyone else, the decision is harder. Many users are uneasy about handing over broad access to inboxes, work accounts and personal files. That hesitation slows down adoption and forces product teams to build much clearer controls, explanations and permission flows.
How the setup process can frustrate users
Even OpenAI’s own attempts to simplify access have run into friction. The setup flow for connecting cloud storage and other services can be confusing, and in some cases the system appears to require broader access than users expect before it functions properly.
That creates a classic AI product dilemma: the model becomes more capable as it gets more context, but each additional permission increases the user’s sense of vulnerability.
OpenAI says it is still refining those experiences. The company’s product and engineering teams acknowledge that the default settings and reasoning controls are not yet intuitive enough for newcomers.
How did OpenAI fall behind Anthropic in coding agents?
OpenAI was early to the idea of agentic coding, but Anthropic’s more interactive approach won the initial round. The difference came down to how much the product asked of the user versus how much it asked of the model.
OpenAI’s first version of Codex leaned heavily toward autonomy. Anthropic’s Claude Code, by contrast, was built to check in with users more often, present options and keep the conversation moving in smaller, safer steps. That human-in-the-loop design proved more practical when models and harnesses were less mature.
OpenAI now says that earlier version of its product assumed the model was more capable than it really was at the time. In hindsight, the company says, Anthropic’s interface was better matched to the limitations of the models available then.
OpenAI engineers described the earlier approach as too ambitious for the state of the technology at the time, while Anthropic’s format kept the user more involved and reduced the chance of errors.
As models improved, OpenAI shifted toward a more interactive desktop and mobile experience. That helped Codex regain ground, and recent usage trends suggest it has now moved slightly ahead of Claude Code in demand, at least by some download-based proxies and enterprise indicators.
What competition is doing to product design
Competition is not just about branding. It shapes how these products are built. If one vendor’s agent feels safer, clearer or easier to direct, that can affect adoption as much as raw model quality.
OpenAI’s engineers insist the biggest advantage still lies in the strength of its latest models. But rival products have forced the company to think harder about the balance between autonomy and supervision.
That tension is likely to define the next phase of the agent market. Too much autonomy risks confusion and mistakes. Too much hand-holding makes the product feel less magical and less efficient.
What is a good harness, and why does it matter?
A good harness is the thin layer of product design that helps a model use the right tools, access the right information and avoid unnecessary complexity. In AI engineering terms, it is the bridge between raw model capability and practical utility.
OpenAI’s harness team says the best systems are often the simplest ones. Their view is that if the model is strong enough, developers should not bury it under too many rules or brittle shortcuts. The job is to expose just enough context and just enough tools for the model to solve the task effectively.
That philosophy comes from a long-standing principle in AI research: general-purpose models tend to improve so fast that overly customized add-ons can become obsolete quickly. Build too much fixed logic, and the next model release may make it redundant.
Why the “bitter lesson” matters here
The so-called bitter lesson in AI is the idea that scalable learning and general models ultimately outperform hand-engineered tricks. In the context of agents, that means product teams may get only a short-lived advantage from complicated chains of if-then rules or excessive workflow scripting.
OpenAI’s harness engineers say they are trying to strike a balance between giving the model enough structure and preserving its ability to reason flexibly.
Still, there is no consensus that the market is ready for so much autonomy. Some experts believe users prefer systems that ask more questions, show more comparisons and leave more room for oversight.
Ethan Mollick, a Wharton professor who studies workplace AI, has argued that ChatGPT’s style is more aggressively automatic, while Claude’s approach feels more iterative and user-directed.
How OpenAI measures whether Work is actually useful
OpenAI says it evaluates these systems with a benchmark called GDPVal, which draws on dozens of occupations and hundreds of knowledge-work tasks. It also uses user feedback to see where the product succeeds or fails in practice.
That matters because unlike coding, much of office work is hard to measure cleanly. A program either compiles or it does not. A good memo, presentation or strategy document is far more subjective.
OpenAI’s internal teams say they still rely heavily on their own usage to figure out what works. That means employee habits may continue to shape the product, at least in the short term, even as the company tries to design for a much broader audience.
| Milestone | What happened | Why it matters |
|---|---|---|
| Early Codex | OpenAI launched an early version focused on autonomous coding | Proved agentic software could help developers, but the interface was too ambitious |
| Anthropic’s response | Claude Code used a more interactive, step-by-step design | Showed that user involvement could improve reliability and adoption |
| OpenAI redesign | Codex became more interactive on desktop and mobile | Helped OpenAI recover market traction |
| ChatGPT Work launch | OpenAI extended agents to broader office tasks | Marks the company’s attempt to move beyond developers into mainstream white-collar work |
What OpenAI still has to solve
OpenAI’s challenge is not simply building a more powerful model. It is creating a product people trust enough to use daily, even when it touches private messages, business systems and personal schedules.
That means solving for discoverability, clearer permissions, better defaults and more useful guidance when the agent gets stuck. It also means deciding how much autonomy users actually want. For some tasks, the appeal is obvious. For others, people may prefer a system that asks before acting.
The company also has to answer a deeper market question: if the majority of professionals only need occasional help, will they pay for an agent that works continuously in the background? OpenAI appears to believe the answer is yes, especially if the tool saves enough time and makes enough information actionable.
But broader adoption will likely depend on something more delicate than raw intelligence. It will depend on whether the AI feels understandable, reliable and worth trusting with the parts of digital life that users currently manage themselves.
What happens next?
The next phase of OpenAI’s agent strategy will likely be less about demonstration and more about workflow integration. The company is trying to make its tools feel less like experiments and more like assistants that can be dropped into real work without extensive setup.
That may require more education, better onboarding and a clearer answer to the question at the center of all of this: how much control are people willing to surrender to an AI system in exchange for convenience?
For OpenAI, the answer will shape both product design and revenue. For workers, it may determine whether AI agents become everyday collaborators or remain a niche tool for power users and the most automation-friendly teams.
For now, OpenAI is making the case that the future of AI is not just a better chatbot. It is software that can actually do the job.
Frequently asked questions
What is ChatGPT Work?
ChatGPT Work is OpenAI’s workplace-focused AI product designed to help users complete multi-step office tasks across apps like email, calendars, Slack, documents and spreadsheets. It goes beyond chat by using AI agents that can take actions, not just generate answers.
Why is OpenAI focused on AI agents now?
OpenAI is focused on AI agents now because they can deepen user engagement, consume more tokens and expand the company beyond coding into mainstream white-collar work. That opens a much larger market than software development alone.
How is ChatGPT Work different from regular ChatGPT?
ChatGPT Work is different from regular ChatGPT because it is built to connect to external tools and carry out workflows across them. Instead of simply responding to prompts, it can retrieve information, organize data and perform tasks on a user’s behalf.
What is the biggest risk with AI agents in the workplace?
The biggest risk with AI agents in the workplace is trust and access. To be useful, they often need broad permissions across sensitive systems, which can expose private messages, business data or unwanted actions if the user setup is too loose.
Is OpenAI ahead of Anthropic in AI agents?
OpenAI is not clearly ahead across the board, but it is catching up. Anthropic’s Claude Code helped define the market early, while OpenAI’s Codex and ChatGPT Work are now gaining ground thanks to product changes and stronger model performance.









