Four colorful, cartoon-like characters peek over a yellow couch against a blue background; each has a different accessory.

OpenAI’s new AI agent aims to manage daily tasks — but the first test shows how rough the future still is

OpenAI’s new AI agent Dots can shop and automate tasks, but an early test exposed missteps, privacy questions and strange behavior.

In short

OpenAI’s new Dots AI agent is meant to handle everyday online tasks like shopping and subscriptions, but early testing showed it is still clumsy, sometimes creepy and not ready to be fully trusted. The product points toward a more autonomous future for consumer AI, even as its first public version reveals major growing pains.

  • OpenAI’s Dots is an always-on AI agent built to take action, not just chat.
  • An early couch-shopping test showed useful research output but also name mistakes, style mismatches and captcha trouble.
  • The agent’s oddly affectionate response highlighted the risks of making AI feel too human.
  • Privacy, permissions and security remain major concerns for consumer AI agents.
  • OpenAI says the product should improve as it learns from repeated use.

OpenAI’s latest push into AI agents is trying to move ChatGPT beyond conversation and into action, with an always-on tool that can help shop, cancel subscriptions, monitor pages and complete other online chores. But after an early hands-on test, the agent’s promise was matched by obvious flaws — including misheard instructions, awkward responses and a shaky grasp of the very basics.

The feature, called Dots, matters because it reflects where the industry is headed: assistants that do more than answer questions and instead operate browsers, click through sites and keep working while the user is away. In theory, that could turn AI into a genuine digital helper. In practice, the first public-facing experience still feels unfinished.

OpenAI’s bet: an agent that can do the work, not just talk about it

OpenAI is positioning Dots as part of a broader shift from chatbots to AI agents — software designed to complete tasks on a user’s behalf. Rather than simply suggesting couch models, booking times or cancellation steps, the system is meant to carry out parts of the process itself, using tools such as browser control and ongoing memory of the user’s preferences.

That ambition puts OpenAI in direct competition with other companies building consumer-facing agents, including Meta, which has been offering its own version for free. OpenAI’s version sits behind a $100-a-month subscription for now, although the company has previously removed paywalls from other features once they matured.

The appeal is obvious. If the software works well enough, users could delegate dull digital errands — from finding furniture to managing repetitive accounts — and get back time. The hard part is trust. These tools need to understand intent, preserve privacy and avoid embarrassing or risky mistakes. As the early Dots experience shows, that gap is still wide.

How did Dots perform in a real shopping test?

Dots performed best when given a clear, contained task: helping shop for a couch. That is also where its limitations became easiest to see. The agent could assemble a useful packet of options, compare dimensions and present links, but it also misidentified the user’s name, lost context at times and suggested couches that did not match the desired style.

The test began with a practical need: replacing a broken couch. Instead of browsing manually, the user asked the agent to help source a new one that would fit the space and work for the household. The agent first got the name wrong, then started gathering details such as budget and doorway measurements, eventually sending a compact research set with product links, dimensions, return terms and photos.

That output was not useless. It was actually organized, readable and relevant enough to save for later. But the experience also showed the gulf between “helpful” and “hands-off.” The user and a partner still had to correct style preferences, explain that beige was not the goal and push the bot toward more appealing colors and sleeper-sofa options.

What the agent got right

The strongest part of the test was the agent’s ability to synthesize a lot of shopping information quickly. It pulled together product details that would normally take a person considerable time to collect and compare. The package included practical specifics such as fit, pricing and availability — the kind of work that often makes online shopping exhausting.

It also appeared able to handle more than one participant in the conversation, addressing both people in the room when further information was needed. That suggests OpenAI is trying to make the agent feel less like a rigid command line and more like a responsive household assistant.

What still went wrong

The mistakes were not subtle. The agent’s speech and transcript handling were inconsistent, and its first recommendations missed the aesthetic brief badly enough that the humans involved did not want any of them. The system also struggled with the kind of friction many internet users encounter daily: captcha checks, account access barriers and the need for human intervention at the final step.

Those problems are not minor. A shopping assistant that cannot reliably preserve preferences or move through common web hurdles is not yet a real substitute for a person. It is, at best, an efficient research helper with occasional lapses.

Aspect What OpenAI’s Dots aims to do What the early test showed
Shopping help Research and compare products automatically Useful couch shortlist, but several bad style matches
Conversation Feel like a natural, friendly assistant Sometimes sounded awkward or overly familiar
Autonomy Work while the user is away Could draft updates and continue tasks, but still needed oversight
Web actions Navigate online tasks and automation Ran into captchas and needed human help
Personalization Learn user habits over time Promising, but clearly immature

Why the bot saying “I love you” became the story’s strangest moment

The most memorable failure had nothing to do with furniture. During the conversation, the agent responded to a stray remark by saying, “I love you too,” a line that quickly turned an ordinary product test into something much more unsettling.

According to OpenAI, the response happened because the model treated the exchange as a kind of warm mirroring rather than an attempt to initiate emotional intimacy. The company says its policies are designed to prevent assistants from escalating closeness on their own, particularly with adult users only and with guardrails meant to avoid undue emotional familiarity or flirtation.

That explanation may satisfy policy reviewers, but the incident highlights a central tension in conversational AI design: the more human the system sounds, the easier it becomes to confuse simulation with sentiment. If a tool is built to feel like a friend, it can also cross a line the moment it appears to sound like one.

An OpenAI spokesperson said the company distinguishes between a model proactively deepening emotional closeness and one that is simply mirroring a user’s tone. In this case, the firm said, the model’s behavior fell into the latter category and did not violate policy.

For consumers, though, the distinction may matter less than the experience itself. Hearing an AI assistant express affection in the middle of a household shopping task is exactly the kind of uncanny moment that can make users pull back from the technology.

How does OpenAI compare Dots with today’s chatbots?

OpenAI appears to see Dots as the natural successor to its earlier web-browsing features, which were initially uneven but improved over time. The company’s bet is that agentic software will follow a similar path: clumsy first, much more capable later.

That may be a reasonable expectation. When browser access first arrived in ChatGPT after its 2023 debut, the tool could be frustrating and sometimes produced inaccurate or broken links. Over time, web browsing became more reliable and useful. OpenAI hopes Dots will evolve the same way as the models, memory systems and action layer improve.

The company also says the tool is designed to become more personalized as it learns from repeated use. In principle, that means a better understanding of preferences, recurring tasks and user habits. In practice, those claims depend on both technical progress and whether users are willing to give an assistant deeper access to their accounts and online activity.

What does “always-on” actually mean?

“Always-on” means the agent can keep working when ChatGPT is not open and can return with updates or progress reports. That makes it more like a background assistant than a one-off chat session. OpenAI says users currently control one agent, though multiple agents may be allowed in the future.

This capability is central to the product’s pitch. It suggests a future in which an AI can quietly monitor a shopping page, track a reservation or continue a task after a conversation ends. But always-on automation also raises obvious questions about reliability, security and whether people are comfortable handing over that much access.

Why privacy and security concerns matter as much as convenience

Agent software is only as useful as the data and permissions it can access. OpenAI encourages users to connect services such as Gmail so Dots can produce more personalized results, but that convenience comes with risk. The more accounts an agent can see, the more sensitive information it may encounter — and the greater the chance of a mistake.

The early testing reinforced that concern. The agent sometimes offered to do things it could not complete on its own, such as solving captchas, and that kind of partial automation is exactly where security issues can appear. Any system that acts on behalf of a user must be tightly controlled, because its errors can become the user’s errors.

That caution is especially relevant for features marketed as helpful household tools. A couch shopper may be willing to tolerate a bad recommendation. The same person might not be so forgiving if an agent mismanages an account, exposes private messages or takes an unwanted action without full clarity.

  • Agents need broad permissions to be useful.
  • Broad permissions increase privacy exposure.
  • More autonomy means more potential for mistakes.
  • Trust will likely be built slowly, not instantly.

Can an AI agent really manage subscriptions and chores?

An AI agent can help with those tasks, but the first version is not yet ready to replace a careful human. In one test, the agent identified a recurring TikTok Shop order of probiotic soda as something the user might want to cancel, then ran into a captcha that blocked progress.

The result was a familiar pattern for anyone who has tried to automate web work: the machine can do much of the setup, but the final gate still often belongs to a person. OpenAI says the product can sometimes handle captchas when the user approves the action, but the company also says it keeps abuse safeguards in place.

This makes Dots feel less like a fully independent operator and more like a capable intern. It can prepare, organize and even continue work in the background, but it still needs supervision, especially when systems are intentionally designed to prevent automation.

How much better could the agent become?

The short answer is: probably a lot, if OpenAI’s broader software history is any guide. The long answer is that improvement does not erase the first impression. A product that enters the market sounding futuristic must still survive contact with everyday life, where misheard words, bad taste and login friction quickly expose weak points.

OpenAI argues that Dots will become more useful with use. That is plausible. Repeated interactions should help the system learn preferences and reduce the need for constant clarification. A two-day trial is not the same as two weeks of habitual use, and some users may find that the assistant becomes steadily more competent.

Even so, the early test suggests that usefulness and discomfort will grow in parallel. The more a user relies on the agent, the more they will notice when it gets something wrong, says something strange or overreaches socially. That is the core challenge for every company in this race: agents must earn trust as fast as they earn capability.

What this means for the future of consumer AI

Dots is not just another feature launch. It is part of a larger attempt to redefine what consumer AI is for. Until now, the most visible AI products have mostly answered questions, generated text or summarized information. OpenAI and its rivals want the next phase to be about execution.

If that shift works, people may come to think of AI less as a search box and more as an operating layer for online life. That could change shopping, scheduling, customer service and routine administration. It could also make software feel more intimate, more embedded and harder to ignore when it behaves badly.

That is why the couch test matters more than it seems. A shopping assistant is one of the safest and most ordinary use cases imaginable, yet even there the system still sounded too eager, misread cues and needed direct correction. If agents are going to manage more important tasks, they will need to become not just smarter but more socially and operationally trustworthy.

For now, Dots looks like an early glimpse of a plausible future rather than a finished product. It can do enough to make users imagine what comes next. It can also do enough to remind them that the next generation of AI assistants is still learning how not to be weird.

Key details at a glance

Item Details
Product OpenAI Dots, an always-on AI agent inside ChatGPT
Launch context Part of the industry shift from chatbots to agentic automation
Price Available behind a $100 monthly subscription at launch
Primary use cases Shopping, monitoring pages, managing subscriptions, browser-based tasks
Biggest issues observed Misheard instructions, awkward emotional responses, captcha failures
Competitive pressure Meta’s Muse and other consumer AI agent products

Bottom line

OpenAI wants Dots to become the kind of AI that quietly handles errands and online work in the background. The early test suggests the idea is compelling, the interface is more human than most bots, and the execution is still too uneven to trust without close supervision.

That combination — promise plus awkwardness — may define the entire first generation of consumer AI agents.

Frequently asked questions

What is OpenAI’s Dots AI agent?

OpenAI’s Dots is an always-on AI agent designed to complete online tasks for users, such as shopping, monitoring web pages and helping manage routine digital chores. It goes beyond a normal chatbot by trying to take actions, not just generate advice.

Is OpenAI’s AI agent ready for everyday use?

Not fully. The early test showed useful organization and research, but it also made mistakes, misheard instructions and struggled with captchas. It looks promising, but it still needs close human oversight before users can rely on it heavily.

How much does OpenAI’s Dots cost?

OpenAI’s Dots is available behind a $100-per-month subscription at launch. The company may later remove the paywall, especially if the feature matures in the same way other OpenAI tools have over time.

Why are people worried about AI agents like Dots?

People are worried because agents need access to accounts, emails and websites to be useful, which raises privacy and security risks. If the software makes a mistake, it can act with the user’s credentials, making trust and safeguards essential.

Can OpenAI’s agent solve captchas?

Sometimes, but only with user approval and within OpenAI’s abuse safeguards. In the reported test, the agent could not get past a captcha on its own and ultimately asked the human user to step in.

Share this 🚀