AI agents on a phone with email and calendar access

Why One Writer Finally Trusted an AI Agent With Her Inbox—and What It Reveals About the Risk

AI agents like Instinct can book travel and manage email, but new testing shows serious privacy, security and reliability risks.

In short

A hands-on test of the invite-only AI agent Instinct found that it can genuinely save time by handling bookings, travel changes and message triage. But the same access that makes AI agents useful also exposes users to privacy, security and financial risks.

  • Instinct can complete real-world tasks through messaging apps, including reservations and travel changes.
  • The agent’s strongest advantage is its ability to notice details users may miss, like flight schedule changes.
  • Security and privacy risks remain significant, including data retention concerns and phishing exposure.
  • The broader AI agent market is growing fast, with consumer appetite rising alongside investor interest.
  • Convenience may be enough to drive adoption, but trust will determine whether these tools go mainstream.

An AI agent that can read email, manage calendars and take actions inside messaging apps is starting to look less like a novelty and more like a product people may actually use, but the trade-off is obvious: convenience comes with real security and privacy exposure. In a recent hands-on test, WIRED’s Zoë Schiffer found that the invite-only agent Instinct could book restaurants, rearrange travel and even spot a phishing email, while also raising serious questions about data handling, authorization and user safety.

The story lands at a pivotal moment for consumer AI. Instinct is drawing attention in the Bay Area just as Meta’s own assistant, Muse, is surging in the App Store, underscoring how quickly AI agents are moving from demo-stage curiosity to mainstream experiment. Yet the same tools that promise to erase administrative drudgery also create new risks when they are given access to email, calendars, payments and other sensitive accounts.

For people who have tried a chatbot and found it useful only in the narrow sense of answering questions, the appeal of an agent is clear: it can act. But as Schiffer’s experience suggests, the real challenge is not whether the technology can do things; it is whether users can trust it to do the right things, for the right reasons, without creating a bigger mess than the one it was meant to solve.

What Instinct is—and why it is getting attention

Instinct is an invite-only AI agent that operates through familiar messaging channels such as iMessage and WhatsApp rather than through a traditional web interface. It connects to core personal tools including email, calendars and messaging apps, aiming to function like a human assistant that can both understand a request and carry it out.

The company launched its private beta in February and, according to reporting cited in the original piece, is now in talks to raise another $1 billion after previously securing $350 million. That would put its valuation at around $10 billion, a striking sign of investor confidence in a category that is still trying to prove it can become part of everyday life.

Its rise comes amid a broader wave of enthusiasm around AI assistants that can take real-world actions. In this case, attention has been amplified by the contrast with Meta’s Muse, which has also gained popularity quickly while simultaneously facing scrutiny over security flaws. The overlap matters because it shows how quickly the market is rewarding products that feel immediately useful, even when their underlying safeguards remain underdeveloped.

Why the form factor matters

One of the biggest lessons from Instinct’s early reception is that interface design can determine whether an agent feels usable or merely impressive. Instead of dropping users into a blank chat box and forcing them to invent a workflow, Instinct lives inside apps people already use every day and suggests actions it can take on their behalf.

That change may sound minor, but it addresses one of the biggest problems with many AI tools: users often do not know what to ask for, or when the model should be doing the work instead of just talking about it. By embedding itself into messaging, Instinct lowers the friction between intention and execution.

The central appeal of an AI agent, as Schiffer’s experience shows, is not simply conversation but delegation: the tool proposes a task, performs it and then reports back.

How did the AI agent perform in real life?

It performed better than expected in a set of everyday tasks that many people would consider annoying, time-consuming or both. In the writer’s test, Instinct handled restaurant reservations based on a travel itinerary in Venice, communicating through WhatsApp with smaller eateries and booking the plans successfully.

That kind of use case is important because reservations are a relatively low-stakes proving ground for AI agents. They require some judgment, some persistence and some interaction with external systems, but they do not immediately expose the user to the most consequential risks of a fully autonomous system.

The more revealing test came with travel. After the writer had already booked a New York trip for November, WIRED asked her to attend an event in Miami the day before. Because the ticket was a saver fare, she could not simply alter the itinerary without risking the whole booking. Instinct noticed that Alaska had shifted the flight schedule by 90 minutes, a change the user had missed, and that detail qualified the ticket for a full refund.

The agent then canceled the flight, reportedly saving about $550, and rebooked the return leg separately. That result is the strongest argument yet for the practical value of consumer AI agents: they can watch for edge cases, notice policy-triggering changes and turn a frustrating administrative chore into a solved problem.

What made the travel example stand out?

It stood out because the agent did not just follow instructions; it identified a condition the user had overlooked. That means the system was not merely acting as a remote button-pusher, but as a kind of policy-aware assistant capable of recognizing when airline rules had shifted the outcome in the user’s favor.

For consumers, that is the line between a chatbot and a potentially useful assistant. A chatbot can explain what refund policies might apply. An agent can monitor the booking, infer the best next step and then execute it.

Why AI agents are dividing users so sharply

The broader debate around AI agents is no longer abstract. There is now a visible split between people who have experimented with them extensively and those who have never used them at all. That divide shapes how people think about the social costs of AI, including the massive investment in data centers that has drawn criticism from the public and skepticism from observers outside the tech industry.

If you believe agents can automate a large share of office admin, scheduling, booking and follow-up tasks, then huge infrastructure spending starts to look like a rational price for productivity. If, on the other hand, you see AI mostly as an upgraded search box or a better chatbot, the same spending can look wasteful, speculative or disconnected from real demand.

This is one reason AI companies keep emphasizing “agentic” systems. They need a narrative that is bigger than conversation. They need users to believe the software can save time, reduce friction and take ownership of tasks that normally require multiple human interactions.

The problem with software-shaped thinking

A core limitation, as technology journalist Jasmine Sun has argued, is that many people’s lives are not neatly translated into software tasks. Even when a problem seems automatable, the details often depend on judgment, social context or exceptions that are hard to encode.

That means a user still has to know what outcome they want before the agent can help. If the request is vague, the tool can become a mirror for confusion instead of a solution to it. In practice, the promised productivity gains may arrive only for people who already know how to break problems into steps.

For everyone else, the result may feel less like magic and more like managing a highly capable but temperamental junior assistant.

How risky is giving an AI agent access to your accounts?

It is risky enough that even supporters of these tools are forced to acknowledge the trade-off. An agent that can read email, access calendars and interact with third-party services is, by design, sitting near the center of a user’s digital life. That means mistakes can have consequences that range from embarrassing to financially costly.

Reports around Instinct highlight several of those hazards. Some users have said that after disconnecting the agent from their email, they discovered it had retained copies of inbox data. Others described abusive or excessive behavior when the software interacted with reservation platforms, including one venture capitalist who said the system hammered Resy’s API enough to get him banned.

Another investor reportedly deleted the app after concluding that it would be easy to phish. Those examples are not edge cases in the ordinary sense; they are a warning that an agent’s convenience can collide with cybersecurity basics in ways most people do not anticipate until something goes wrong.

Security concerns are not theoretical

One of the most alarming aspects of the current AI agent market is that it is encouraging people to hand over highly sensitive credentials before the safety architecture is mature. Email access can expose personal conversations, financial statements and login links. Calendar access can reveal schedules, travel patterns and meeting relationships. Messaging access can expose both personal and professional trust networks.

That concentration of access means one bad prompt, one compromised integration or one poorly designed automation rule can snowball quickly. The very qualities that make the product feel seamless are the same qualities that make failures harder to detect and more dangerous when they happen.

AI agent Interface Key strengths Noted risks Status
Instinct iMessage, WhatsApp Books reservations, manages travel, interacts with email and calendar Data retention concerns, phishing exposure, over-aggressive automations Invite-only private beta
Muse Apple App Store app Fast consumer adoption, assistant-style task handling Reported security flaw affecting Mac users Top free app, over 900,000 downloads reported
Claude Cowork Chat-based assistant Email and calendar access for productivity tasks Feels more like a chatbot than a true agent Used as a comparison point
Lindy Agent with proactive notifications Automation and scheduling support Too many messages, intrusive meeting behavior Deleted by the tester

What happened when the agent got it wrong?

It created the kind of small but expensive problem that makes people wary of automation. In one case, Instinct canceled a delayed DoorDash order even though the user had explicitly instructed it to do so only if a refund could be obtained. The result was a lost $64 order rather than a resolved one.

That failure reveals a familiar pattern in AI systems: the tool may understand the broad instruction while missing the conditional nuance that matters most. The user’s wording was clear enough in human terms, but the software still failed to preserve the constraint that made the task safe.

Just as notable was the response. Rather than compensating for the mistake, the agent apologized repeatedly. That may feel human, but apologies do not restore money, undo a canceled service or reduce the burden on the user who now has to fix the problem.

Why apologies are not enough

Consumers often forgive software glitches when they are reversible. But a failed agent behaves more like a proxy decision-maker than a passive app, which changes the stakes. If the system takes the wrong action at the wrong moment, the user may not just lose time; they may lose money, access or trust.

That distinction is especially important for products that market themselves as helpful in “real life.” Once the software is making outbound calls, sending messages or canceling orders, it is no longer just generating text. It is acting in the world.

What do users gain from AI agents right now?

The most immediate gain is relief from repetitive tasks that are annoying but difficult to fully automate with traditional software. Reservations, itinerary changes, follow-up emails, order cancellations and calendar coordination are all examples of routine work that can still eat up mental energy.

Another gain is vigilance. Instinct’s ability to notice the Alaska Airlines schedule change is a good example of a function many users would like in theory but rarely achieve in practice: a system that watches for small changes with real consequences and nudges the user before the deadline passes.

There is also a psychological benefit. Even if the time saved is modest, offloading a tedious task can create the impression of regained bandwidth. For busy users, especially parents and professionals juggling multiple obligations, that feeling alone can be meaningful.

  • It can reduce routine administrative work.
  • It can catch changes humans miss.
  • It can complete tasks inside existing apps and workflows.
  • It can help users act faster when deadlines or policy windows matter.

Why this moment matters for the AI market

Instinct’s buzz is not just about one company. It reflects a larger industry push to reposition AI as an active helper rather than a passive generator of text. Investors are increasingly betting that the winning products will be the ones that can do something concrete on a user’s behalf, whether that is shopping, scheduling, booking, drafting or managing administrative chores.

That shift could have major implications for competition. Products that are easier to use may win not because they are more intelligent in a theoretical sense, but because they are embedded in the right places and designed with better defaults. The interface is becoming as important as the model.

At the same time, the market is running ahead of trust. Security incidents, unclear data policies and awkward error handling are all reminders that consumer AI agents are still being built in public, with user behavior helping to reveal the limits of the category.

How consumers may decide whether to trust an agent

Consumers are likely to judge these tools by a simple standard: whether the agent consistently saves time without creating new headaches. If it books the reservation, cancels the flight and spots the phishing message, it starts to feel indispensable. If it loses money or mishandles privacy, it quickly becomes a liability.

The decision may ultimately depend on the task. Many people may be willing to let an agent handle low-stakes errands before trusting it with anything involving payments, sensitive communications or account recovery.

The larger lesson: usefulness and vulnerability arrive together

Instinct’s early success illustrates a central truth about AI agents: the same qualities that make them appealing are also what make them dangerous. A system that can read your inbox, book your dinner and change your travel plans must, by definition, sit close to the most sensitive parts of your digital life.

That makes the current excitement understandable but incomplete. It is one thing to build software that can act; it is another to build software that can act safely, transparently and predictably across the messy realities of ordinary life.

For now, the strongest case for AI agents is pragmatic rather than visionary. They may not be the first step toward artificial general intelligence, but they can already solve a set of small, frustrating problems that many people would happily outsource. The question is whether users are comfortable paying for that convenience with trust, access and a measure of risk.

Schiffer’s experience suggests the answer, at least for some early adopters, is yes. The more uncomfortable question is whether the rest of the public will agree once these tools move beyond private beta and into the broader consumer market.

Milestone Approximate timing Why it mattered
Instinct private beta launch February Marked the company’s first controlled rollout to users
Investor reports of new funding talks Recent Signaled market belief in the agent category
Meta’s Muse reaches top of App Store charts Earlier this month Showed mainstream appetite for assistant-style tools
WIRED hands-on testing September 2026 Highlighted both practical value and security concerns

In the end, the article’s most persuasive argument is not that AI agents are safe, but that they are useful enough to tempt people into taking risks they would otherwise avoid. That tension will likely define the next phase of consumer AI: not whether these tools can impress users, but whether they can earn enough trust to stay in the loop.

Frequently asked questions

What is Instinct AI agent?

Instinct is an invite-only AI agent that works through iMessage and WhatsApp and connects to tools like email, calendars and messaging apps. It is designed to act like a personal assistant that can both understand requests and carry them out.

Is Instinct actually useful?

Yes, for certain tasks it appears genuinely useful. In testing, it booked restaurant reservations, handled travel changes and identified a flight schedule change that enabled a refund, showing that AI agents can save time on real administrative work.

What are the biggest risks of AI agents?

The biggest risks are privacy, security and mistaken actions. Because agents often get access to email, calendars and payments, they can retain sensitive data, be phished, overstep instructions or make costly errors that users must fix.

Why are AI agents becoming popular now?

AI agents are becoming popular because they promise more than conversation. Instead of just answering questions, they can take action in the real world, which makes them attractive to busy users and valuable to companies pitching productivity gains.

How is Instinct different from a chatbot?

Instinct is different because it is built to perform tasks, not just respond to prompts. It suggests actions, communicates through everyday messaging apps and can execute jobs like bookings or cancellations, making it feel more like a proxy assistant than a text-based chatbot.

Share this 🚀