A person with braids and glasses stands against a plain background, wearing a dress with newspaper print. Synthesia logo i...

Inside Synthesia’s digital twin push: a journalist tests the future of AI avatars

Synthesia created a digital avatar of a journalist, spotlighting how AI twins could reshape PR, training and journalism.

In short

Synthesia built a digital avatar of a journalist to demonstrate how far AI twins have advanced. The experiment highlights new enterprise uses for avatars while raising questions about trust, authenticity and the future of media work.

  • Synthesia created an interactive digital avatar of a journalist trained on one reported story.
  • The company now offers scripted avatar videos, interactive Sessions and an API platform.
  • Its avatars combine voice-to-text, language, text-to-voice and video models.
  • The demo showed both the practical appeal and the eerie feel of digital twins.
  • The experiment raised broader questions about AI in journalism, PR and corporate communications.

Synthesia has moved beyond generic talking-head videos and into personalized AI twins: the company recently created an interactive avatar of a journalist, showing how far synthetic likenesses have advanced and how quickly they could reshape media, PR and corporate communications. The experiment matters because it highlights both the convenience and the unease of giving a machine your face, voice and conversational style.

The avatar, built during a visit to Synthesia’s new New York office in September, was trained on one of the journalist’s reported stories and could answer questions only within that narrow topic. That made it a controlled demonstration rather than a fully open-ended chatbot, but it still offered a vivid preview of what “digital doubles” may look and sound like as the market matures.

What Synthesia is showing off now

Synthesia is no longer just selling a way to turn scripts into polished avatar videos. The London-founded company is now pushing interactive products that can listen, respond and even evaluate people in real time, positioning itself at the intersection of enterprise training, AI video and conversational systems.

The company says its platform now spans three distinct offerings: standard avatar video generation, interactive Sessions products for roleplay and surveys, and an API layer that lets customers build their own avatar-based tools using Synthesia models alongside other vendors’ technology.

From static presenters to responsive digital twins

Earlier Synthesia use cases were largely one-way. A customer would type a script, choose an avatar and publish a video. The newer products are more dynamic: users can ask questions, practice sales calls or complete assessments while an avatar speaks back and reacts.

That shift matters because it takes Synthesia out of the “presentation software” bucket and moves it toward agentic software and interactive communications. In other words, the product is no longer just about showing a face on screen. It is increasingly about simulating a conversation.

Item What it is Why it matters
Synthesia classic avatars Scripted videos where an avatar reads supplied text Still the company’s core content-creation product
Sessions Interactive AI avatars for roleplay, training and surveys Moves Synthesia into two-way enterprise engagement
API platform Tools for building custom avatar applications with third-party models Lets customers assemble their own stack and workflows
Personal digital twin An avatar built from photos and voice samples of a real person Shows how close synthetic identity is getting to mainstream use

How the digital twin was made

To create the avatar, Synthesia brought the journalist into a small studio setup inside its office and collected photos plus a short voice recording. The process took only a couple of days, according to the account, and required consent before the likeness could be used.

The result was not just one avatar but several. Synthesia produced simple reading avatars and interactive versions that could answer questions. Some were rendered with glasses, others without, underscoring how quickly a person’s digital presence can be repackaged into alternate styles.

What powers the system?

The company’s stack combines multiple AI systems rather than relying on a single model. Speech is converted to text, a language model interprets the request and decides what to say, text is turned back into audio, and Synthesia’s own video model animates the face and mouth movements.

Synthesia also allows customers to swap in outside vendors for some layers. That includes voice and language tooling from other companies such as Cartesia, ElevenLabs, Google and OpenAI. Enterprises can host the avatars themselves or have Synthesia manage the infrastructure for them.

This modular approach is important because it suggests the market is evolving from a single-product video company into a platform business that can plug into broader enterprise workflows.

According to Synthesia’s demonstration, the avatar was intentionally limited to one article and would redirect nearly every query back to that story, rather than improvising beyond its training.

Why the experiment felt both useful and unsettling

The journalist behind the avatar described an immediate sense of novelty, followed by something closer to discomfort. The likeness was good enough to be amusing to friends and family, but also strange enough to raise the larger question of where these systems go next.

That reaction is likely to be familiar to anyone watching the rise of AI-generated faces and voices. At first, the technology feels like a clever demo. Then the implications begin to sink in: if a person can be copied for public explanation, customer service, training or sales, what stops companies from cloning workers more broadly?

In the specific case of journalism, the test raised a sharper concern. A newsroom depends on trust, reporting and the human relationships that make audiences believe the work. An avatar may help distribute information, but it is harder to imagine it replacing the credibility built by a reporter’s judgment, sourcing and accountability.

What friends and family noticed

People who saw the avatars reportedly reacted differently depending on which version they encountered. The simple reading avatar sounded more convincing than the interactive one, while the interactive version was considered close enough to trigger unease despite limitations in voice quality and realism.

The journalist’s mother, meanwhile, responded with enthusiasm and tried to challenge the model with personal questions only family members would know. The avatar did not answer those prompts, instead steering back to the article it had been trained on. That constraint made the system safer, but also underscored how scripted the experience remained.

How far has Synthesia come as a company?

Synthesia has become one of the best-known names in AI avatar software, joining competitors such as D-ID, HeyGen and Colossyan in a crowded market for synthetic presenters. The company says it crossed $100 million in annual recurring revenue last year and reached a $4 billion valuation earlier this year, signaling that enterprise demand for polished, scalable video tools remains strong.

That growth matters because it shows the avatar category has moved well past novelty. Enterprises are not merely experimenting; they are paying for systems that can reduce production time, localize content, automate internal training and standardize communication across teams.

What companies are buying

  • Internal training videos created without on-camera production crews.
  • Sales roleplay tools that simulate customer objections.
  • Survey and assessment systems using conversational avatars.
  • Custom video workflows built through API integrations.

The pitch is straightforward: if a company needs a presenter in 15 languages, on short notice and at scale, an AI avatar can be cheaper and faster than booking a studio, hiring talent and re-editing dozens of variants.

What makes digital avatars different from chatbots?

Digital avatars differ from chatbots because they add a face, a voice and a performance layer to the conversation. A text model can answer a question. An avatar can appear to answer it in a human-like way, which makes the exchange feel more personal, persuasive and memorable.

That realism is exactly why companies see commercial value in the format. It is also why critics worry about manipulation, identity misuse and the emotional pull of synthetic humans. When a machine looks like a colleague, executive or reporter, the line between tool and impersonation gets thinner.

Deterministic versus open-ended systems

One of the most important distinctions in the demonstration was that the journalist’s interactive avatar was deterministic. It did not invent new topics or roam freely across the internet. It stayed tethered to the story it was built around and consistently returned to that source material.

That design choice reduces the risk of hallucinations and misrepresentation, but it also reveals a product strategy. Synthesia appears to be emphasizing controlled, enterprise-safe interactions over open-ended companionship. For corporate buyers, that restraint may be essential.

Why journalists and creators may care most

For journalists, creators and public-facing professionals, the rise of digital twins changes the economics of being visible online. A clone can answer routine questions, recap published work and handle repetitive outreach without requiring the original person to be present every time.

That could save time. It could also create a deeper brand identity around an individual’s face and voice. But it would bring new questions about consent, context and ownership. If a reporter’s likeness can be used to explain a story, who controls the script, and who is responsible when the clone gets it wrong?

There is also the issue of audience expectations. Viewers may accept avatars for product demos or training modules, but when it comes to news, they may still prefer a live human being who can be questioned, challenged and held accountable.

Could avatars replace on-camera talent?

They could replace some routine appearances, but full replacement appears unlikely in the near term. The clearest near-term use cases are repetitive, templated and operational: onboarding, compliance, FAQ responses, sales rehearsal and multilingual announcements.

For high-trust roles, especially in journalism and executive communication, the stronger case is augmentation. An avatar can extend reach without fully replacing the person behind it.

How the market may evolve next

The next stage of avatar adoption will likely depend on whether companies can make the experience both believable and safe. That means better lip-sync, more natural voices, stronger guardrails and clearer disclosure when viewers are interacting with a synthetic likeness.

It also means more choices around infrastructure. By letting customers combine its video models with third-party speech and language systems and host them in different cloud environments, Synthesia is trying to meet enterprise demands for flexibility and control.

The broader challenge is social, not technical. As avatars become easier to produce, businesses will have to decide where authenticity still matters more than scale. In some settings, a synthetic spokesperson may be efficient and useful. In others, a real person may remain the only acceptable option.

Timeline Milestone Why it matters
Last year Synthesia said it passed $100 million in ARR Signals enterprise traction
Earlier this year Company hit a $4 billion valuation Shows investor confidence in AI video tools
Summer Interactive avatar of Synthesia’s corporate affairs head was shown to the journalist Demonstrated internal use of digital twins in PR
September Journalist’s own avatar was created during a New York office visit Marked the company’s first journalist digital twin demo

What this says about AI in PR, media and work

AI avatars are quickly becoming more than a novelty in corporate communications. The technology now supports public relations outreach, employee training, roleplay, customer education and internal support. The next step is likely to be personalized digital agents that look and sound like real employees.

That is where the promise and the tension collide. A digital twin can work around the clock, repeat approved messaging and answer predictable questions. But the more lifelike it becomes, the more it raises concerns about trust, disclosure and the emotional boundary between a real person and a modeled version of that person.

The underlying question raised by the demo is not whether avatars can speak, but whether audiences will accept them as substitutes for the people they imitate.

For now, Synthesia’s demo suggests that the market is still in the early stage of figuring that out. The tools are already good enough to feel strange, helpful and commercially viable all at once. What remains uncertain is how much of daily professional life people are willing to hand over to a digital version of themselves.

Frequently asked questions

What did Synthesia do with the journalist’s likeness?

Synthesia created a digital avatar based on the journalist’s photos and voice sample, then trained it on one article so it could answer only questions about that story. The company used the demo to show how interactive AI twins can work in practice.

How does Synthesia’s interactive avatar technology work?

It works by combining voice-to-text, a language model, text-to-voice and a video animation model. The system listens, converts speech into text, generates a response, speaks it aloud and animates the avatar’s face and mouth in real time.

Why does this matter for journalism and PR?

It matters because avatars could help reporters and companies answer routine questions, scale outreach and save time, but they also raise concerns about trust, disclosure and authenticity. In journalism especially, the human relationship and accountability still matter a great deal.

Is Synthesia only making talking-head videos?

No. Synthesia now sells several products, including standard scripted avatar videos, interactive Sessions for roleplay and surveys, and an API platform for building custom avatar applications. That suggests the company is expanding into broader enterprise communication tools.

Share this 🚀