Robotic arm grasping an orange and white marker in a lab setting

XDOF races toward a $1.2 billion valuation as robot data demand surges

Robotics data startup XDOF is in talks for a $1.2B valuation as investors race to back the company building training data for robots.

In short

XDOF is reportedly in late-stage talks to raise a Series B at about a $1.2 billion valuation just months after coming out of stealth. The startup sells robotics data infrastructure and is benefiting from surging demand for real-world training data.

  • XDOF is in late-stage talks for a Series B at roughly a $1.2 billion valuation.
  • The startup emerged from stealth only about three months ago and raised a $70 million Series A in June.
  • Its business centers on collecting, labeling and organizing real-world robot training data.
  • Revenue is reportedly approaching a $50 million annualized run rate, helping fuel investor interest.
  • The company is positioning itself as an infrastructure layer for general-purpose robotics.

XDOF, the robotics-data startup that emerged from stealth only about three months ago, is in advanced talks to raise a Series B at a roughly $1.2 billion valuation. The round, which is being led by 8VC, underscores how quickly investors are bidding up companies that supply the training data needed to build general-purpose robots.

The Palo Alto-area startup has become one of the clearest beneficiaries of a new robotics bottleneck: while large language models were trained on massive swaths of the internet, physical robots still lack comparable real-world datasets. XDOF wants to solve that by building the collection systems, teleoperation pipelines and annotation tools that robotics companies and frontier AI labs need to train machines to work in the real world.

The company’s momentum is striking. XDOF raised a $70 million Series A in June from a high-profile group that included Thrive Capital, Andreessen Horowitz, Lux Capital and Spark Capital. According to people familiar with the business, it was not planning another fundraising process so soon. But the startup’s rapid commercial traction — with annualized revenue reportedly nearing $50 million — changed the picture and drew investor interest for an even larger follow-on round.

While the size of the new financing has not been disclosed and the final structure could still change, the discussion alone places XDOF among the most closely watched startups in the robotics infrastructure market.

Why investors are piling into robotics data

XDOF’s rise reflects a broader shift in AI investing: the most valuable companies are increasingly those that control scarce inputs, not just model architecture. In robotics, that scarce input is not compute or text data but high-quality physical interaction data — the kind needed to teach machines how to grasp objects, sort items, fold clothes, move boxes and perform other everyday tasks.

For years, robotics researchers have argued that progress has been slowed less by software brilliance than by the absence of a large-scale, standardized training set. Unlike chatbot systems, which could learn from billions of pages of text scraped from the internet, robots need examples of the physical world in motion. That includes arm positions, object manipulation, hand-eye coordination and sensor feedback from human demonstrations.

XDOF is trying to make that data pipeline repeatable and industrialized. Instead of relying on ad hoc lab recordings, the company is building a networked collection system designed to generate structured datasets at scale. Investors see that as a foundational layer for the next generation of robotics companies, much the way data-labeling platforms supported the rise of large language models.

Industry backers have increasingly described XDOF as a robotics analogue to Scale AI or Mercor, the data companies that helped power the generative AI boom by making training data easier to source and organize.

How XDOF got here

The company’s roots go back to academic research at UC Berkeley. Co-founder Philipp Wu, now the chief executive, was studying how robots learn from large datasets when he ran into the field’s core limitation: there simply was not enough large-scale data available to make the research move faster.

Wu joined forces with Fred Shentu, now the chief technology officer, and together they built a project called GELLO, a lower-cost teleoperation system that lets a person control a robotic arm from a distance. That work led to an influential robotics paper and ultimately became the technical and commercial basis for XDOF.

The startup’s approach is built around a simple idea: if the best robot training data is scarce, create the infrastructure to produce it efficiently. That means not only remote control of machines but also systems for capturing, labeling and organizing the resulting data so it can be used by model developers.

What does XDOF actually sell?

XDOF sells the infrastructure behind the data, not just the data itself. That includes collection tools, remote teleoperation systems, annotation workflows and the operational machinery needed to turn human demonstrations into usable robotics training sets.

The company is positioning itself as an outsourced data supply chain for physical AI. In practice, that means robotics labs and AI developers can focus on model training while XDOF handles the difficult work of gathering the motion, sensor and task data that those models require.

  • Teleoperation systems for controlling robots remotely
  • Data collection pipelines for physical tasks
  • Annotation and labeling workflows
  • Human operator training and deployment
  • Sensor-based capture for movement data

What is ABC, and why does it matter?

ABC is XDOF’s planned large-scale robot training dataset, and the company says it is building it with UC Berkeley’s AI Research lab. The startup says the collection could become the largest high-quality robotics training corpus assembled to date, though independent verification of that claim will depend on what is released and how it is benchmarked.

The project matters because the robotics sector has long lacked an equivalent to the vast datasets that transformed language AI. If ABC becomes a widely used benchmark or training resource, it could help standardize the way robots are taught across tasks and hardware platforms.

XDOF has said the dataset effort combines remote teleoperation with human collectors wearing sensors. The idea is to capture both robot-control demonstrations and embodied movement from people performing ordinary tasks. By mixing those streams, the company hopes to build richer data that can teach robots not just what to do, but how to do it in different physical settings.

How the data is collected

Data collection at XDOF is designed to mirror real-world movement as closely as possible. Remote teleoperators guide robotic arms from afar, while other workers wear body sensors that record motions such as folding laundry or flattening cardboard boxes.

Those tasks may sound mundane, but they matter enormously for robotics development. Everyday activities contain the subtle movement patterns, timing and environmental variability that robots must learn if they are to move beyond narrow industrial applications and into flexible general-purpose use.

Milestone What happened Why it matters
2024 XDOF was founded by UC Berkeley researchers Philipp Wu and Fred Shentu Established the company’s academic and technical base
June 2026 XDOF raised a $70 million Series A Signaled strong investor confidence in robotics data infrastructure
Three months after stealth exit Late-stage talks began for a Series B at about a $1.2 billion valuation Shows the company’s unusually fast valuation growth
Current phase Expansion of global data-collector and teleoperator teams Supports larger-scale robot training data production

Why the valuation is rising so fast

In venture capital, speed often follows proof of demand, and XDOF appears to have both. People familiar with the startup said annualized revenue is approaching $50 million, a figure that would be notable for a company only recently out of stealth and still scaling its operational footprint.

The company also reportedly works with 20 customers, including several frontier AI labs. That customer mix is important because it suggests demand is coming not only from robotics startups but from the largest and most ambitious AI builders looking for physical-world training data.

Venture capitalists have been particularly eager to back companies that sit at the intersection of AI infrastructure and scarce data assets. In the text-model era, data labeling and curation businesses became essential. The same pattern may now be repeating in robotics, where the business value lies in collecting hard-to-get examples faster and more consistently than others can.

The valuation also reflects a simple market dynamic: if investors believe robotics is entering its own scaling phase, then the companies that gather and package training data may become the picks-and-shovels winners.

Who else is competing in robot data?

XDOF is not alone in targeting this opportunity. A number of startups are trying to capture the same market, either by building robotics data systems directly or by broadening existing human-data businesses into physical AI.

Among the companies working in the same direction is Mecka AI, while larger data platforms such as Scale AI and Micro1 have also been extending beyond language-model work. The field is still early, but the strategic direction is clear: whoever can assemble the best real-world training data may gain an advantage in shaping how robots learn.

What makes physical data harder than text data?

Physical data is harder because it is slow, expensive and inconsistent to collect. Text can be gathered at internet scale with software; robot training data requires hardware, human labor, sensor systems, calibration and often repeated demonstrations of the same task in slightly different conditions.

That difficulty is exactly why XDOF is attracting attention. The startup is not trying to solve robotics with a single model breakthrough. Instead, it is trying to industrialize the production of the very material that future robotics models will need to improve.

The broader robotics race

Interest in robot data is rising at the same time that robotics itself is moving closer to mainstream commercial deployment. Warehouses, logistics firms, manufacturing lines and consumer-facing companies are all pushing for machines that can handle more varied tasks without exhaustive custom programming.

General-purpose robots, however, remain limited by their training. They still struggle with dexterity, object variation and unpredictable environments. That makes data quality and diversity critical. Every additional hour of well-labeled physical demonstration can help models learn behaviors that are difficult to encode by hand.

XDOF’s growth suggests that investors now see the data layer as a core part of the robotics stack, not a side service. If the company can continue scaling its collection network and maintain customer demand, it could become one of the defining infrastructure businesses in physical AI.

The startup’s challenge will be execution. Building a global workforce of teleoperators and sensor-based collectors is operationally complex, and the company will need to maintain quality as volume rises. It must also prove that the datasets it produces are not just large, but valuable enough to improve robot performance in measurable ways.

Timeline of XDOF’s rapid rise

XDOF’s path from research project to unicorn candidate has moved faster than most robotics startups. The following timeline captures the key public milestones known so far.

Date Event Significance
2024 Founded by Philipp Wu and Fred Shentu Converted academic research into a startup
June 2026 Closed a $70 million Series A Gained backing from major venture firms
Late summer 2026 Revenue reportedly approached $50 million annualized Showed strong early commercial traction
Sept. 2026 Entered late-stage Series B talks at about a $1.2 billion valuation Marked a dramatic valuation jump in a short period

What happens next?

The terms of the Series B are not yet final, and the total amount being raised has not been disclosed. That means the valuation, the round size and even the lead investor mix could still shift before any deal is completed.

Still, the fact that XDOF is already in late-stage discussions so soon after its Series A indicates strong market enthusiasm. If the financing closes near the reported valuation, the company will become one of the most prominent startups in the emerging robotics data category.

For the robotics industry, the deal would be another sign that the most valuable companies may be the ones building the plumbing behind the next wave of automation. Just as cloud infrastructure and data-labeling platforms became indispensable to software AI, robot training data may now be the bottleneck that determines who wins in physical AI.

In that sense, XDOF’s story is about more than a fast-growing startup. It is a snapshot of where investor belief is shifting: away from purely model-centric AI bets and toward the systems that make intelligent machines possible in the real world.

If the company can turn its academic roots, customer traction and data infrastructure into a durable platform, its current valuation may end up looking less like an outlier and more like an early signal of a larger robotics breakout.

Frequently asked questions

What is XDOF?

XDOF is a robotics-data startup that builds teleoperation, collection and annotation systems for training general-purpose robots. It was founded by UC Berkeley researchers Philipp Wu and Fred Shentu and is aiming to become the infrastructure layer behind physical AI.

Why are investors valuing XDOF so highly?

Investors are bidding up XDOF because robotics needs scarce, high-quality real-world data, and the company is reportedly showing strong early revenue traction. Its annualized revenue is said to be approaching $50 million, which is unusually fast growth for such a young startup.

How does XDOF collect robot training data?

XDOF collects data through remote teleoperation and sensor-based human demonstrations. Operators steer robotic arms from a distance, while other collectors wear body sensors to record everyday tasks such as folding clothes and flattening boxes, creating structured physical-world training data.

Who is leading the Series B talks?

The Series B talks are being led by 8VC, according to people familiar with the deal. The terms are not final, and the total amount being raised has not yet been disclosed, so the structure could still change.

What companies are competing with XDOF?

XDOF faces competition from startups such as Mecka AI and from broader data platforms like Scale AI and Micro1 that are expanding beyond language-model data. The market is still early, but competition is growing as more firms chase robotics training data.

Share this 🚀