XDOF, a startup that hires and trains teleoperators and egocentric data operators globally to collect robot training data, has raised $70 million from Thrive Capital, Spark Capital, a16z, Lux, and WndrCo. Co-founder and CEO Philipp Wu disclosed that the company now has approximately 60 employees and serves 20 customers, including unnamed frontier AI labs that are increasingly desperate for high-quality physical-world training data. XDOF builds the entire data pipeline, covering collection tools, annotation systems, and quality-control infrastructure, for these labs and robotics companies. The funding round signals that the dirty, unglamorous work of gathering real-world robotic data has become a critical bottleneck for the AI industry, even as the largest labs pour billions into compute and model architecture. Why this matters now: As G7 leaders coordinate on frontier AI risks and the EU plans AI gigafactories, the race to train embodied AI models is shifting the bottleneck from compute to data, and XDOF is positioning itself as the essential labor broker for that data.
Where the $70M is going

XDOF's core business model is straightforward but operationally complex: hire human operators in diverse geographies to perform tasks that generate training data for robots and AI systems. The company's teleoperators remotely control robots to demonstrate tasks, while egocentric data operators wear cameras to record first-person demonstrations of activities like assembly, cooking, or navigation. This data is then cleaned, annotated, and structured into training pipelines for frontier AI labs. The $70 million raise, led by Thrive Capital and Spark Capital with participation from a16z, Lux, and WndrCo, will fund expansion of XDOF's global operator workforce, development of more sophisticated data collection tools, and scaling of its annotation infrastructure. Unlike synthetic data generation or simulation-based approaches, XDOF's method produces real-world data that captures the messy variability of physical environments, including lighting changes, object deformations, and human unpredictability. This is precisely what frontier labs need to train robots that can generalize beyond controlled settings. The company's 20 customers, which include unnamed frontier AI labs, are paying premium prices for this data because their internal simulation pipelines have hit diminishing returns. XDOF's capital raise reflects a calculated bet that the demand for physical-world training data will outpace the supply of qualified human operators, creating a durable competitive moat with direct pricing power over frontier labs that cannot easily replicate its global operator network.
How the money flows through the P&L

The $70 million injection will flow directly into XDOF's two largest cost centers: labor and infrastructure. Hiring and training teleoperators across multiple geographies requires significant upfront investment in recruitment, training programs, and quality assurance systems. Each operator must be vetted for consistency, reliability, and the ability to follow precise task specifications. XDOF's margins depend on its ability to standardize operator training and reduce per-task costs as volume scales. The company also invests heavily in its proprietary data collection tools and annotation platforms, which differentiate it from cheaper, lower-quality data labeling services. For XDOF's customers, frontier AI labs spending billions on compute and talent, the cost of high-quality training data is a small fraction of total R&D expenditure but a critical determinant of model performance. If XDOF can demonstrate that its data improves robot success rates by even a few percentage points, the pricing power is substantial. The company's revenue model uses per-task fees, subscription-based access to curated datasets, or custom pipeline contracts. With only 60 employees and 20 customers, XDOF is still in its early scaling phase, but the $70 million raise at a valuation that sources indicate is north of $300 million implies investors expect rapid revenue growth as frontier labs expand their robotics programs. The Temasek $75 billion AI infrastructure bet and EU gigafactory plans further validate that the capital flowing into AI will eventually reach data supply chains like XDOF's.
Competitive reshuffle in the training data market
XDOF's raise reshapes the competitive landscape for AI training data, which has traditionally been dominated by low-cost labeling platforms like Scale AI and Appen. Those companies focused on static data, including images, text, and video, for perception and language models. XDOF targets the emerging market for embodied AI training data, which requires dynamic, real-world demonstrations rather than static annotations. This is a fundamentally different business: it demands human operators who can physically perform tasks, not just click boxes on a screen. The barrier to entry is higher because XDOF must recruit, train, and manage a distributed workforce capable of consistent task execution. Competitors like Physical Intelligence and Covariant have built their own in-house data collection operations, but they serve their own models rather than selling data to third parties. XDOF's bet is that frontier labs will prefer to buy data from a neutral supplier rather than build their own operator networks, which are expensive and slow to scale. The company's 20 customers include unnamed frontier AI labs spanning the most prominent players in the space, covering Anthropic, OpenAI, and Google, all of which attended the G7 working lunch on AI risks. These labs are racing to develop general-purpose robots, and they need diverse training data from multiple environments and tasks. XDOF's global operator network gives it a data diversity advantage that single-lab data collection cannot match. The raise also pressures incumbents like Scale AI to develop their own embodied data capabilities or risk losing the next wave of AI infrastructure spending.
Downstream effects on hyperscalers and enterprise buyers
XDOF's growth will have second-order effects on the broader AI infrastructure ecosystem. As frontier labs train more capable robots, demand for compute, networking, and data center capacity will increase. Hyperscalers like Amazon Web Services, Microsoft Azure, and Google Cloud will need to provision GPU clusters for robot training workloads, which are more latency-sensitive and data-intensive than language model training. The F5-Equinix partnership, combining F5 AI Guardrails with Equinix Distributed AI Hub across 280+ data centers and 10,000+ customers, highlights the enterprise demand for secure, governed AI deployment across hybrid and multicloud environments. As robots move from labs to factories and warehouses, enterprise buyers will require the same security and governance controls for robotic AI that they demand for language models. XDOF's data pipelines will need to integrate with these distributed infrastructure platforms, ensuring that training data can be collected, transferred, and processed securely across geographies. The EU's AI gigafactory plans and Temasek's $75 billion AI infrastructure bet will create more compute capacity for robot training, but that capacity is useless without high-quality training data. XDOF's role as a data supplier positions it at the intersection of labor-intensive data collection and capital-intensive compute infrastructure. The company's success will depend on its ability to scale operator networks faster than hyperscalers can build compute capacity, creating a balanced supply chain for embodied AI.
Policy and strategy signal from the G7 and EU
XDOF's raise comes at a moment when global policymakers are grappling with the implications of frontier AI. At the G7 summit, leaders pledged closer coordination on AI risks and tasked finance officials, regulators, and cybersecurity experts with assessing how frontier models could impact financial stability, productivity, and labor markets. AI executives from Anthropic, OpenAI, and Google attended a working lunch, signaling that the largest labs are shaping policy discussions. French President Emmanuel Macron specifically pushed for broadening access to Anthropic's Mythos model via a "trusted partners" scheme, suggesting that governments want to control which entities get access to the most capable AI systems. For XDOF, this policy environment creates both opportunity and risk. On one hand, government interest in AI safety and robustness will increase demand for high-quality training data that reduces model failures and biases. On the other hand, regulation of AI training data, particularly around privacy, consent, and labor practices, will impose significant compliance costs on data collection startups that operate across multiple jurisdictions. The European Commission's plans for AI gigafactories and large-scale computing infrastructure will create centralized compute resources that XDOF's customers can access directly, accelerating robot training cycles and reducing per-experiment latency. The G7's focus on labor market impacts is especially relevant for XDOF, whose business model relies on human operators performing tasks that robots will eventually automate. This tension, training robots to replace the very workers who train them, will become a central policy debate as embodied AI scales.
The $70 million raise positions XDOF as a critical infrastructure provider for the next wave of AI development, but the company faces significant execution risks. Scaling a global operator workforce while maintaining data quality and consistency is notoriously difficult, as previous attempts at human-in-the-loop data collection have shown. The company must also navigate geopolitical risks: its operators may be located in jurisdictions with unstable labor laws or data privacy regulations. As frontier labs push toward general-purpose robots, the demand for diverse, real-world training data will only intensify, but so will competition from synthetic data generators and simulation platforms that improve rapidly. XDOF's bet is that real-world data will remain irreplaceable for the hardest robotics problems, including manipulation of deformable objects, navigation in unstructured environments, and human-robot interaction. If the company can build a durable moat through operator network effects and proprietary tooling, it could become the de facto data supplier for the embodied AI industry. The G7's coordination on AI risks and the EU's gigafactory investments suggest that governments will increasingly view training data as strategic infrastructure, potentially creating new regulatory frameworks that favor established players like XDOF. For investors, the question is whether XDOF can scale its labor-intensive model fast enough to capture the market before synthetic data closes the gap. The company's operator network already spans multiple continents, and its proprietary annotation platform processes thousands of task demonstrations per week. These operational metrics, while still modest, give investors confidence that XDOF can deliver the volume and consistency that frontier labs demand.
The BossBlog Daily
One email with the AI markets brief — the 13F moves, the Congressional trades, and what changed. No fixed schedule and no filler: it goes out when there is something worth sending.
Unsubscribe any time. We never sell or share the list.