XDOF, a startup building the data pipeline for physical artificial intelligence, has raised $70 million from Thrive Capital, Spark Capital, a16z, Lux, and WndrCo to scale its global network of teleoperators and egocentric data collectors. The company, which currently employs about 60 people and serves 20 customers including frontier AI labs, plans to hire and train operators worldwide to capture the real-world data that robot models require for training. CEO Philipp Wu argues that physical AI represents the next frontier, and XDOF is positioning itself as the essential infrastructure layer for that transition. The funding round signals that major investors see a lucrative opportunity in the unglamorous but critical work of gathering the high-quality, labeled data that powers autonomous systems. As AI labs race to build robots that can navigate physical spaces, manipulate objects, and perform complex tasks, the bottleneck is shifting from compute to data. XDOF is betting that outsourcing this dirty work will become a permanent feature of the AI supply chain. This matters now because the physical AI market is projected to grow explosively over the next decade, and the companies that control the data pipeline will wield significant leverage over which models succeed.
Building a global teleoperator workforce with the $70M raise

XDOF's core business involves hiring and training teleoperators and egocentric data operators who collect the raw material for robot training. Teleoperators remotely control robots to perform tasks, generating demonstration data that AI models learn from. Egocentric data operators wear cameras and sensors to capture first-person perspectives of human activities, providing the visual and motion data that teaches robots how to interact with the world. The $70 million raise will fund the expansion of this workforce across multiple geographies, allowing XDOF to scale its data collection capabilities to match the growing demands of its customer base. The company's 20 customers include frontier AI labs that are developing general-purpose robot models, as well as specialized robotics companies building systems for logistics, manufacturing, and healthcare. By centralizing data collection, XDOF offers these labs a way to avoid the operational complexity of managing their own teleoperation teams. The model mirrors the shift that occurred in software AI, where companies like Scale AI and Labelbox emerged to handle data labeling for computer vision and natural language models. XDOF is applying the same logic to physical AI, betting that the dirty work of data collection will be outsourced to specialists rather than kept in-house. The company plans to hire operators in regions with strong technical talent pools, including parts of Southeast Asia and Eastern Europe, where labor costs are lower and the availability of English-speaking workers is high. Each new hire undergoes a two-week training program that covers safety protocols, hardware operation, and data quality standards before being assigned to customer projects. The company currently employs about 60 people, and the new funding will allow it to expand its workforce to several hundred operators within the next 18 months, according to a person familiar with the company's plans.
Data collection as a service: how the money flows through the P&L

XDOF's revenue model charges customers per hour of teleoperation or per labeled data point, creating a variable cost structure that aligns with the labs' development cycles. For a frontier AI lab training a new robot model, the cost of data collection can run into the millions of dollars per training run, depending on the complexity of the tasks and the number of demonstrations required. XDOF's pricing reflects the labor-intensive nature of the work: teleoperators must be trained, supervised, and equipped with the necessary hardware, and the data must be validated for quality before it enters the training pipeline. The company's gross margins depend on its ability to optimize the efficiency of its workforce, using software tools to manage scheduling, quality control, and data annotation. The $70 million raise provides a runway of several years, assuming the company maintains its current burn rate of roughly $15 million to $20 million per year. As customer demand grows, XDOF can scale its workforce without proportional increases in overhead, improving unit economics over time. The company's investors are betting that the physical AI market will follow the same trajectory as software AI, where data infrastructure companies achieved significant scale and valuation multiples. XDOF's internal software platform automates parts of the quality assurance process, flagging low-quality demonstrations and routing them for human review, which reduces the cost per validated data point as the workforce expands. The company has already signed multi-year contracts with two of the largest frontier labs, locking in recurring revenue and giving it a first-mover advantage in negotiating exclusive data collection arrangements.
Competitive reshuffle: who gains and who loses in the data pipeline
XDOF's emergence creates a new competitive dynamic in the AI infrastructure stack. Established data labeling companies like Scale AI and Labelbox have focused primarily on software AI, including image classification, natural language processing, and autonomous vehicle data. XDOF targets the physical AI segment, which requires different data types: teleoperation demonstrations, egocentric video, tactile feedback, and multi-modal sensor streams. This specialization gives XDOF an advantage in serving frontier AI labs that are building general-purpose robot models, but it also means the company must invest heavily in domain expertise and hardware infrastructure. The competitive threat to incumbents is real: as physical AI becomes more important, the data labeling giants will need to either acquire or build capabilities in this area. For the frontier AI labs themselves, XDOF offers an alternative to building internal data collection teams, which would require significant capital expenditure and operational complexity. The labs can instead focus their resources on model architecture and training algorithms, leaving the data pipeline to specialists. This division of labor mirrors the broader trend in AI, where companies are increasingly outsourcing infrastructure to specialized providers. XDOF has already signed multi-year contracts with two of the largest frontier labs, locking in recurring revenue and giving it a first-mover advantage in negotiating exclusive data collection arrangements.
Downstream effects: hyperscalers, hardware, and enterprise buyers
The rise of XDOF and similar data pipeline companies has second-order effects across the AI ecosystem. Hyperscalers like Google and Amazon Web Services benefit from increased demand for compute and storage, as the data collected by teleoperators must be processed, stored, and served to training clusters. The hardware supply chain also feels the impact: companies like NVIDIA and AMD see increased demand for GPUs and specialized AI accelerators, while robotics hardware manufacturers benefit from the deployment of teleoperation systems. For enterprise buyers, the availability of high-quality training data from providers like XDOF accelerates the development of physical AI applications in logistics, manufacturing, and healthcare. Companies that might have waited years for robot models to mature can now access pre-trained systems that handle specific tasks. The regulatory environment also plays a role: as governments increase their focus on AI regulation and sovereign AI capabilities, the data pipeline becomes a strategic asset. Countries that want to develop domestic physical AI capabilities will need access to data collection infrastructure, creating opportunities for companies like XDOF to expand internationally. XDOF is already in discussions with government agencies in Japan and Germany about pilot programs for industrial robotics data collection, tapping into sovereign AI initiatives that prioritize domestic control over training data.
Policy and strategy signal: the data pipeline as a strategic asset
XDOF's funding round and business model signal a fundamental shift in how the AI industry thinks about data. For years, the dominant narrative focused on compute as the primary bottleneck, with companies like NVIDIA reaping the benefits of GPU scarcity. But as frontier models approach the limits of available training data, the quality and diversity of data become the critical differentiator. Physical AI requires data that is expensive and difficult to collect, creating a natural monopoly for companies that can build and maintain the infrastructure. The involvement of Thrive Capital and a16z, two of the most prominent venture firms in AI, validates this thesis. Their investment suggests that the data pipeline will be a major source of value creation in the next phase of AI development. The strategic implications extend beyond individual companies: as AI labs race to build physical AI systems, control over the data pipeline becomes a form of leverage. Governments and corporations that want to ensure access to high-quality training data will need to invest in this infrastructure, either through partnerships or direct ownership. The G7's recent discussions about trusted partners and sovereign AI capabilities underscore the geopolitical dimension of this trend. XDOF has already registered subsidiaries in the UK and Singapore to comply with local data residency requirements, positioning itself as a trusted partner for governments that want to keep sensitive training data within their borders.
The physical AI market is still in its early stages, but the infrastructure being built today will shape the competitive landscape for years to come. XDOF's $70 million raise provides the capital to scale its workforce, develop its software platform, and expand its customer base. The company faces challenges: recruiting and training teleoperators at scale is difficult, quality control is expensive, and the technology is evolving rapidly. But the demand for high-quality training data is only going to increase as frontier AI labs push the boundaries of what robots can do. If XDOF can execute on its vision, it will become an essential part of the AI supply chain, generating recurring revenue from the world's most valuable technology companies. The company's success will depend on its ability to maintain quality, manage costs, and stay ahead of competitors. For investors and strategists watching the AI industry, XDOF represents a bet on the thesis that data infrastructure will be as important as compute infrastructure in the physical AI era.
The BossBlog Daily
One email with the AI markets brief — the 13F moves, the Congressional trades, and what changed. No fixed schedule and no filler: it goes out when there is something worth sending.
Unsubscribe any time. We never sell or share the list.