The next leap in artificial intelligence may depend on some decidedly unglamorous work. According to TechCrunch, several AI labs are already paying a company called XDOF to collect the training data needed to build smarter robots.
The issue is what researchers call "physical AI" — systems meant to operate in the real world rather than just generate text on a screen. TechCrunch frames the challenge directly: if physical AI is going to match the accomplishments of large language models, there is a data problem that needs to be solved first.
That problem is data. The chatbots and text generators that have dominated headlines were trained largely on text scraped from the internet — an enormous, ready-made resource. Robots have no equivalent. Teaching a machine to move through and manipulate the physical world requires data drawn from physical actions, and TechCrunch describes the work of gathering it as dirty and unglamorous.
That is where XDOF comes in. Rather than each lab building its own pipeline for this labor-intensive collection, some are outsourcing it — paying XDOF to do the grunt work of assembling the real-world datasets that robotics models need.
Why it matters: the public tends to imagine AI breakthroughs as flashes of algorithmic genius, but this story is a reminder that progress often hinges on the painstaking, behind-the-scenes task of gathering the right data — and a market is already forming around who does that dirty work.