The Next Generation Of AI Relies On High-Quality Data Creation
This voice experience is generated by AI. Learn more.This voice experience is generated by AI. Learn more.Akash Pugalia | Chief Digital Officer at TP.
gettyThe last generation of frontier AI models was trained on the internet: the open web, digitized books, code repositories and the vast universe of public text. The reality today is that major labs have effectively read it. That era is closing, and not because the world is running low on data. It’s because the kind of data that made the last leap possible has largely been used up, and the next leap requires something the internet doesn’t have.
This tension will redraw the entire AI data economy. Models need training material the open internet never had: deep, specialized, real-world records of how expert work gets done. This data is scarce, valuable and largely locked inside the operations of individual organizations.
But scarcity is only half the story and, if you ask me, the less interesting half. The richer question is where this data comes from, because unlike the internet, scraping will not be the solution. In some cases, this data does not even exist yet in any usable form.
Herein lies both the conflict and the opportunity. The companies building frontier AI need new, abundant sources of data to move forward. The organizations that hold that data, or can produce it to specification, are discovering they sit on one of the most valuable raw materials in the modern economy. Understanding what this new data actually is, why it commands such a premium and what separates a usable dataset from a liability is fast becoming essential knowledge for anyone building or buying AI.
The key behind the data needed today is depth and complexity. This particular kind of data has become so relevant and valuable because the job we’re asking AI to do has fundamentally changed.
Models are moving from answering questions to performing tasks; our first approach to genAI was drafting emails, searching for answers and handling basic surface-level chores. Now, agents can inspect records, call software tools, write code, check their own work, escalate exceptions and complete multi-step tasks.
This foundational change has pushed frontier labs, the companies building the most advanced AI systems in the world, toward increasingly specialized capabilities. The goal is no longer a model that answers questions well. It’s a model that can handle real, complex, uncertain scenarios: the messy, high-stakes situations where the right move isn’t obvious, and the work doesn’t fit in a chat box.
This transition requires training data that reflects real, complex working environments. That data takes several forms. The first is Reinforcement Learning environments, sometimes called RL gyms, which are controlled simulations of real software and systems in which a model can practice completing tasks and learn from the outcomes. The second is operational data: the authentic record of how an actual organization functions, across every department and workflow, over years of real activity.
The value of that second category has surfaced in a new trend: frontier labs and specialized data companies buying the internal data of defunct businesses. When a company shuts down, its operational history becomes an asset that can be sold. As one data-industry figure puts it, companies are sitting on decades of records that show how real work gets done, and that record is now among the most valuable material for training and evaluating AI.
Both categories described before share a quiet limitation. An RL environment simulates reality; operational data records reality. Neither is created from scratch to train a model. One is built to imitate the real world; the other is salvaged from it. That distinction matters in determining what data we need to train future models.
The “dead data” trend is revealing the supply constraint. Operational history is finite and available only under very specific conditions. That means no bigger version is coming. The same is true, in its own way, of simulations: an RL environment can only teach what its builders already knew to model. Both categories are ultimately bounded by reality as it exists or as it has already happened.
This surfaces the real frontier and why data creation is the heart of this story. The most valuable data for the next generation of models often doesn’t exist yet in any form; not on the internet, in an archive or in a simulation. It has to be deliberately produced: an expert performing the specific task, reasoning through the specific problem, generating the exact demonstration the model is missing.
Creating data is easy; creating data good enough to train a frontier model is not. It demands genuine domain expertise rather than generalists approximating it, rigorous verification so quality holds across millions of examples and a documented record of where every piece came from, increasingly a legal requirement, not just good practice.
Those are the questions any serious buyer should be asking of a data partner: Who are your experts? How is quality verified? Can you prove the origin of every piece? But the deeper point isn’t the checklist. It’s that data creation, done properly, is the one source of training data that doesn’t run dry, because it manufactures exactly what the frontier needs, precisely when the frontier needs it.
The internet that fueled the last decade of AI has already been used. What comes next will be built on data that must be deliberately made; deep, expert-driven and traceable to its source. This quietly changes who holds the power in AI. The race to accumulate the most data is over. Instead, we are now facing an era of data production that prioritizes complexity and depth over volume.
This means the advantage is shifting from those who have data to those who can make it well.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?
Comments (0)
No comments yet. Be the first to share your opinion!