Business·7 min read

#2 – How Does an AI Training Data Company Make Money?

MYBy MY

Introduction

Just as Warren Buffett is known for favoring 'tollgate' businesses — those you must pay to pass through — whenever a new industry is growing enormously, I find myself wondering what that industry's tollgate business might be. For a long time, when thinking about what business could be done in the AI industry, I thought that securing AI training data would fit Buffett's concept of a 'tollgate business.' Through this post, I want to examine whether the AI training data business is actually a tollgate business capable of generating sustainable profit, and understand how its revenue model works.

🚀 AI Training Data Companies in Korea

To quickly grasp how much attention and scale a market commands, how it has evolved, and the differentiation strategies each player has adopted to survive, looking at the key players is the place to start. I've summarized the top 7 AI training data companies in Korea by investment or revenue scale in the two tables below.

Summary of 7 companies: founding year, service description, investment

Summary of 7 companies: business results and characteristics

From comparing these seven companies, I identified three characteristics common to the players in this industry.

  1. Founded between 2015 and 2017: From the early 2010s, keywords around artificial intelligence — humanoid robots, autonomous driving — began entering popular consciousness, and the March 2016 Lee Sedol vs. AlphaGo match pushed interest in AI to its peak. It appears that entrepreneurs who quickly identified this trend as a business opportunity started building training data businesses at around the same time. As always, even people living in the same era, some have the creative eye to spot opportunities within it.
  2. Tug-of-war between automation and quality: In the major tasks of data construction — data collection, labeling, quality review — human intelligence still accounts for a large share. Companies are each hiring strong AI engineers to develop AI that can automate data labeling and annotation tasks with high accuracy. This capability creates a virtuous cycle of more efficiently reducing labor costs while raising data quality and better satisfying data buyers. Currently, automation and manual work are still intermixed, so companies employ both manual workers and automation R&D staff simultaneously — which means relatively high payroll as a share of costs.
  3. Relatively small investment scale: Compared to the investment rounds of the generative AI startups and larger AI companies drawing enormous attention recently, the capital raised by training data companies with 5 to 8 years of history is modest. Given that I had framed these as 'gateway' businesses at the outset, the relatively small investment and revenue scale was quite striking. My hypothesis for why: the dominant business model is B2G (government-led services), and since the ownership of the final data produced doesn't belong to the data construction companies — making their revenue one-off rather than recurring — this has held down expectations.

💸 Demand Driving the Market

The AI training data construction business is a foundational service for the AI industry and a business type needed in the early stages of market formation — which makes investment from market-forming entities like government necessary. For now, government-led demand is the most visible, while private and corporate data demand is likely growing quietly below the surface. Current AI training data demand can be broken down as follows:

  1. Government-led B2G contracts: The largest government program is the Ministry of Science and ICT's AI Training Data Construction Support Program. In 2023, a total budget of 218.8 billion KRW was allocated; 538.2 billion KRW was allocated in 2022, and 292.5 billion KRW in 2021. The largest budget items in the 2023 designated solicitation included: joint/arthritis data (5.1 billion KRW), fairy tale data (4.5 billion KRW), and live streaming video translation data (4.2 billion KRW). In particular, medical and content categories took up a large and varied number of slots — which tells you which AI sectors the government is prioritizing. Beyond this, local governments (Seoul, Daejeon, etc.) and public research institutions (NIPA, ETRI) also post data construction projects.
  2. Government R&D contracts: Looking through projects listed on the National R&D Portal, there are many contracts in the AI data space for consulting, design research, and guideline development. These tend to be small in scale (around 100 million KRW each) and, while numerous, require narrow and deep specialization across diverse fields — making them a demand segment that's hard for companies to concentrate on at scale.
  3. Corporate B2B contracts: Web research didn't surface clear data on corporate B2B contracts. Looking at the footnotes in audit reports that are publicly disclosed for a few of the seven companies, it appears that specialized data contracts from large corporations can, in sizable cases, generate 2–3 billion KRW in revenue from a single contract annually. Personally, I think specialized corporate data demand (medical, logistics, autonomous driving, etc.) is likely to become the largest share of total revenue going forward.

🔮 Three Strategies for the Future

Running through all this, I found myself doing an interesting thought experiment: if I were the CEO of an AI training data company right now, what strategies would I need to pursue to survive and grow? I organized my thinking into three core directions:

(1) Shift from services to data sales

As the saying goes, data is gold — and good data holds value as intellectual property, which makes the question of who owns the constructed data a critical one. It would be hard to capture ownership of all data from the start. But steadily reinvesting the cash generated from simple service contracts back into the company — proactively building data in domains where demand is expected to grow — and then increasing the share of revenue from selling that data to multiple buyers would become a meaningful competitive advantage.

(2) Build data in specialized domains with a moat

Looking at current government project announcements and major revenue sources for these companies, autonomous driving and healthcare appear consistently. Some industries are growing their markets through AI adoption; in others, the cost of AI integration still outweighs the benefits. From the perspective of a data company, the strategy is to target sectors where AI utilization is likely to surge, in order, and build specialized data stacks in those domains. Sectors that will require high-purity, high-quality data in the near term include autonomous vehicles, industrial robotics, smart cities, logistics, and defense. Winning a strong reputation by focusing intensively on key domains could create a business with deep moats.

(3) Services for the 'next' stage after data construction

For now, 'quantitative' expansion is the urgent issue in this space, and at least through 2025, scaling data volume seems likely to be the main focus. But alongside quantitative expansion, the next stage to prepare for is upgrading to data 'management' services — ML Ops (machine learning operations) and data SaaS. Just as humans need lifelong learning to grow smarter, AI must also be trained on data updated over time. And the data that exists must be managed to be used as efficiently as possible in AI systems. So AI training data companies need to prepare ML Ops and data SaaS services that target the same customer base they already have, enabling additional revenue streams and — over the long term — evolving toward a platform model.

Closing Thoughts

With all the attention AI has been receiving recently, I had assumed that the entire value chain would have grown into large-scale industries. Looking more closely, while expectations for the final product stage of AI (applications used directly by consumers) are high, the businesses in the earlier stages of the value chain — those that raise the overall quality of AI — are being valued in a more straightforwardly honest way. And just as the saying goes that "education is a hundred-year plan," it struck me that it will take quite a long time before the good data contributing to good AI shows its true worth. For an entrepreneur who believes AI will be the single most transformative technology of the future, can find satisfaction in steadily accumulating high-quality data while generating revenue, and whose ultimate mission is to have a positive impact on humanity — this is a genuinely compelling business model.

More from this category

Subscribe to the blog, and we'll email you when a new post goes up.

Business