Encord co-founder Ulrik Stig Hansen raised $50M by identifying a critical gap in the AI development stack: data quality, not just models. By conducting over 1,000 customer interviews, he and his co-founder found a deep need for better data annotation, quality control, and model evaluation tools. This article breaks down their playbook for finding a market wedge, leveraging Y Combinator, and crafting a fundable narrative for an AI infrastructure company.
Key takeaways
- Your AI moat isn't the model, it's your data engine.
- Interview 100+ potential users to find the real, painful bottleneck.
- Shift focus from data quantity to data quality and curation.
- A non-linear career path can be a strategic advantage.
- Use accelerators like YC to avoid common mistakes and build a network.
- Frame your fundraising pitch around a market shift, not just your product.
Your AI Moat Isn't the Model, It's the Data Engine
In the gold rush of AI, most founders are chasing the same thing: a better model. But as foundation models become powerful and accessible, the model itself is ceasing to be a durable competitive advantage. The real, defensible moat is the engine you use to create, curate, and improve your data.
Encord, a data infrastructure platform that has raised over $50M, is a masterclass in this thesis. Co-founder Ulrik Stig Hansen and his team built their company not by creating a new model, but by solving the painful, unglamorous problems of the AI data stack. Their journey from a small town in Denmark to Y Combinator to a $30M Series B offers a tactical roadmap for any founder building in AI.
The Insight: Find the Market Gap by Asking the Right Questions
Before writing a line of code, Ulrik and his co-founder Eric conducted over 1,000 interviews with AI engineers, researchers, and data scientists. They didn't ask, "Would you buy our product?" They dug into the core of the workflow to find the most painful, expensive part of the process.
This is the critical first step most founders skip. They build a solution for a problem they assume exists, rather than the one the market is desperate to solve.
Common Mistake: Building for a Theoretical Problem
Many founders, especially technical ones, fall in love with a technology and then search for a problem. Encord did the opposite. They started with the ecosystem and interrogated its daily workflow.
How to Do It Right: A Playbook for Customer Discovery
Your goal is to find the bottleneck that costs teams time and money. Steal these questions for your own discovery process:
"Walk me through your end-to-end process for shipping a new model, from data collection to deployment." · "Which step in that process is the most manual, frustrating, or prone to error?" · "How do you currently measure the quality of your training data? What do you do when you find a quality issue?" · "Tell me about a time a model failed in production. What was the cause? How did you debug it?" · "If you had a magic wand to fix one part of your data pipeline, what would it be?"
For Encord, the answer became clear: the bottleneck was shifting from the quantity of data to the quality. Teams were drowning in data but starving for high-quality, correctly labeled, and well-curated datasets to feed their models.
From Quantity to Quality: The Three Pillars of a Data Engine
The insight from their user research led Encord to build a platform focused on the full data lifecycle. Bad data doesn't just produce bad results; it can "poison" a model, degrading its performance over time in subtle ways that are hard to debug. A self-driving car that misinterprets a stop sign isn't failing because the algorithm is bad, but because its training data was flawed.
Data Annotation and Creation: Building the tools to label vast amounts of data accurately and efficiently. This is the foundational layer. · Data Curation and Feedback: Giving humans a systematic way to provide feedback on model performance (a process related to Reinforcement Learning from Human Feedback, or RLHF). This creates a tight loop where the model constantly learns from real-world corrections. · Model Evaluation: Creating a "test harness" to diagnose where and why a model is failing. This moves teams from "the model is 89% accurate" to "the model consistently fails to identify pedestrians in rainy conditions at dusk," which is an actionable insight.
By solving all three, Encord provides a complete data engine, allowing teams to treat their data as a product, not an afterthought.
The Founder's Path Is Not Linear
Ulrik’s background wasn't that of a typical AI researcher. He started his career in finance at JP Morgan in London. Many founders would see this as a detour, but it was a strategic advantage. Working in finance gave him a deep understanding of market dynamics, enterprise needs, and how to build a real business—not just a research project.
He paired this business acumen with a Master's in Computer Science from Imperial College London, diving deep into AI precisely as transformer models were beginning to show their world-changing potential. It was there he met his co-founder, Eric, who brought a background in high-frequency trading and production-grade systems.
This combination of commercial awareness and deep technical expertise is the ideal founder pairing for an AI infrastructure company. One understands the market and the customer; the other understands how to build a robust, scalable system to serve them.
Common Mistake: The "Academic" Founder
A common failure mode for AI startups is being founded by brilliant researchers who have never shipped a product or sold a deal. They build fascinating technology that nobody will pay for. Ulrik’s path shows the value of having one foot in the world of business and the other in deep tech.
Leveraging Y Combinator to Avoid Rookie Mistakes
Encord joined Y Combinator in the Winter 2021 batch, which was fully remote due to the pandemic. While some founders might devalue a remote accelerator experience, Ulrik found it immensely valuable for avoiding common first-time founder traps.
An accelerator like YC is less about magic formulas and more about mistake prevention. For an AI company, the advice is often brutally effective:
Stop writing papers, start talking to users. Your success is measured in paying customers, not citations. · Find one person to pay you for a terrible V1. The path to product-market fit starts with a single, desperate customer. · Your first hires must be builders. Avoid hiring specialized researchers or marketing teams until you have a product that works and customers who love it.
The YC network also proved invaluable for fundraising. But even with the YC stamp of approval, Ulrik emphasizes that fundraising is never easy—especially in AI, where investors are wary of science projects masquerading as businesses.
How to Raise $50M for an Infrastructure Company
Encord’s fundraising journey—culminating in a $30M Series B—was built on a powerful narrative. Fundraising for an infrastructure company requires you to sell a story about a fundamental market shift.
The Fundraising Narrative: From Insight to Inevitability
Your pitch isn't just a set of slides; it's a compelling argument in three acts:
The World has Changed: "The AI industry has moved from being model-constrained to being data-constrained. The best models are now available via API, so the only defensible asset is a proprietary, high-quality data pipeline." · The Unsolved Problem: "But building this data engine is incredibly hard. It requires specialized tooling for annotation, quality control, and evaluation that no off-the-shelf solution provides. This is the pain point we heard from 1,000 engineers." · Our Inevitable Solution: "We built the platform that solves this entire workflow. Our customers are building better models faster, and we are becoming the essential data layer for the next generation of AI. Here’s the enterprise traction to prove it."
Raising a Series B of this magnitude requires more than a story; it requires proof. By this stage, investors need to see evidence of a scalable go-to-market motion, significant annual recurring revenue (ARR), and strong net revenue retention (NRR) from happy enterprise customers.
How to Apply This This Week
You don’t need to have it all figured out today. Here are three concrete actions inspired by Encord's journey:
Map your data pipeline. Whiteboard the journey of a single piece of data from its raw state to its use in your model. Identify the most manual, slow, or unreliable step. That's your starting point. · Schedule five "problem discovery" calls. Reach out to 5-10 people in your target user group. Use the questions above and just listen. Don't pitch, don't sell—just learn. · Write your "market shift" paragraph. In 100 words or less, articulate why your corner of the world has fundamentally changed and what new problem has been created as a result. This is the seed of your fundraising narrative.
Frequently asked questions
- What is an AI 'data engine'?
- It's the complete system for sourcing, cleaning, labeling, and improving the data your AI models train on. A good data engine turns raw information into a proprietary asset that gives you a competitive edge.
- How much should a pre-seed AI startup raise?
- A typical range is $1M to $3M. This capital is primarily for hiring a small team of engineers, securing initial GPU compute resources, and landing your first few pilot customers to prove the concept.
- Do I need a PhD in AI to start an AI company?
- No. It's often more valuable to have a co-founder with deep experience shipping production systems, like Encord's co-founder Eric. The key is building a product customers will pay for, not just publishing research.
- What's the biggest mistake early-stage AI founders make?
- Focusing exclusively on the model architecture while ignoring the data pipeline. Many founders discover too late that their biggest bottleneck isn't the algorithm, but the quality and management of their training data.