Bar chart displaying Elo ratings for various AI models, with GPT Image 2 leading at 1388. DesignArena logo at top left.

DesignArena’s Parent Raises $7.9 Million as AI Labs Pay for Human Taste

DesignArena creator Intelligence raises $7.9M as AI labs pay for human feedback, with 5.3M users and reported $60M ARR.

In short

Intelligence, the company behind DesignArena, raised $7.9 million to expand a platform that uses human rankings to help AI labs improve generative products. The startup says the service has 5.3 million users and $60 million in ARR, underscoring investor interest in human feedback infrastructure.

  • Intelligence raised $7.9 million in a seed round led by Index Ventures.
  • The company says DesignArena has 5.3 million users and $60 million in annual recurring revenue.
  • Its core product turns human preference rankings into training and evaluation data for AI labs.
  • The funding reflects rising demand for human-led AI evaluation as automated benchmarks become easier to game.

Intelligence, the company behind the AI feedback platform DesignArena, has raised $7.9 million in seed funding as it positions human judgment as a critical layer in training and improving generative AI products. The round, announced Monday, was led by Index Ventures and comes as the startup says its platform already reaches 5.3 million users and is generating $60 million in annual recurring revenue.

The financing highlights a growing belief in the AI industry: even as models improve at speed, companies still need people to decide whether outputs are genuinely useful, attractive, or fun. Intelligence first emerged from a failed attempt to build an AI game engine, but that effort led the founders to a new business — a system for collecting large-scale human preference data from everyday users and selling that signal to frontier AI labs.

That market is becoming more valuable as companies look for better ways to evaluate text, images, websites, and other generated content. The pitch is simple: automated benchmarks can be gamed or miss the subtleties of taste, while human rankings can reveal what people actually prefer.

How DesignArena turned a product problem into a business

DesignArena began as a side project among college friends who were trying to make an AI game engine work in the weeks before graduation in 2025. The prototype could generate playable games, but the founders kept running into the same issue: the results functioned, yet they lacked the quality and appeal that would make people want to play them.

That gap forced a more fundamental question. If a model can produce a working product, how do you measure whether it is actually good? For the founders, the answer was not another benchmark. It was scaleable human preference data.

Grace Li, Intelligence’s co-founder, has described the early realization as a search for the missing piece in model improvement. In her telling, the company quickly discovered that there was real demand from AI labs that wanted more than generic performance metrics — they wanted evidence about what users preferred in real settings.

Li said the platform filled a bottleneck for AI systems trying to improve on design-related tasks, and that the company landed its first major frontier-lab customer within about a week of reaching that insight.

What started as an internal frustration evolved into a consumer product and then into enterprise infrastructure. That progression matters because it reflects a broader trend in AI: tools that look playful or experimental at first are increasingly becoming core data pipelines for the companies building the next generation of models.

What exactly does DesignArena do?

DesignArena is built to turn user preference into training signal. On the consumer side, the experience resembles a model-routing or comparison interface rather than a standard chatbot. Users enter prompts and choose the format they want — such as websites, images, or other visual outputs — and the system produces multiple responses for comparison.

Instead of asking users to grade one result in isolation, the platform presents a series of side-by-side decisions. Users pick between “A” and “B,” then keep ranking outputs until a preference order emerges. That structure helps the company collect richer data on what people consider better, not just what they consider acceptable.

The design is important for another reason: many participants are not loyal to a particular underlying model. They are typically focused on finding the strongest result. That makes their choices especially useful as an indicator of what ordinary users value, rather than what model fans or technical evaluators might prefer.

Why enterprises care

For AI companies, those rankings can be far more valuable than a one-off usability test. They create a steady stream of feedback that can help refine image generators, website builders, design systems, and other media tools.

In theory, automated evaluation can process far more outputs than people can. In practice, however, automated scoring often fails to capture nuance — especially in creative tasks where style, layout, and aesthetic judgment matter. Human taste remains difficult to model at scale, which is precisely why Intelligence believes customers will pay for access to it.

The company says the value proposition is strong enough that the business is already producing significant revenue. If accurate, the reported $60 million ARR would put Intelligence among the more commercially successful startups in the AI-evaluation category, a space that is beginning to draw serious venture attention.

Key metric Reported figure Why it matters
Seed funding $7.9 million Fresh capital to expand the platform and enterprise sales
Lead investor Index Ventures Signals strong institutional confidence in the category
Reported users 5.3 million Shows consumer reach and potential data scale
Reported ARR $60 million Suggests the product has converted attention into meaningful revenue
Primary use case Human preference ranking Supports model evaluation and improvement

Why human feedback is becoming more valuable

The AI industry has reached a point where raw benchmark scores are no longer enough. As models become more capable, the weaknesses in standard tests become easier to exploit. A model can optimize for a benchmark without truly getting better at the real-world task users care about.

That is particularly true in design-heavy applications. Whether the output is a website, an image, a landing page, or another visual asset, taste is hard to quantify. A score can tell a company which model is technically faster or cheaper. It cannot always tell the company which result people would actually choose.

Intelligence is betting that these subjective choices are not a weakness in the evaluation process but the point of it. If people can reliably identify the better result, then their preferences become training data. In that sense, the platform is not just a poll or a feedback form. It is part of the model-development stack.

How the platform measures taste across regions

One feature of the system is that users must log in before receiving their output. That gives Intelligence the ability to analyze preferences over time and by geography, which can reveal how aesthetic tastes differ from one market to another.

Li has said the company sees differences across continents and described some regional design tendencies in broad strokes, including a more maximalist approach on dashboards in parts of Asia. Those kinds of patterns can matter a great deal to products that aim to localize style rather than simply translate language.

For AI builders, this type of data can support more than model tuning. It can inform product decisions about layout, verbosity, composition, and presentation — all the small judgments that shape whether an AI-generated experience feels polished or generic.

What does the funding round say about investor appetite?

The $7.9 million round suggests investors see a new category emerging around AI evaluation and taste collection. Index Ventures led the deal, with participation from Conviction — the firm associated with Sarah Guo and Mike Vernal — as well as A*, Valkyrie, and others.

That roster matters because it shows the market is not just backing model developers or infrastructure companies. It is also funding the tools that help determine whether those models are any good. In a crowded AI market, the picks-and-shovels layer is increasingly attractive when it directly connects to enterprise spend.

The startup’s growth also points to the monetization potential of consumer-generated evaluation data. By pairing free or low-friction user participation with paid enterprise access, Intelligence appears to have built a hybrid business that can scale both data collection and revenue.

How it compares with other AI evaluation startups

DesignArena is not alone in trying to industrialize human judgment. A number of startups have explored similar ideas, with mixed results. Some have attracted major funding and fast adoption, while others have discovered that enthusiasm does not always translate into durable economics.

That tension is one of the biggest questions in this corner of the market. If human evaluation is so valuable, why do some companies struggle to stay alive? The answer may lie in execution, customer concentration, and the challenge of turning one-off demand into recurring contracts.

Still, the category is clearly gaining credibility. LM Arena, which focuses on text-based responses, raised $150 million in a Series A in January, only a few months after introducing its paid product. That kind of fundraising suggests investors believe the market for comparative evaluation is more than a short-lived trend.

Company Focus Funding Status
Intelligence / DesignArena Visual and design preference ranking $7.9 million seed Ramping enterprise business
LM Arena Text response evaluation $150 million Series A Rapidly scaling paid offering
Yupp Crowdsourced feedback $33 million raised Shut down after less than a year

Why did Yupp fail where others are growing?

Yupp is a reminder that demand for user feedback does not automatically create a lasting business. The company raised $33 million, including backing from a16z crypto co-founder Chris Dixon, and reported more than 1.3 million users. Even so, it closed earlier this year after struggling to find a sustainable path.

The failure suggests that scale alone is not enough. A startup in this space must do several things at once: attract enough users to generate meaningful data, persuade AI labs to pay for that data, and maintain a business model that can survive beyond the initial wave of excitement.

That is a difficult balance. Consumer participation tends to be cheap at the start but expensive to maintain. Enterprise demand may be strong, but large customers expect differentiated value, reliability, and data quality. If one of those pieces falters, the business can quickly become fragile.

What makes DesignArena different?

Based on the company’s reported traction, Intelligence appears to have found a more specific use case than some predecessors. Rather than trying to be a generic feedback marketplace, it has focused on visual and design generation, where subjective judgment is especially difficult for machines to replace.

That narrower focus may give the startup an advantage. AI labs building websites, mockups, images, and other design-adjacent products have an immediate need for high-quality preference data. If the company can own that niche, it may avoid some of the broader competition that undermined earlier efforts.

It also helps that the product itself can be engaging for users. Ranking outputs can feel more interactive than filling out surveys, and that interactivity may help the company gather enough signal to make the system commercially useful.

What happens next for Intelligence?

The new capital will likely be used to expand the company’s platform, deepen enterprise relationships, and sharpen the data products it sells to AI labs. The bigger question is whether Intelligence can convert a hot niche into a durable category before larger platform players build similar tooling in-house.

That question matters because AI companies are increasingly trying to control as much of the training and evaluation pipeline as possible. If a startup like Intelligence becomes essential, it could enjoy a strong moat. If not, its customers may decide to recreate the system internally.

For now, the fundraising is a strong vote of confidence. It suggests that investors think human-led evaluation remains one of the most underbuilt and commercially relevant parts of the AI stack.

Timeline of the company’s rise

  • Early 2025: The founders begin building an AI game engine around graduation.
  • Within weeks: They realize the model can make functional games, but not enjoyable ones.
  • Shortly after: The team pivots toward collecting human preference data at scale.
  • About a week later: Intelligence says it closes its first major frontier-lab deal.
  • Now: DesignArena reports 5.3 million users and $60 million in ARR, alongside a new $7.9 million seed round.

Why this matters for the broader AI market

This funding round is about more than one startup. It reflects a larger shift in how the AI industry thinks about progress. For years, the focus was on bigger models, more data, and more compute. Now, companies are spending more attention on whether those models can be judged accurately and improved in ways that match human preference.

That shift has implications across the sector. A better evaluation layer can change which models get adopted, which products win, and which companies can prove that their systems are actually getting better. In markets where style and subjective quality matter, that may be as important as raw intelligence.

It also reinforces a key reality about AI: some of the hardest problems are not purely technical. They are human. Taste, preference, and judgment remain central to the experience of using generative tools, and companies that can capture those signals may hold an outsized advantage.

For Intelligence, the challenge now is to prove that its early momentum can last. The startup has an eye-catching user count, a notable revenue claim, and a fresh infusion of capital. Whether it can turn those ingredients into long-term category leadership will depend on one thing the AI industry still struggles to automate: knowing what people actually like.

Frequently asked questions

What is DesignArena?

DesignArena is a human feedback platform built by Intelligence that lets users compare AI-generated outputs and rank them by preference. The company says those rankings help AI labs improve design, image, and website-generation models by revealing what people actually think is better.

How much funding did Intelligence raise?

Intelligence raised $7.9 million in seed funding. Index Ventures led the round, and the company said other participants included Conviction, A*, Valkyrie, and additional backers.

Why do AI companies pay for human feedback data?

AI companies pay for human feedback because automated benchmarks often miss subjective qualities like taste, usefulness, and visual appeal. Human rankings can show which outputs real users prefer, giving model builders a more practical signal for improving generative systems.

How big is DesignArena’s business?

Intelligence says DesignArena has reached 5.3 million users worldwide and is generating $60 million in annual recurring revenue. Those figures suggest the company has already turned consumer participation into a substantial enterprise business.

Did other startups try the same idea?

Yes, other startups have pursued similar human-evaluation businesses, with mixed results. LM Arena has raised significant funding for text-response evaluation, while Yupp, which also focused on crowdsourced feedback, shut down after raising $33 million.

Share this 🚀