Duply
Back to blog

28 September 2026

Design Arena AI: what it is and what it cannot tell you

Searches for "design arena ai" have gone from roughly 110 a month a year ago to 1,300 a month now, and most of the pages ranking for it are directory listings that do not say what the thing measures.

Quick answer: Design Arena is a crowdsourced benchmark that ranks AI models on design quality. You give it a prompt, four hidden models build your thing, you vote on which output looks best, and those votes become an Elo-style leaderboard. It is free to use, it is run by a company called Intelligence, and frontier labs buy the resulting human preference data. What it measures is which model wins a first impression on a blank page. It cannot tell you which model will build in your product's style.

Minimal illustration: four thin outlined rectangles arranged in a bracket, one of them highlighted in muted amber

What Design Arena actually is

Design Arena is at designarena.ai. Its Y Combinator profile (YC S25) describes it as "the first crowdsourced benchmark for AI-generated design" and says it evaluates "how models perform on real-world design tasks - frontend, audio, image" and more. Its own about page calls it "a subjective framework for evaluating AI design capabilities" and says the project "explores whether AI can exhibit taste".

The parent company is called Intelligence. It raised a $7.9 million seed round announced on 3 August 2026, led by Index Ventures with Conviction, A* and Valkyrie participating, reported by TechCrunch under the headline "Design Arena creators raise $7.9 million to bring taste to AI models". That article puts usage at "5.3 million people around the world, providing critical human evaluations to frontier labs".

Index Ventures' own write-up of the investment adds the customer list: "Google used Design Arena to test Gemini 3 ahead of launch", and "OpenAI, Thinking Machines Lab, xAI, and Meta have used it to evaluate their models". It also reports the platform grew "from $5 million to $60 million in ARR in six months".

The growth curve is visible on Reddit too. A July 2025 r/LocalLLaMA post from the team opens "We just hit 15K users!". That was about a year before the 5.3 million figure.

How the tournament works

The mechanics are documented on the about page and they matter, because the scoring method is what people argue about.

  1. You submit a prompt in one of the categories. The site lists fourteen, including websites, web apps, fullstack, mobile apps, games, UI components, data visualisation, 3D, images, logos, SVG, slides and text to speech.
  2. Four models are drawn from the active pool and each builds your thing.
  3. You vote blind. Model identities stay hidden "to prevent brand bias", and the bracket resolves through five battles, so one session produces several pairwise votes and a full first-to-fourth ranking.
  4. Votes become ratings through the Bradley-Terry model, a standard statistical method for pairwise comparison data. Strengths are normalised, then converted with the formula Rating = 400 × log₁₀(strength), which is why the numbers look like chess Elo.
  5. Every vote counts the same. The site states each comparison is "weighted equally with no filtering or editorial adjustment".

There is also a Builder Arena. The r/OpenAI launch thread describes it as covering "AI agents and builders like Devin, Lovable, Bolt.new, Same.new". That is the leaderboard most relevant to anyone shipping with a vibe-coding tool rather than a raw model.

Why the name confuses everybody

Three different things share this name space, and the search results mix all of them.

  • Design Arena (designarena.ai) is the design benchmark described here.
  • Arena AI (arena.ai) is the former LMArena, the Berkeley chatbot leaderboard. An r/aicuriosity thread records the rename: "Arena.ai which everyone knew as LMArena or Chatbot Arena just got a fresh new name." It ranks on page one for "design arena ai".
  • Designar.ai is an unrelated Arabic-language design app that also surfaces on the query, and a Dallas company at designsarena.com has its own Trustpilot page in the results.

Google is confused too. On the live results for "design arena ai" pulled on 28 September 2026, the People Also Ask box asks "Is design Arena AI completely free?", then "How much does Arena AI cost?", "How is Arena AI making money?" and "Who is behind Arena AI?" Two products, one question block.

The second source of confusion is the product itself. Design Arena's homepage now reads "the place to create anything with the world's best AI models", and the related searches for the query are "Design Arena AI apk", "Design Arena AI app download" and "Design Arena AI video generator". A benchmark that lets you generate things for free looks, to most searchers, like a free generator that happens to keep score. Both readings are correct. The voting is the product; the generation is the bait.

What the leaderboard is good for

✅ Comparing raw model taste on a cold start, with brand names hidden ✅ Spotting a new model that is genuinely better at front-end output before the marketing lands ✅ Tracking open-weight models against closed ones on visual quality ✅ Seeing how builders like Lovable and Bolt stack up on the same prompt

What it cannot tell you

❌ Whether a model will match your design system ❌ Whether the winning output was the one that followed the prompt ❌ Whether the design survives past the first screen ❌ Which model to standardise on for a product that already has a look

The second one is the substantive criticism, and it came up in the founders' own Launch HN thread. A commenter pointed out that contests get won not by the entry that best adhered to the prompt but by the best looking one, citing a brutalism competition won by a magenta and yellow design that contradicted the brief. Another raised the local maximum problem: picking the best of four mediocre options rewards whichever is least bad.

A thread from last week on r/LocalLLaMA titled "What happened to Design Arena?" is blunter. The snippet Google shows is one commenter's accusation that the team used "dodgy statistical methodology". Treat that as an opinion from a single thread, not a finding. The verifiable part is the underlying question: an unweighted public vote on a one-shot prompt is a measure of first impressions, and first impressions favour contrast, colour and density over restraint.

That is the whole gap for anyone building a real product. A model that wins a blind vote by producing the most striking landing page is not the model that will quietly match your spacing scale.

What Reddit actually uses it for

The threads that get traction are not about design. They are about the open-weight race.

Nobody in those threads is using the leaderboard to make their own app look better. They are using it to track the model race.

The part the leaderboard leaves to you

Pick the top-ranked model today and your output still looks like the model, not like your product. That is the finding behind why AI-generated apps all look the same: with no design constraints in context, every model falls back to its defaults, and the defaults are shared.

The fix is not model selection. It is giving the model something specific to build in:

  1. Pick a design you want to build in. Browse the duply library and choose a real product's system, for example Linear, Stripe or Vercel.
  2. Copy the DESIGN.md. One file with the tokens as data and the usage rules as text. See what a DESIGN.md file is if the format is new to you.
  3. Paste it into the agent's context in Cursor, Claude Code, Lovable or v0, and reference it in the prompt. The full method is in give your AI agent a design system.
  4. Re-run the same prompt and compare. The model changes less than the constraints do.

Use Design Arena to decide which model to reach for. Use a design system to decide what it builds.

FAQ

What is Design Arena AI? A crowdsourced benchmark that ranks AI models on design quality. You submit a prompt, four hidden models generate an output, you vote on which is best, and the votes feed an Elo-style leaderboard across fourteen categories including websites, UI components, logos and video.

Is Design Arena free? Yes, generating and voting are free. The company monetises by selling the resulting human preference data and version testing to AI labs, which is why Index Ventures reports Google used it to test Gemini 3 before launch.

Is Design Arena the same as Arena AI or LMArena? No. Arena.ai is the renamed LMArena, the Berkeley chatbot leaderboard for general model quality. Design Arena is a separate company, Intelligence, benchmarking design and front-end output specifically. Google's own results page mixes the two.

How are Design Arena rankings calculated? With the Bradley-Terry model for pairwise comparisons. Strengths are normalised and converted to ratings with Rating = 400 × log₁₀(strength), which produces chess-style Elo numbers. Each vote is weighted equally, with no filtering or editorial adjustment.

Who owns Design Arena? A company called Intelligence, founded in 2025 and led by co-founder Grace Li. It raised $7.9 million in seed funding announced on 3 August 2026, led by Index Ventures.

Can I trust the Design Arena leaderboard? For what it measures, which is blind public preference on a single prompt, the method is transparent and documented. The known criticism, raised in the founders' own Launch HN thread, is that voters reward the best looking output rather than the one that best followed the prompt.

Does the top model on Design Arena make the best app? Not necessarily. The leaderboard scores one-shot output with no constraints. Once you hand a model your tokens and rules, the gap between the top few models narrows sharply, and the constraints do more work than the ranking does.

What is the Builder Arena? The leaderboard for agents and app builders rather than raw models, covering tools like Devin, Lovable, Bolt.new and Same.new on the same head-to-head format.

Summary

  • Design Arena ranks AI models on design by blind public vote, scored with Bradley-Terry and displayed as Elo
  • It is free, run by Intelligence, funded with a $7.9M seed led by Index Ventures, and used by Google, OpenAI, xAI and Meta for evaluation
  • The name collides with Arena.ai, the renamed LMArena, and Google's own results page mixes them
  • The documented weakness is that a blind vote rewards the best looking output, not the one that followed the prompt
  • Use it to pick a model; use a design system to control what that model builds

Pick a design you want your agent to build in, copy one file, and stop relying on model defaults. Browse the duply library.

Related reading