top of page

Whoever writes the evals defines the product

AI features broke both of the ways product teams describe intended behaviour at once: the design handoff and the PRD. The artefact replacing them is the eval. And the people best placed to write the ones that matter most already know what good looks like.

AI

PRODUCT

DESIGN

TOOLS & TEMPLATES

3 person desk.png

The design handoff: the package design gives engineering when a design is finished. Annotated screens, interaction flows, state tables: the interface's intended behaviour, drawn.

PRD (Product Requirements Document): the product team's written description of what a feature will do and why, agreed before build.

Every team that builds software has a way of saying, before the development work starts, what should happen when someone uses the thing. Design teams mostly draw it. Product teams mostly write it down. AI features are quietly breaking both.​

​

Currently a design team hands off an AI feature the way they hand off everything else: annotated Figma frames, an interaction flow, a state table. Every state drawn, every edge case labelled. Six weeks later, something none of those frames describe goes live. The summary panel tells an already-angry customer their refund is "pending" when it was issued days ago, in a breezy tone nobody chose. Nothing in the handoff was wrong. The handoff was built for a system that does the same thing every time, and this one doesn't.​

​

Across the hall, a product manager opens the team's Product Requirements Document (PRD) template for their first AI feature and reaches the section that used to be the point of the whole document: expected behaviour. For a deterministic feature, you write down what the product will do, and the product does it. For a feature that hands part of the job to an AI model, the same sentences describe a product that doesn't exist.

​

Two different artefacts: the annotated handoff and the PRD. One failure: both specify fixed outcomes for a system that doesn't produce them.

noun-tandem-575479-148.png

Two lineages of saying what should happen

“It doesn’t matter how good your engineering team is if they are not given something worthwhile to build."

Marty Cagan, Inspired: How To Create Products Customers Love

Those two artefacts aren't rivals. They're parallel traditions of the same job: defining intended behaviour before it reaches users.​

​

Design's lineage is criteria-based. Critique, heuristic evaluation, usability tasks: an artefact judged against articulated criteria by people who know what good looks like. It's also deliberately light on documents, and that was the field's explicit best practice rather than a maturity gap. Marty Cagan told product teams from 2006 onwards that written requirements "take too long to write, they are seldom read", and that the high-fidelity prototype was the spec. If you're a designer thinking 'we never used PRDs': correct. You weren't behind; you were doing what the discipline said to do.

​

Product's lineage is promise-based: the PRD contains commitments about what the product will reliably do. When AI arrived, the diligent renovated it with new sections for fallback behaviour, confidence thresholds, escalation triggers, data dependencies, drift monitoring and failure modes. Necessary work, but still prose, making promises about a system that doesn't take instructions from prose.​
 

Both lineages have now run out of road for the same reason. And both are converging on the same replacement.​

noun-convergence-195965-148.png

The artefact both lineages were heading towards

Eval (evaluation): a defined input plus definitions of pass and fail, run against an AI feature before launch and on every change after it.

Slice: the same eval run for a specific user group or condition, checking quality holds for everyone.

For AI products and AI assisted development, many teams are moving away from these older artefacts and towards evals (short for evaluation). These are a way to define success based on outcome.

 

An eval typically contains:

An input - What information is being input into the AI system

A definition of pass - What success looks like

A definition of fail - What failure looks like

A slice - the same test, run for a specific group of users or type of input, to check quality holds for everyone, not just the "average" case

​​​​

Here's an example:​​

Feature

AI-suggested replies in a customer-support inbox.

Input

A complaint from an angry customer who has not been offered a refund.

A definition of pass

The draft reply acknowledges the problem and proposes next steps, without committing to anything that hasn't been approved.

A definition of fail

The draft promises a refund or credit nobody signed off, or breezes past the customer's frustration.

A slice

The same complaint, written in non-native English. Is the draft's quality and tone just as good?

As Braintrust's CEO Ankur Goyal puts it in Aakash Gupta's write-up:

"The prompt is temporary. The eval is permanent."

The whole thing runs before launch, after launch, and on every change. Unlike a promise, it can't quietly stop being true because it either passes or it doesn't. And remember the suggested reply from the start of this piece, that promised an angry customer a refund nobody approved? That is exactly the fail condition above. With this eval in the pipeline, that failure is caught in testing. Without it, it happens in front of a customer.

​

Now let's look back at the renovated PRDs that people use. Each of of the new sections that have been added is an eval in prose costume.
 

"Fallback behaviour": what the feature should do when the AI isn't sure - This is a fail condition waiting to be written as a test.
 

"Failure modes": the ways it can go wrong - These are your fail definitions, listed but not yet runnable.


"Drift monitoring": checking quality doesn't quietly degrade after launch - This is just evals, run continuously in production. The renovated PRD describes the tests. The eval suite is those descriptions, made to run.


"Escalation triggers": when something goes noticeably wrong - This is a threshold test for when the feature should hand off to a human

​

This isn't theoretical. Every change to GitHub Copilot has to pass more than 4,000 automated tests before it ships. Notion uses a eval suite to keep a 70-engineer AI team agreed on what good looks like. And when OpenAI shipped the ChatGPT update that turned it sycophantic and rolled it back within days, their own postmortem was blunt: the pre-launch evaluations "generally looked good". The behaviour that failed was the one no eval covered. 

noun-change-7266592-148.png

The pivot designers are already positioned to make

OpenAI's Chief Product Officer Kevin Weil says:

"Writing evals is going to become a core skill for product managers."

A $4,200 course on evals is reported as Maven's highest-grossing.

Demand for AI skills in design and UX postings grew 225% year on year, while only 5.8% of applicants to AI-requiring UX roles actually met the qualifications.

'Get on Board's Impact of AI 2025' report, drawn from more than two million applications across Latin American tech hiring

The evals discourse has a blind spot, and it's exactly designer-shaped.

 

There are many blogs and essays out there, even literally titled 'Evals are the new PRD'. But all of it is written for product managers, and most PM guides to evals don't mention designers at all. ​

​

This is strange, because if you look at what an eval actually contains, it sounds like design work: a realistic scenario, a definition of what good looks like for a user, and a judgment call about which failures are unacceptable. That's design-shaped work, and designers already produce it all the time. A usability task has the eval language built in: defined input, articulated success criteria, observed pass/fail. Anthropic's standard for a good eval, that "two domain experts would independently reach the same pass/fail verdict", is a description of a functioning critique culture. Or, in Goyal's version: "When you do a vibe check, you are using your brain as a scoring function… that is an eval. It is just the version that does not scale."


So this isn't a new discipline being imposed on design, it's the natural extension of work that designers should already be doing: researching what users need and expect, and setting out what good looks like. Evals formalise that judgment into criteria that can run automatically. Anthropic's guidance is clear that "the people closest to product requirements and users" are best positioned to define success. And the pivot has already begun, with content designers leading LLM evaluation activity at Indeed, and Microsoft's UX researchers treating the choice of eval metrics as UX research.

​

There are two honest boundaries though.

First, scope. User-driven evals are design's natural territory:
task success, tone, never-events (the things a feature must never do to a user), fairness across user groups. Things like latency, security and robustness evals belong with engineering.

Second, many designers need to learn how to operationalise this into evals. Coverage, versioning, criteria that hold up without you in the room. That gap is also why the opportunity is real: demand for this kind of judgment is running well ahead of the people who can supply it.

​

There's a payoff for the whole team, not just design. Handoff is the most complained-about moment in product development (92% of designers and 91% of developers say it needs improvement), and what developers consistently ask designers for is context and intent, not more pixels.

 

An eval is design intent in a format that survives the handoff. Development stays aligned to the intended design outcome rather than a spec document without context, and engineers get the "why" that they need when the model surprises everyone. Nielsen Norman Group's line about what now differentiates practitioners, "critical thinking, creativity, and taste — the ability to discern and curate a series of outputs and decisions", is, almost word for word, the content of a good eval.

noun-star-document-5974947-148.png

Who holds the pen

Autodesk's 2026 AI Jobs Report: Design skills are the number-one in-demand skill in AI job listings, and 'AI UX Designer' is new to its fastest-growing roles, up 145%.

So should anyone still write spec documents? To be fair to that lineage, it has serious, well-founded defenders. OpenAI's Sean Grove argues that the written specification is the thing that lasts, and the code is just one imperfect translation of it, in his words "a lossy projection from the specification". GitHub has built tooling on the same conviction: the spec is the source of truth, and code gets generated from it and checked against it. (Both are talking about AI writing code, not product PRDs — but the instinct carries over.) And they're right about one thing the eval camp under-weights: evals don't automatically carry intent. Prose and pictures still own why and who for. So the practical question isn't spec or eval, it's about how much of the story needs writing down before the evals take over as the contract.

 

Three questions settle most cases:

  • Could you already write three pass examples that engineering wouldn't argue with?

  • Is the behaviour still being discovered in prototyping?

  • If this feature gets it wrong, is the damage serious enough that leadership needs the full written case, not just the tests?

 

Evals have limits too, and they're worth naming. They guarantee the floor, not the ceiling: they will catch the reply that promises a refund nobody approved; they will not make the product delightful. And an eval doesn't make design quality objective, it just takes a judgment that used to live in someone's head and makes it visible, testable and open to challenge.

 

But the floor is where the stakes are. If designers and PMs don't write the user-driven evals, someone else's judgment can quietly governs the user experience, for example engineering optimising for whatever is easiest to measure. If engineering writes your evals alone, engineering decides which users' experience gets its own slice. (The warning cuts both ways. As one Mind the Product writer puts it, if the product team locks the eval set away from everyone else, you've just recreated the PRD problem: rules handed down without their context.)

 

Nobody owns this artefact yet. There are already courses and certifications are already teaching the skill, but the field hasn't settled how it should work, which means the ownership question is being answered right now, team by team. Whoever picks up the pen defines the product. It's a bigger job for both roles, not a smaller one: designers move from judging the experience after it ships to writing the definition of good that it has to meet before it ships; PMs move from describing what they intend to making it testable.​

noun-check-3550759-148.png

Start on Tuesday

If you're on a design team that has never written a PRD, don't start by writing one. Take your most recent usability test protocol, or your feature's most feared failure mode, and convert one criterion into eval structure: input, pass, fail. Bring it to engineering and ask, "would you run this?" If yes, you've started your eval set. If no, then listen to the objections: "we'd need more examples", "how would we score that?", "what happens when the input is X?" That's not a rejection; it's a to-do list. Each objection names the next thing to learn about making your criteria runnable.

​

If you're a PM with a PRD, check what's actually protecting you. Think of the worst thing your AI feature could plausibly do to a user. Now, three questions.

  • Is there anything in your process (a test, a check, a review step) that would catch it before a user sees it?

  • Who wrote that check?

  • When did it last actually run?

If the honest answers are "nothing", "nobody" and "never", your PRD isn't governing the product , it's just describing it.

 

Either way, write your never-events. What should this feature never do to a user? Write three; those are your first three evals. Make one of the three about a user group rather than a behaviour, and you've started thinking in slices.​


The obvious shortcut is to ask AI to write the evals for you. It can help: a model can grade outputs against criteria you wrote (the technique known as LLM-as-judge), and it can draft candidates for you to judge. But an auto-generated eval encodes the tool's generic judgment. It tests that the model behaves like a model, not that your feature behaves like your product. The criteria are the part you shouldn't delegate, because your reasoning about what good looks like is the product decision.

PHOTO-2026-08-06-17-55-33 1.png

Our template can guide you through each of the elements of an eval: the inputs , pass and fail definitions, slices, and the questions engineering will ask you. With this, you'll be be setting your own evals in no time. ⤵

Wormhole brings AI-powered guided workflows, integrated contextual bite-sized learning, and artefact generation together so you can 'design the right thing' before you 'design the thing right'. 

© wormhole | Design3 Network 2026

If you'd like to know more or contribute

bottom of page