Tasks / Design

One-shot prototype

Can the configuration produce a working, constraint-compliant prototype in one attempt?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models.

Reliably right

  1. Avoids template layout and copy tells100% pass
  2. Shows the no-slots state100% pass
  3. Respects explicit constraints86% pass

Where it slips

  1. Says what it tests and what's faked43% pass
  2. Avoids the generic AI palette71% pass
  3. Stays in scope75% pass

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Build a clickable prototype of a flow where an employee captures a receipt, checks the amount the app extracted, corrects it, and saves the expense. Deliver a single self-contained HTML file that opens straight into the flow in a browser: no design canvas, artboard or "play" step. Use made-up data.

What the model was given2 items: Scenario, Constraints
ScenarioAlex Morgan is submitting a taxi receipt from 14 September for £38.40. The app's automatic extraction misreads the amount as £83.40.
ConstraintsDesktop web. Receipt capture is simulated: a control such as "Use sample receipt" stands in for the camera or file upload, with no real upload needed. Show the extracted amount and let the employee correct it before saving; the saved expense shows the corrected amount.
What a strong answer doesThe answer key the graders mark against

Capture (simulated), a visible extracted amount, an easy correction, and a saved expense showing the corrected amount.

Critical failures (cap the score at 40)
  • Core interaction broken
Case

Mid-level · v1.1 · synthetic · fintech

The outputs

7 setups answered this brief. Pick up to three to read side by side.

3 of 6 checks passed

Output

Open the prototypeThe file the model delivered, in a new tab

Here is the clickable prototype in a single self-contained HTML file.

Grades and run details

Decision model 58
Decision model checks
  • partialStays in scope26%
  • failSays what it tests and what's faked91%
  • passRespects explicit constraintsby hand100%
  • failCore interaction worksby hand100%
  • passAvoids the generic AI paletteby hand100%
  • passAvoids template layout and copy tellsby hand100%
Artefacts
Run
Run
#1
API response time
46 s
Submitted
1 Oct 2026

Check by check

Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Got wrong 2

Core interaction worksWrong

No reason given.

Says what it tests and what's fakedWrong

No reason given.

Mixed 1

Stays in scopeMixed

No reason given.

Got right 3

Respects explicit constraintsRight

No reason given.

Avoids the generic AI paletteRight

No reason given.

Avoids template layout and copy tellsRight

No reason given.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 85% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI92.9–2None
2GPT-6.1 SolwithAPI88.7–2None
3GPT-6 LunawithAPI84.5–2None
4Opus 5.5withClaude76.8–2None
5Gemini 3.8 FlashwithAPI75.6–2None
6GPT-6 AstrawithChatGPT69.6–2None
7Gemini 3.5 Flash-LitewithGemini49.4–2None

About the task

The PM job

Getting a clickable prototype in front of users or stakeholders fast.

Why it matters

A prototype that looks right but whose core interaction is broken wastes a research session. This task rewards working behaviour over polish.

What good looks like

  • The core interaction works end to end
  • Respects the stated constraints and scope
  • Handles the empty and error states named in the brief
  • Says what it's testing and what's faked
  • Looks like something you'd share: clear hierarchy, none of the stock AI-template look

Deliberately not measured

  • Production code quality
  • Design to a requested style: the brief never asks for one, so this measures default taste
Capability tested

Turning a spec into a usable interactive artefact

The failure we’re looking for

An attractive shell with a broken or absent core interaction

Grading

Browser checks (interaction and design) and decision model, calibrated against a blind PM review

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.