Tasks / Design

One-shot prototype

Can the configuration produce a working, constraint-compliant prototype in one attempt?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models.

Reliably right

  1. Avoids template layout and copy tells100% pass
  2. Shows the no-slots state100% pass
  3. Respects explicit constraints83% pass

Where it slips

  1. Says what it tests and what's faked46% pass
  2. Avoids the generic AI palette67% pass
  3. Stays in scope75% pass

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Build a clickable prototype of a flow where an employee captures a receipt, checks the amount the app extracted, corrects it, and saves the expense. Deliver a single self-contained HTML file that opens straight into the flow in a browser: no design canvas, artboard or "play" step. Use made-up data.

ScenarioAlex Morgan is submitting a taxi receipt from 14 September for £38.40. The app's automatic extraction misreads the amount as £83.40.
ConstraintsDesktop web. Receipt capture is simulated: a control such as "Use sample receipt" stands in for the camera or file upload, with no real upload needed. Show the extracted amount and let the employee correct it before saving; the saved expense shows the corrected amount.
What a strong answer does

Capture (simulated), a visible extracted amount, an easy correction, and a saved expense showing the corrected amount.

Critical failures (cap the score)
  • Core interaction broken
Case

v1.1 · synthetic · fintech

3 of 6 checks passed

Output

Open the prototypeThe file the model delivered, in a new tab

I've built the prototype as a single clickable desktop screen, 1280 wide, in the design canvas. Alex Morgan goes from capture to review to saved expense. I haven't run through it in a browser yet, so the clicks and the amount field still need a first test.

Capture: "Use sample receipt" stands in for the camera or upload. The "Upload file" button is greyed out, and a line says it's simulated. Reading: a short scanning animation plays, then the app moves on by itself. Review: the receipt image sits on the left and shows a total of £38.40. The amount field comes pre-filled with the misread £83.40, marked amber with "Check · low confidence", and a banner asks Alex to check it against the receipt. While the amount field is selected, the TOTAL line on the receipt is highlighted. Correcting: typing 38.40 turns the field green with "Changed from £83.40". There's an "Undo change" button, the save button updates to "Save expense · £38.40", and saving is blocked if the amount isn't a valid number. Saved: a confirmation appears, and the new expense is at the top of the September list with the corrected £38.40. Under it a note reads "Amount corrected by you · read as £83.40", and the month total includes the new amount.

Where I filled gaps:

Only the amount is editable. Merchant, date, category and card are shown read-only, since the flow is about the amount. Saving without correcting is allowed, as it would be in a real app. In that case the £83.40 is what gets saved. Made-up details: the app is called "tally", the taxi firm is "CityLine Cars" (£34.40 fare + £2.50 booking fee + £1.50 tip), the approver is Priya Shah, and there are three earlier September expenses.

"Add another expense" or "Use a different receipt" resets the flow.

Check by check

Mixed · 3
  • Respects explicit constraintsThe two graders disagreed on this one.
  • Stays in scopeThe two graders disagreed on this one.
  • Avoids the generic AI paletteThe two graders disagreed on this one.
Got right · 3
  • Core interaction works
  • Avoids template layout and copy tells
  • Says what it tests and what's faked

Grades and run details

Decision model 75
Decision model checks
  • partialRespects explicit constraints23%
  • passCore interaction worksby hand100%
  • partialStays in scope29%
  • partialAvoids the generic AI palette25%
  • passAvoids template layout and copy tells48%
  • passSays what it tests and what's faked35%
Artefacts
Run
Run
#1
Time to output
2.6 min
Submitted
24 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 86% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI92.9–2None
2GPT-6.1 SolwithAPI88.7–2None
3GPT-6 LunawithAPI84.5–2None
4Opus 5.5withClaude76.8–2None
5GPT-6 AstrawithChatGPT69.6–2None
6Gemini 3.5 Flash-LitewithGemini49.4–2None

About the task

The PM job

Getting a clickable prototype in front of users or stakeholders fast.

Why it matters

A prototype that looks right but whose core interaction is broken wastes a research session. This task rewards working behaviour over polish.

What good looks like

  • The core interaction works end to end
  • Respects the stated constraints and scope
  • Handles the empty and error states named in the brief
  • Says what it's testing and what's faked
  • Looks like something you'd share: clear hierarchy, none of the stock AI-template look

Deliberately not measured

  • Production code quality
  • Design to a requested style: the brief never asks for one, so this measures default taste
Capability tested

Turning a spec into a usable interactive artefact

The failure we’re looking for

An attractive shell with a broken or absent core interaction

Grading

Browser checks (interaction and design) and decision model, calibrated against a blind PM review

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.