Tasks / Design

One-shot prototype

Can the configuration produce a working, constraint-compliant prototype in one attempt?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models.

Reliably right

  1. Avoids template layout and copy tells100% pass
  2. Shows the no-slots state100% pass
  3. Respects explicit constraints83% pass

Where it slips

  1. Says what it tests and what's faked46% pass
  2. Avoids the generic AI palette67% pass
  3. Stays in scope75% pass

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Build a clickable prototype of a flow where an employee captures a receipt, checks the amount the app extracted, corrects it, and saves the expense. Deliver a single self-contained HTML file that opens straight into the flow in a browser: no design canvas, artboard or "play" step. Use made-up data.

ScenarioAlex Morgan is submitting a taxi receipt from 14 September for £38.40. The app's automatic extraction misreads the amount as £83.40.
ConstraintsDesktop web. Receipt capture is simulated: a control such as "Use sample receipt" stands in for the camera or file upload, with no real upload needed. Show the extracted amount and let the employee correct it before saving; the saved expense shows the corrected amount.
What a strong answer does

Capture (simulated), a visible extracted amount, an easy correction, and a saved expense showing the corrected amount.

Critical failures (cap the score)
  • Core interaction broken
Case

v1.1 · synthetic · fintech

6 of 6 checks passed

Output

Open the prototypeThe file the model delivered, in a new tab

This prototype opens straight on the capture screen. Click Use sample receipt to start the flow.

  • Extraction: the app reads the taxi receipt and gets the total wrong, showing £83.40 instead of £38.40.
  • Correction: the amount is flagged as low confidence. The receipt sits next to the form with a callout on the total, and the fare, fee and tip add up to £38.40.
  • Saving: you can't save until you either correct the amount or confirm it as it is. A corrected amount saves as £38.40 and appears at the top of "My expenses". The confirmation notes that it was corrected from £83.40.

Notes on the prototype:

  • Data: everything is made up (Metro Cabs Ltd, Alex Morgan, the other expenses in the list). The date is set to 14 Sep 2025.
  • Amount states: the amount field moves through a low-confidence warning, a corrected state with a "Restore" link, and a validation error. The tag on the receipt changes with it.
  • Confirming without editing: if someone clicks "It's correct" on £83.40 without changing it, the expense saves with that amount and the confirmation says they confirmed it.
  • Resetting: "Add another expense" and "Start over" reset the form to the misread £83.40 so you can run the flow again. Expenses saved during the session stay in the list until you reload the page.

Check by check

Got right · 6
  • Respects explicit constraints
  • Core interaction works
  • Stays in scope
  • Avoids the generic AI palette
  • Avoids template layout and copy tells
  • Says what it tests and what's faked

Grades and run details

Decision model 100
Decision model checks
  • passStays in scope36%
  • passSays what it tests and what's faked73%
  • passRespects explicit constraintsby hand100%
  • passCore interaction worksby hand100%
  • passAvoids the generic AI paletteby hand100%
  • passAvoids template layout and copy tellsby hand100%
Artefacts
Run
Run
#1
API response time
1.9 min
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 86% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI92.9–2None
2GPT-6.1 SolwithAPI88.7–2None
3GPT-6 LunawithAPI84.5–2None
4Opus 5.5withClaude76.8–2None
5GPT-6 AstrawithChatGPT69.6–2None
6Gemini 3.5 Flash-LitewithGemini49.4–2None

About the task

The PM job

Getting a clickable prototype in front of users or stakeholders fast.

Why it matters

A prototype that looks right but whose core interaction is broken wastes a research session. This task rewards working behaviour over polish.

What good looks like

  • The core interaction works end to end
  • Respects the stated constraints and scope
  • Handles the empty and error states named in the brief
  • Says what it's testing and what's faked
  • Looks like something you'd share: clear hierarchy, none of the stock AI-template look

Deliberately not measured

  • Production code quality
  • Design to a requested style: the brief never asks for one, so this measures default taste
Capability tested

Turning a spec into a usable interactive artefact

The failure we’re looking for

An attractive shell with a broken or absent core interaction

Grading

Browser checks (interaction and design) and decision model, calibrated against a blind PM review

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.