Tasks / Design

One-shot prototype

Can the configuration produce a working, constraint-compliant prototype in one attempt?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models.

Reliably right

  1. Avoids template layout and copy tells100% pass
  2. Shows the no-slots state100% pass
  3. Respects explicit constraints83% pass

Where it slips

  1. Says what it tests and what's faked46% pass
  2. Avoids the generic AI palette67% pass
  3. Stays in scope75% pass

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Build a clickable prototype that lets a patient move an existing appointment to another available slot in one flow. Deliver a single self-contained HTML file that opens straight into the flow in a browser: no design canvas, artboard or "play" step. Use made-up data.

ScenarioSam Patel has a GP appointment with Dr Amira Khan at Harbour Medical Centre on Tuesday 6 October at 09:30 and needs to move it. Show the next seven days of availability; at least one of those days is fully booked.
ConstraintsDesigned for a mobile screen, about 390px wide. No login, sign-up or account screens. Show a clear "no slots available" state for a fully booked day. The original appointment stays booked until the patient confirms, and the confirmation shows the new time.
What a strong answer does

A working mobile rebooking flow: pick a new slot, confirm, and see the new time confirmed; a fully booked day says so clearly; nothing beyond rebooking.

Critical failures (cap the score)
  • Core interaction broken
Case

v1.1 · synthetic · health, mobile

5 of 7 checks passed

Output

Open the prototypeThe file the model delivered, in a new tab

I've built the prototype as one clickable phone screen (390 × 844) on a Design canvas. I haven't clicked through it in a browser yet, so do a quick run-through before you show it to anyone.

Flow: appointment details → choose a new time → check and confirm → confirmation → Done, which goes back to the appointment. After a move, the appointment screen shows the new time and says "Moved from Tue 6 Oct, 09:30". There are no login or account screens.

How it meets the constraints:

Next seven days: a row showing Fri 25 Sep to Thu 1 Oct, each marked with how many slots are free, "Full" or "Closed". It opens on the first day with free slots, and times are split into morning and afternoon. Fully booked day: Mon 28 Sep shows "No slots available" and a button that jumps to the next day with free slots. Original stays booked: the choose and check screens both say the 6 Oct 09:30 appointment is still booked. It's only cancelled on "Confirm new time", and Continue stays disabled until a slot is picked. Confirmation: the new day and time are the largest text on the screen, with a line saying the old appointment was cancelled.

Choices I made that you may want to change:

"Next seven days" counts from today (24 Sep). That means every slot shown is earlier than the original 6 Oct appointment. Sunday 27 Sep shows as "Surgery closed", which looks different from the fully booked Monday. If you'd rather keep the example to a single empty state, I can make Sunday bookable instead. All slots are with Dr Khan only, and there's no option to pick another GP. The practice address and the "mobile ending 214" confirmation text are made up. It uses my own simple colours and fonts, not NHS branding, since no design system was available.

Check by check

Got wrong · 1
  • Avoids the generic AI palette
Mixed · 1
  • Core interaction worksThe two graders disagreed on this one.
Got right · 5
  • Respects explicit constraints
  • Shows the no-slots state
  • Stays in scope
  • Avoids template layout and copy tells
  • Says what it tests and what's faked

Grades and run details

Decision model 79
Decision model checks
  • passRespects explicit constraintsbrowser95%
  • partialCore interaction works26%
  • passShows the no-slots statebrowser100%
  • passStays in scope60%
  • failAvoids the generic AI palettebrowser100%
  • passAvoids template layout and copy tellsbrowser100%
  • passSays what it tests and what's faked81%
Artefacts
Run
Run
#1
Time to output
2.7 min
Submitted
24 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 86% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI92.9–2None
2GPT-6.1 SolwithAPI88.7–2None
3GPT-6 LunawithAPI84.5–2None
4Opus 5.5withClaude76.8–2None
5GPT-6 AstrawithChatGPT69.6–2None
6Gemini 3.5 Flash-LitewithGemini49.4–2None

About the task

The PM job

Getting a clickable prototype in front of users or stakeholders fast.

Why it matters

A prototype that looks right but whose core interaction is broken wastes a research session. This task rewards working behaviour over polish.

What good looks like

  • The core interaction works end to end
  • Respects the stated constraints and scope
  • Handles the empty and error states named in the brief
  • Says what it's testing and what's faked
  • Looks like something you'd share: clear hierarchy, none of the stock AI-template look

Deliberately not measured

  • Production code quality
  • Design to a requested style: the brief never asks for one, so this measures default taste
Capability tested

Turning a spec into a usable interactive artefact

The failure we’re looking for

An attractive shell with a broken or absent core interaction

Grading

Browser checks (interaction and design) and decision model, calibrated against a blind PM review

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.