Tasks / Design

One-shot prototype

Can the configuration produce a working, constraint-compliant prototype in one attempt?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models.

Reliably right

  1. Avoids template layout and copy tells100% pass
  2. Shows the no-slots state100% pass
  3. Respects explicit constraints83% pass

Where it slips

  1. Says what it tests and what's faked46% pass
  2. Avoids the generic AI palette67% pass
  3. Stays in scope75% pass

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Build a clickable prototype that lets a patient move an existing appointment to another available slot in one flow. Deliver a single self-contained HTML file that opens straight into the flow in a browser: no design canvas, artboard or "play" step. Use made-up data.

ScenarioSam Patel has a GP appointment with Dr Amira Khan at Harbour Medical Centre on Tuesday 6 October at 09:30 and needs to move it. Show the next seven days of availability; at least one of those days is fully booked.
ConstraintsDesigned for a mobile screen, about 390px wide. No login, sign-up or account screens. Show a clear "no slots available" state for a fully booked day. The original appointment stays booked until the patient confirms, and the confirmation shows the new time.
What a strong answer does

A working mobile rebooking flow: pick a new slot, confirm, and see the new time confirmed; a fully booked day says so clearly; nothing beyond rebooking.

Critical failures (cap the score)
  • Core interaction broken
Case

v1.1 · synthetic · health, mobile

3 of 7 checks passed

Output

Open the prototypeThe file the model delivered, in a new tab

I will summarize what was built:

Direct Entry Flow: The prototype boots immediately into the rescheduling flow with patient context (Sam Patel) and a prominent status badge reminding the patient that the original appointment (Tuesday 6 October at 09:30 with Dr Amira Khan) is safely reserved until the change is confirmed.

7-Day Interactive Horizon: Built a horizontally scrollable chip selector covering the 7 calendar days starting October 7th.

Explicit Fully Booked State: Thursday 8 October (and Sunday 11 October) have clear "No appointments available" feedback with a single-tap jump to the next available slot.1

Structured Slot Selection: Categorized into Morning and Afternoon sessions with immediate highlight feedback and live drawer label updates.

Confirmation & Safeguards: Selecting a slot activates the CTA button (Confirm Change to [Date, Time]). Upon confirmation, it delivers a high-fidelity confirmation view highlighting the new appointment time, location details (Room 4 / Reception A), booking reference (HMC-7928-PAT), calendar export action, and a disclaimer noting the release of the previous Tuesday slot.

What a PM had to fix

  1. 1Artefact unusableFix the artefactQuick edit

    Source: ConstraintsShow a clear "no slots available" state for a fully booked day.

    What we’d changeFix the jump from a fully booked day: it suggests the wrong next available date. Everything else is ready to test.

Check by check

Got wrong · 2
  • Stays in scope
  • Says what it tests and what's faked
Mixed · 2
  • Respects explicit constraintsThe two graders disagreed on this one.
  • Avoids the generic AI paletteThe two graders disagreed on this one.
Got right · 3
  • Core interaction works
  • Shows the no-slots state
  • Avoids template layout and copy tells

Grades and run details

Decision model 57
Decision model checks
  • partialRespects explicit constraints29%
  • passCore interaction works35%
  • passShows the no-slots stateby hand100%
  • failStays in scope30%
  • partialAvoids the generic AI palette22%
  • passAvoids template layout and copy tellsby hand100%
  • failSays what it tests and what's faked52%
Artefacts
Run
Run
#1
Time to output
46 s
Submitted
24 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 86% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI92.9–2None
2GPT-6.1 SolwithAPI88.7–2None
3GPT-6 LunawithAPI84.5–2None
4Opus 5.5withClaude76.8–2None
5GPT-6 AstrawithChatGPT69.6–2None
6Gemini 3.5 Flash-LitewithGemini49.4–2None

About the task

The PM job

Getting a clickable prototype in front of users or stakeholders fast.

Why it matters

A prototype that looks right but whose core interaction is broken wastes a research session. This task rewards working behaviour over polish.

What good looks like

  • The core interaction works end to end
  • Respects the stated constraints and scope
  • Handles the empty and error states named in the brief
  • Says what it's testing and what's faked
  • Looks like something you'd share: clear hierarchy, none of the stock AI-template look

Deliberately not measured

  • Production code quality
  • Design to a requested style: the brief never asks for one, so this measures default taste
Capability tested

Turning a spec into a usable interactive artefact

The failure we’re looking for

An attractive shell with a broken or absent core interaction

Grading

Browser checks (interaction and design) and decision model, calibrated against a blind PM review

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.