Tasks / Design

One-shot prototype

Can the configuration produce a working, constraint-compliant prototype in one attempt?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models.

Reliably right

  1. Avoids template layout and copy tells100% pass
  2. Shows the no-slots state100% pass
  3. Respects explicit constraints86% pass

Where it slips

  1. Says what it tests and what's faked43% pass
  2. Avoids the generic AI palette71% pass
  3. Stays in scope75% pass

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Build a clickable prototype that lets a patient move an existing appointment to another available slot in one flow. Deliver a single self-contained HTML file that opens straight into the flow in a browser: no design canvas, artboard or "play" step. Use made-up data.

What the model was given2 items: Scenario, Constraints
ScenarioSam Patel has a GP appointment with Dr Amira Khan at Harbour Medical Centre on Tuesday 6 October at 09:30 and needs to move it. Show the next seven days of availability; at least one of those days is fully booked.
ConstraintsDesigned for a mobile screen, about 390px wide. No login, sign-up or account screens. Show a clear "no slots available" state for a fully booked day. The original appointment stays booked until the patient confirms, and the confirmation shows the new time.
What a strong answer doesThe answer key the graders mark against

A working mobile rebooking flow: pick a new slot, confirm, and see the new time confirmed; a fully booked day says so clearly; nothing beyond rebooking.

Critical failures (cap the score at 40)
  • Core interaction broken
Case

Mid-level · v1.1 · synthetic · health, mobile

The outputs

7 setups answered this brief. Pick up to three to read side by side.

6 of 7 checks passed

Output

Open the prototypeThe file the model delivered, in a new tab

Here is the complete, self-contained prototype in a single HTML file. It opens directly into the reschedule flow on mobile (or centered as a 390px mobile frame on desktop) with full interactivity, realistic states, a fully booked day condition, and a confirmation view.

Flow Notes

  • Starting State: Opens directly onto Sam Patel's reschedule interface showing Dr Amira Khan and the currently reserved slot (Tuesday 6 October at 09:30).
  • Safe Hold Guarantee: An explicit badge clarifies that the current Tuesday slot is safely preserved until a final confirmation occurs.
  • 7-Day Window & Fully Booked State:
  • Day 3 (Friday 9 October) and Day 5 (Sunday 11 October) demonstrate the explicit "fully booked / closed" state with guidance to quickly navigate to available days.
  • The other days provide realistic morning and afternoon appointment options.
  • Confirmation & Release: Selecting a slot enables the sticky CTA, opens a quick confirmation sheet showing both old and new times, and confirms the move with clear visual confirmation that the original 6 October slot has been released.

Grades and run details

Decision model 93
Decision model checks
  • passStays in scope94%
  • partialSays what it tests and what's faked34%
  • passAvoids the generic AI paletteby hand100%
  • passAvoids template layout and copy tellsby hand100%
  • passShows the no-slots stateby hand100%
  • passCore interaction worksby hand100%
  • passRespects explicit constraintsby hand100%
Artefacts
Run
Run
#1
API response time
79 s
Submitted
1 Oct 2026

Check by check

Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Mixed 1

Says what it tests and what's fakedMixed

No reason given.

Got right 6

Respects explicit constraintsRight

No reason given.

Core interaction worksRight

No reason given.

Shows the no-slots stateRight

No reason given.

Stays in scopeRight

No reason given.

Avoids the generic AI paletteRight

No reason given.

Avoids template layout and copy tellsRight

No reason given.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 85% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI92.9–2None
2GPT-6.1 SolwithAPI88.7–2None
3GPT-6 LunawithAPI84.5–2None
4Opus 5.5withClaude76.8–2None
5Gemini 3.8 FlashwithAPI75.6–2None
6GPT-6 AstrawithChatGPT69.6–2None
7Gemini 3.5 Flash-LitewithGemini49.4–2None

About the task

The PM job

Getting a clickable prototype in front of users or stakeholders fast.

Why it matters

A prototype that looks right but whose core interaction is broken wastes a research session. This task rewards working behaviour over polish.

What good looks like

  • The core interaction works end to end
  • Respects the stated constraints and scope
  • Handles the empty and error states named in the brief
  • Says what it's testing and what's faked
  • Looks like something you'd share: clear hierarchy, none of the stock AI-template look

Deliberately not measured

  • Production code quality
  • Design to a requested style: the brief never asks for one, so this measures default taste
Capability tested

Turning a spec into a usable interactive artefact

The failure we’re looking for

An attractive shell with a broken or absent core interaction

Grading

Browser checks (interaction and design) and decision model, calibrated against a blind PM review

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.