One-shot prototype
Can the configuration produce a working, constraint-compliant prototype in one attempt?
The PM job
Getting a clickable prototype in front of users or stakeholders fast.
Why it matters
A prototype that looks right but whose core interaction is broken wastes a research session. This task rewards working behaviour over polish.
What good looks like
- The core interaction works end to end
- Respects the stated constraints and scope
- Handles the empty and error states named in the brief
- Looks like something you'd share: clear hierarchy, none of the stock AI-template look
Deliberately not measured
- Production code quality
- Design to a requested style: the brief never asks for one, so this measures default taste
Turning a spec into a usable interactive artefact
An attractive shell with a broken or absent core interaction
Browser checks (interaction and design), decision model and blind PM review
This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.
Results
Every setup we’ve tested on this task, across all cases and repeats.
| # | Model · Harness | Task score | Decision model | LLM judge | PM review | Runs | Critical failures | Cost / run | Latency |
|---|
Case viewer
Read the brief, then put up to three outputs side by side. The outputs are the point; the scores just tell you where to look.
Build a clickable prototype of a flow where an employee captures a receipt, checks the amount the app extracted, corrects it, and saves the expense. Deliver a single self-contained HTML file that opens straight into the flow in a browser: no design canvas, artboard or "play" step. Use made-up data.
Capture (simulated), a visible extracted amount, an easy correction, and a saved expense showing the corrected amount.
- Core interaction broken
v1.0 · synthetic · fintech