The task library

Every task is a job PMs actually do in a normal week, with a realistic brief and messy inputs. 9 are in the core suite. The rest are planned, and they’ll join only once the current ones prove they can tell good work from bad.

Define

Deciding what to build and why

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

2 cases0 models · 0 setupsDifficulty Updated 12 Sept 2026Measures the model
Not yet testedMaking product behaviour, uncertainty and eval requirements executable

Develop product strategy

Can the model make a coherent choice grounded in the evidence, rather than list aspirations?

2 cases0 models · 0 setupsDifficulty Updated 12 Sept 2026Measures the model
Not yet testedChoosing where to play and what not to do, from supplied evidence

Build a roadmap

Can the model sequence bets against capacity and dependencies, and explain the order?

PlannedTests: sequencing under constraints
Not yet testedSequencing under constraints

Write a GTM plan

Can the model pick a segment, a message and a channel, and say why?

PlannedTests: launch focus
Not yet testedLaunch focus

Discover

Learning from customers and markets

Extract discovery insights

Can the model separate evidence, themes and hypotheses without inventing consensus?

2 cases0 models · 0 setupsDifficulty Updated 12 Sept 2026Measures the model
Not yet testedFaithful synthesis of qualitative research

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

PlannedTests: question design
Not yet testedQuestion design

Interview discussion guide

Can the model structure a stakeholder or customer research interview?

PlannedTests: interview structure
Not yet testedInterview structure

Market research

Can the model size and describe a market from supplied sources without fabricating figures?

PlannedTests: source-faithful research
Not yet testedSource-faithful research

Competitor analysis

Can the model find where competitors are actually weak rather than list features?

PlannedTests: competitive insight
Not yet testedCompetitive insight

Design

Shaping the experience

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

2 cases0 models · 0 setupsDifficulty Updated 12 Sept 2026Measures the model
Not yet testedConsequential critique

One-shot prototype

Can the configuration produce a working, constraint-compliant prototype in one attempt?

2 cases0 models · 0 setupsDifficulty Updated 24 Sept 2026Measures the system
Not yet testedTurning a spec into a usable interactive artefact

Wireframe generation

Can the model lay out a flow's screens with the right information hierarchy?

PlannedTests: information hierarchy
Not yet testedInformation hierarchy

Experiment

Testing and reading results

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

2 cases0 models · 0 setupsDifficulty Updated 12 Sept 2026Measures the model
Not yet testedInterpreting results within their limits

Experiment specification

Can the model design a test that could actually change the decision?

PlannedTests: test design
Not yet testedTest design

Challenge

Stress-testing ideas

Challenge an idea

Can the model find the strongest reason an idea may fail, backed by evidence?

2 cases0 models · 0 setupsDifficulty Updated 12 Sept 2026Measures the system
Not yet testedEvidence-based critique

1000x an idea

Can the model expand an idea's ambition while keeping it tethered to a real mechanism?

PlannedTests: ambitious expansion
Not yet testedAmbitious expansion

Operate

Running launches and reporting progress

Write a stakeholder update

Can the model turn messy project status into an honest update that leads with what the reader needs to know or decide?

2 cases0 models · 0 setupsDifficulty Updated 24 Sept 2026Measures the model
Not yet testedHonest, decision-first status communication

Make the launch call

Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?

2 cases0 models · 0 setupsDifficulty Updated 24 Sept 2026Measures the model
Not yet testedDeciding against agreed launch criteria

Agent interaction & voice

Can the configuration hold a consistent product voice across a multi-turn agent session?

PlannedTests: voice consistency
Not yet testedVoice consistency

Stick to an agreed plan

Can the configuration complete agreed scope without silently expanding or changing it?

PlannedTests: scope discipline across a multi-step build
Not yet testedScope discipline across a multi-step build