Write a PRD
Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?
Every task is a job PMs actually do in a normal week, with a realistic brief and messy inputs. 9 are in the core suite. The rest are planned, and they’ll join only once the current ones prove they can tell good work from bad.
Deciding what to build and why
Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?
Can the model make a coherent choice grounded in the evidence, rather than list aspirations?
Can the model sequence bets against capacity and dependencies, and explain the order?
Can the model pick a segment, a message and a channel, and say why?
Learning from customers and markets
Can the model separate evidence, themes and hypotheses without inventing consensus?
Can the model write a guide that uncovers behaviour rather than opinions?
Can the model structure a stakeholder or customer research interview?
Can the model size and describe a market from supplied sources without fabricating figures?
Can the model find where competitors are actually weak rather than list features?
Shaping the experience
Can the model find the friction that matters most and prioritise the fixes?
Can the configuration produce a working, constraint-compliant prototype in one attempt?
Can the model lay out a flow's screens with the right information hierarchy?
Testing and reading results
Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?
Can the model design a test that could actually change the decision?
Stress-testing ideas
Can the model find the strongest reason an idea may fail, backed by evidence?
Can the model expand an idea's ambition while keeping it tethered to a real mechanism?
Running launches and reporting progress
Can the model turn messy project status into an honest update that leads with what the reader needs to know or decide?
Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?
Can the configuration hold a consistent product voice across a multi-turn agent session?
Can the configuration complete agreed scope without silently expanding or changing it?