Usable with a quick edit9 of 12 checks passed
The unsupported claim is a minor wording issue that can be quickly edited to 'The features shipped didn't move the goal metrics' without rechecking evidence.
Output
Annual Review: Theo Brandt
Prepared by: Ana Ruiz, Group PM Overall rating: Meets expectations
Theo, I want this review to be direct, because you deserve a clear picture. I also need to own a mistake of mine, which I cover below.
Results against your goals
| Goal | Start | Target | Result |
|---|---|---|---|
| Teams using Timeline view weekly | 22% | 30% | 23% |
| Time for new teams to build first plan | 3.0 days | 1 day | 2.6 days |
Neither goal was met. Time to first plan improved by about 13%, which is real progress but a small fraction of the gap. Timeline usage was essentially flat.
Strengths
Delivery and team health. Your squad shipped 14 features, 12 on time. Your engineering lead says you keep the team unblocked, write clear specs, and never miss a sprint review. That reliability is hard to build, and the team trusts you.
Collaboration. Your designer called you "great to work with." Sales called you "very responsive." You treat partners as partners, and that is a real asset.
Responsiveness to stakeholders. You moved quickly on requests and kept people informed.
What to work on
1. Connecting what you ship to the outcomes you own. Your goals were about Timeline adoption and onboarding speed. Most of the year's output didn't aim at either. Six of the 14 shipped features are used by fewer than 2% of teams. Five of those six were Sales requests, and two of those were built for a single customer each. Responsiveness is a strength, but here it substituted for prioritization. Saying yes to five Sales asks used capacity that could have gone toward your goals.
2. Validating before building. Your designer said, "I wish we'd tested more before building; we shipped three things I'm not sure anyone uses." I agree. Moving fast on unvalidated work produces output, not impact. Fewer, better-tested bets would serve you and the team better.
3. Using data to drive decisions. Our 1:1s focused on delivery and stakeholder updates, not on whether your metrics were moving. Part of your job is to look at usage numbers regularly and change course when they don't move. I'd like you to make that a habit.
Where I fell short
In your Q3 1:1 you asked whether you were on track for your goals. I told you yes, citing great velocity. That wasn't accurate. Velocity was great, but the goal metrics weren't moving, and I never raised it. You asked the right question and got a misleading answer. I should have flagged this months earlier, when you had time to change course. Some of this outcome is on me, and it's a main reason I'm rating you Meets rather than Below expectations.
Rating rationale
Meets expectations reflects strong execution and collaboration, weighed against two missed outcome goals and a significant share of shipped work with minimal adoption. Without the missing mid-year feedback, I'd have had a harder conversation with you about the lower end of this rating. I don't see evidence for Exceeds, and I think the evidence supports Meets rather than Below, given partial progress on onboarding and the feedback gap on my side. To be clear, this is a solid-but-not-yet-sufficient year on outcomes, and the next half needs to show a shift.
Goals for the next half
- Onboarding: Reduce time to first plan from 2.6 days to 1.8 days. Identify the two or three biggest drop-off points in the first-plan flow within the first month, and bring me a proposal for addressing them.
- Timeline adoption: Raise weekly Timeline usage from 23% to 27%. Start with research into why teams who try Timeline don't return, and share findings with me by the end of month one.
- Validation before build: Every feature above a small-effort threshold needs a written success metric and some evidence of demand (user tests, prototype feedback, or data) before engineering starts. Requests from a single customer need a case for broader value.
- Post-launch review: Check usage of each shipped feature 60 days after launch and report the results to me. If something is underperforming, tell me what you'll do about it.
- Sales partnership: Work with Sales on a lightweight way to evaluate requests so that "yes" is reserved for asks tied to our goals or to broad demand, and "not now" comes with a clear reason.
How we'll work together
We'll spend the first part of every 1:1 on your goal metrics, not just delivery. I'll tell you plainly where you stand each quarter, including when the news is uncomfortable. If you disagree with anything here, I want to hear it. I'd also like to hear what support would help most.
You've built a team that executes well and trusts you. The next step is pointing that capability at the metrics that matter. I'm confident you can do it.
Ana
Claims the judge couldn’t find in the brief
- Most of the year's output didn't aim at either.
Grades and run details
Decision model 92 · LLM judge 10 of 13 checks
Decision model checks
- passUses the supplied evidence correctly63%
- passAddresses the actual decision93%
- passRespects explicit constraints34%
- partialIdentifies material uncertainty17%
- partialAvoids unsupported claims32%
- passProduces the required deliverable96%
- passOwns the manager's part100%
- passSets next goals as outcomes, with support95%
- passClear and direct, with care100%
- passJudges outcomes, not activity98%
- passNames gaps you could see92%
- passWeighs the whole period93%
Run
- Run
- #1
- API response time
- 21 s
- Submitted
- 5 Oct 2026