Usable with a quick edit12 of 12 checks passed
The output is clear, evidence-based, and complete; a PM could use it with minor wording edits without re-checking facts or redoing reasoning.
Output
1. Lukas’s review summary and proposed rating
Proposed rating: Meets expectations, with a significant launch-readiness gap to address.
Lukas delivered strong results on two of his three annual goals:
- Carrier onboarding: Reduced onboarding time from 21 days to 9 days, exceeding the 10-day target.
- Live tracking adoption: Increased adoption from 40% to 57%, exceeding the 55% target.
- Dispatch v2: Launched in October, meeting the timing target, but was rolled back after six days. This was not a successful, sustained delivery and should not receive full credit simply because it launched on schedule.
Dispatch exposed a serious readiness failure: support was not briefed, and 340 support tickets arrived in the first week. That outcome warrants explicit negative feedback and corrective action. However, the record does not establish how responsibility was divided among Lukas, his deputy, and other launch stakeholders. Lukas was on approved leave the week before launch, and his deputy ran the final checks. His leave is not itself a performance issue; his planning and handoff before leave remain legitimate areas to assess.
Peer feedback is positive from three of four peers, particularly on onboarding and collaboration. The Head of Operations’ criticism of Dispatch communication is relevant and consistent with the launch failure. It should be addressed directly rather than generalized into an unsupported conclusion about Lukas’s overall capabilities.
Review summary: Lukas exceeded the onboarding and live-tracking goals and received broadly positive feedback on collaboration. Dispatch v2 was a material setback: the release was not adequately prepared for support and was rolled back. As a Senior PM, Lukas needs to demonstrate more reliable launch planning, explicit delegation, and cross-functional readiness. We should establish his specific responsibility for the failure and assess improvement through the relaunch.
“Meets expectations” is the most defensible overall rating on the current record. It recognizes substantial results without minimizing Dispatch. It is also consistent with Priya’s rating last year after a similar rollback and support failure. That precedent is not automatically determinative, but a lower rating here would require a documented, material distinction that we do not currently have.
2. PIP recommendation
Do not initiate a PIP now.
HR policy requires earlier documented feedback on the specific gap and an opportunity to improve. There is no written feedback this year about strategy or product sense. The only recorded conversation concerns launch readiness in November, after Dispatch. On the available record, that does not establish both the required prior documentation and a subsequent opportunity to improve.
The proposed PIP therefore lacks the required foundation. “Not strategic enough” and “lacks product sense” are also too broad to serve as actionable performance gaps without concrete examples and defined expectations.
Instead, put a documented, non-PIP development plan in place now:
- Clarify accountability: Review the Dispatch launch plan, handoff, readiness checks, and go/no-go decision with Lukas, his deputy, Operations, and Support. Identify individual and process failures.
- Set observable expectations: For the relaunch, require named owners for every readiness item, documented support briefing and training, escalation and rollback procedures, and explicit cross-functional approval before release.
- Check progress regularly: Hold biweekly check-ins with Lukas and record feedback, completed actions, and any remaining gaps.
- Assess the relaunch: Evaluate readiness execution and post-launch support outcomes against thresholds agreed with Support before launch.
This should be genuine coaching, not a PIP under another name. If Lukas subsequently fails to meet clearly documented expectations after a reasonable opportunity to improve, reassess with HR. Any later PIP must run for 60 days, with measurable success criteria and the support needed to meet them.
3. Note to Ines
Ines — I agree that Dispatch’s readiness failure deserves clear accountability. I recommend Meets expectations and no PIP at this stage.
The pre-read should distinguish established facts from judgments. Lukas exceeded two goals; Dispatch met its launch date but failed as a sustained release. We have evidence of a launch-readiness and communication problem, but not enough to substantiate “not strategic enough” or “lacks product sense.” We also need to clarify Lukas’s handoff responsibilities versus the final checks his deputy owned.
Before calibration, I’ll verify the launch ownership, the November feedback record, and whether there are material differences from Priya’s case. Unless that produces materially different evidence, we should revise the pre-read accordingly. HR’s PIP prerequisites are not demonstrated in the current record; the appropriate next step is specific documented feedback and a fair opportunity to improve.
Grades and run details
Decision model 100 · LLM judge 13 of 13 checks
Decision model checks
- passUses the supplied evidence correctly71%
- passAddresses the actual decision96%
- passRespects explicit constraints63%
- passIdentifies material uncertainty96%
- passAvoids unsupported claims73%
- passProduces the required deliverable95%
- passHolds the PIP to the policy and the record100%
- passTreats Lukas like Priya99%
- passClear and direct, with care81%
- passJudges outcomes, not activity98%
- passNames gaps you could see45%
- passWeighs the whole period99%
Run
- Run
- #1
- API response time
- 25 s
- Submitted
- 5 Oct 2026