Usable with a quick edit12 of 12 checks passed
The output is evidence-based, directly addresses the decision, and requires at most minor wording tweaks; no re-checking of facts or reasoning is needed.
Output
1. Lukas’s review summary and proposed rating
Proposed rating: Meets expectations, with a significant launch-readiness gap to address.
Lukas delivered strong results against two measurable goals. Carrier onboarding time fell from 21 days to 9 days, exceeding the 10-day target. Live-tracking adoption increased from 40% to 57%, exceeding the 55% target. Three of four peers gave positive feedback, particularly on his onboarding work and collaboration.
Dispatch v2 was a substantial delivery failure. Although it launched in October, it was rolled back after six days; the launch objective should therefore not be treated as successfully completed. Support had not been briefed, and 340 support tickets arrived in the first week. The Head of Operations’ feedback reinforces the specific concern about launch communication and operational readiness. Relaunch is planned for next quarter, so the intended outcome remains outstanding.
As the responsible Senior PM, Lukas should be assessed on how he established launch-readiness requirements, cross-functional ownership, and delegation. However, the record does not establish which safeguards he put in place or who approved the final launch. He was on approved leave the preceding week, and his deputy ran the final checks. Approved leave is not a performance failure; equally, delegation does not automatically remove accountability for the preparation and handoff. We should establish those facts before assigning Lukas sole responsibility.
The evidence supports a specific execution and communication gap, not the broader conclusions that Lukas is “not strategic enough” or “lacks product sense.” No feedback on those broader concerns was documented this year, and the supplied record contains no concrete examples substantiating them.
On balance, two above-target outcomes, positive collaboration feedback, and one serious but bounded delivery failure support Meets expectations. This is also consistent with Priya’s rating last year after a similar rollback and support failure. That precedent does not dictate Lukas’s rating, but a harsher outcome requires a material, evidenced distinction—not stronger language about the incident.
2. PIP recommendation
Do not initiate a PIP on the current record.
HR requires earlier documented feedback on the specific gap and an opportunity to improve. The record contains one November conversation about launch readiness, after Dispatch, but does not establish that it was documented or that Lukas subsequently had a meaningful opportunity to demonstrate improvement. There is no documented feedback on strategy or product sense. A PIP on those grounds would therefore lack the required foundation.
Instead, put a documented development plan within normal performance management in place now. This is not a disciplinary PIP or a predetermined path to one. Agree with Lukas on:
- Readiness ownership: Before the relaunch go/no-go review, document owners, decision rights, coverage during absences, escalation routes, and rollback criteria.
- Support preparedness: Before launch approval, secure Support and Operations sign-off on training, customer communications, FAQs, and the escalation runbook.
- Launch monitoring: Before relaunch, agree measurable adoption, reliability, and support-volume thresholds with Engineering, Support, and Operations; review results against them after launch.
- Manager support: Provide a readiness template, help resolve cross-functional ownership gaps, and hold regular check-ins through relaunch.
Document the expectations, support provided, and subsequent results. If a specific performance gap persists after clear feedback and a genuine opportunity to improve, reassess with HR. Any later PIP must meet the policy’s 60-day duration and measurable-success-criteria requirements.
3. Note to Ines
Ines—
I recommend changing Lukas’s proposed rating to Meets expectations and removing the PIP recommendation on the current evidence.
We should retain a direct account of Dispatch’s failure: Support was not briefed, 340 tickets arrived in the first week, and the release was rolled back after six days. But “not strategic enough,” “lacks product sense,” and “disaster” should be replaced with specific, attributable observations. The current draft also omits his two above-target results.
Before the pre-read is finalized, I’ll review the launch checklist, handoff and approval record, clarify the November feedback and subsequent improvement opportunity, and give Lukas a chance to respond to the factual account. I’ll also check whether there are material differences from Priya’s case.
If that review produces additional evidence, we should consider it explicitly. Otherwise, the defensible outcome is Meets expectations, clear launch-readiness feedback, and documented follow-through—not a PIP unsupported by our policy.
Grades and run details
Decision model 100 · LLM judge 13 of 13 checks
Decision model checks
- passUses the supplied evidence correctly71%
- passAddresses the actual decision96%
- passRespects explicit constraints65%
- passIdentifies material uncertainty99%
- passAvoids unsupported claims60%
- passProduces the required deliverable97%
- passHolds the PIP to the policy and the record100%
- passTreats Lukas like Priya97%
- passClear and direct, with care92%
- passJudges outcomes, not activity99%
- passNames gaps you could see94%
- passWeighs the whole period100%
Run
- Run
- #1
- Time to output
- 29 s
- Submitted
- 5 Oct 2026