Your plan looks solid. So did the last one that shipped broken.
Plan Court seats independent AI critics around your implementation plan — each hunting through a different lens — while a judge folds the fixes in, round after round, until the table converges. Then it rules: PASS, CONDITIONAL, or FAIL.
runs cost $0.35–$2 · plans are never used for training · open-source core, zero dependencies
Almost every expensive bug starts as a line in a plan nobody challenged.
The plan is written by the person who believes it — then executed, unchanged, by someone (or something) who doesn't question it.
No hostile readers
Plans are advocacy documents, written by whoever is most convinced they're right. Nobody is paid to attack one before work starts — so the blocker nobody imagined ships on schedule.
— nobody, before todayAgents execute bad plans faithfully
Claude Code and Codex don't ask whether your migration is reversible. A missing rollback window isn't a question to them — it's a to-do. They will implement your plan perfectly, straight into an incident.
— the agent, faithfullyDiscovery comes at the worst moment
The flaws surface mid-deploy, mid-migration, mid-incident — the most expensive possible times to learn them. Five minutes of adversarial review before the first line of code is the cheapest QA you will ever run.
— 2 a.m., mid-deployThe defense never rests — because there is no prosecution. Plan Court is the prosecution.
A literal roundtable. Five critics, one judge, one defendant: your plan.
Seat independent critics
Two to eight models — your keys or ours — each assigned an adversarial lens. Lenses rotate every round so no seat repeats itself:
The presumption of brokenness
Critics are instructed to assume the plan is broken until proven otherwise. Praise is forbidden; vague hand-waving counts as silence. Maximum five issues each — specificity or nothing.
The judge folds the fixes in
Every critique gets a disposition: addressed (folded into the plan), deferred (valid but out of scope — on the record), or rejected (wrong, with a reason). Untouched sections are copied verbatim. The judge may not invent structure the plan doesn't name.
Converge, then rule
Rounds repeat — re-review, revise — until the table passes unanimously or rounds run out. In the final ceremony each surviving critique is UPHeld or OVERRULEd with a one-sentence reason, and the verdict is stamped into the record.
The table found this in a plan that was about to ship.
This is a real critique, quoted verbatim from a live Plan Court session over a typical migration plan — the kind of plan that looks fine in a code review.
Active Stripe subscriptions are never cancelled, causing double billing
"Phase 3 flips renewals to the ledger cron using PaymentIntents, but nowhere in the plan are the existing Stripe Subscription objects paused or cancelled. Because Stripe's subscription engine runs autonomously on Stripe's infrastructure, both Stripe Subscriptions and the new internal cron will charge customers concurrently."
Judge's disposition — addressed: "Stripe Subscriptions are canceled/paused as each cohort moves ledger-owned" — plus an exactly-one-billing-engine-per-subscription invariant, verified before every cutover step.
And in round two, a critic caught a flaw in the judge's own round-one fix — "Rollback protocol relies on reactivating a canceled Stripe Subscription, which Stripe does not support." Red Team · DeepSeek V4 Flash · Round 2 — the judge amended the rollback again: pause_collection, a documented recreation flow, and overlap reconciliation
Three ways the trial can end. All of them are progress.
No upheld blockers.
The critics tried and failed. Ship the plan as written — the report says so, on the record.
Ship with amendments.
The judge's exact changes are folded into a finalized plan your coding agent can execute directly. The report lists every amendment and why.
Rework before code.
The upheld blockers are named, severities assigned, fixes suggested. Fix them, then convene again. Cheaper than learning them mid-deploy.
- every critique, with severity and a suggested fix
- every judge disposition — addressed, deferred, rejected — with reasons
- every final ruling: UPHOLD or OVERRULE, on the record
- the finalized amended plan, ready to hand to an agent
- honest accounting: tokens in / out, estimated cost per seat
- the whole debate, streamable live while it happens
One credit ≈ one plan hardened on the standard model mix.
You see the estimated cost before every run — and the true token cost after it. No surprises, ever. Teams bring their own keys and run unlimited at raw model cost.
| Free | Pro | Team | |
|---|---|---|---|
| Price | $0forever | $19per month | $149per month |
| Credits | 2 trial credits | 15 credits / month — never expire | Bring your own API keys — unlimited runs |
| Critics | up to 5 seats | up to 6 seats, all lenses | up to 8 seats, priority models |
| Members | 1 | 1 | 5 members, shared presets |
| History | 7 days | 90 days, PDF export | 90 days, PDF export |
| API & CI | — | CLI + exit codes | API tokens, CI gates, webhooks |
Launch pricing — waitlist members lock in 20% off Pro for the first year. Managed runs are credit-capped and metered honestly; BYOK runs cost exactly what the providers charge. Credit packs ($25 / $100 / $500) for pay-as-you-go.
Take your plan to the table.
Sign in with a one-time link, get two free trial credits, and run your first adversarial review in under a minute. Upgrade to Pro or Team when you're ready.
Already a member? Sign in →