Skip to content

Solo Founder Product Engineering Handbook

Kill Criteria Checklist

Write evidence-based stop, narrow, change, and pause rules that still hold when an experiment is almost convincing.

The Result That Is Almost Enough

Clear failure is easy to stop. The dangerous result is the one that contains a real success, misses the condition that justified the work, and offers several plausible reasons to continue. One customer paid. Three people returned. The demo worked after an hour of repair. A larger segment might respond. Another feature might remove the objection.

Every sentence may be true. Together they can keep a weak path alive indefinitely.

Kill criteria are decisions written while the founder can still afford an unwelcome result. They belong on the MVP Experiment Card before recruiting begins and beside the Experiment Results Memo when the observation window closes. Their job is not to make judgment automatic. Their job is to prevent the standard from moving after the evidence arrives.

Use this record for a product thesis, segment, channel, offer, feature, integration, pilot, MVP, or manual service. Judge one of those objects at a time. “Kill the product” is usually too vague to guide a test and too broad to be a responsible consequence.

Write a Rule That Can Survive the Result

A usable kill criterion names five things.

First, name the object under judgment: the weekly-summary workflow for small architecture studios, outbound email as the first channel, self-serve onboarding, or automated claim generation. If the rule is triggered, this is what loses permission to consume more effort.

Second, name the belief at risk. A rule about repeat use tests recurring value. A rule about founder review time tests whether the workflow can become a product rather than remaining a service. A rule about unsupported claims tests whether the promise can be made safely. Without the belief, a disappointing number does not reveal what should change.

Third, name observable evidence with its boundary. State the qualified segment, denominator, behavior, threshold, and observation window. “Users do not engage” is movable. “Fewer than three of eight qualified project leads send a summary to a real client by the end of their second weekly cycle” can be applied to the experiment that actually ran.

Fourth, name the consequence. Choose stop, narrow, change direction, or pause, and say what work that decision forbids. A criterion that merely requires “reassessment” leaves every implementation option open.

Finally, name the exception boundary, if one is justified. An unresolved question may earn one smaller test, but specify it now: which uncertainty, how many participants, how much founder time, which cutoff, and what result closes the path. “Try one more version” is not an exception boundary.

Safety and trust rules do not need to wait for a timebox. If continued exposure could mishandle customer data, produce an unreviewed consequential claim, or leave customers without a credible recovery path, write an immediate pause rule. Resuming then requires evidence that the named failure is contained, not renewed optimism about the product.

Find the Condition That Governs

Do not score the categories or require several weak criteria to fail before acting. One governing constraint can outweigh encouraging evidence elsewhere. Strong activation cannot compensate for a promise that is unsafe. Praise from users cannot supply buyer authority. Repeat use that requires bespoke founder rescue does not prove a repeatable product.

Write rules only for risks that could change the decision:

  • Customer and pain. What costly behavior will the right customer perform if the problem is urgent—share representative work, change a live workflow, introduce the buyer, or accept a trial? What refusal would weaken this specific segment or problem belief?
  • Value and return. What exact event shows the promised result reached the customer? When does the workflow naturally recur, and what behavior must happen without a reminder or rescue?
  • Buyer and payment. Which commitment is appropriate at this stage: a budget conversation, deposit, paid pilot, renewal, or expansion? Whose decision counts?
  • Reach. How many qualified prospects must the tested channel produce within a fixed cash and founder-time budget? A weak channel does not by itself disprove the customer problem.
  • Founder work. Which setup, delivery, review, correction, support, and exception load is allowed per value event? What amount would make the tested product model uneconomic or impossible for one person to operate?
  • Trust and failure. Which data, permission, accuracy, reliability, and recovery conditions must hold before another customer is exposed?
  • Opportunity cost. Which stronger, already observed path would displace this one? Name its evidence and the decision date. Vague awareness of other ideas is not a criterion.

Not every experiment needs a rule in every category. Choose the few conditions that can actually revoke or narrow the next commitment.

Copy the Record

KILL CRITERIA RECORD

DECISION BOUNDARY
Object under judgment:
Decision this experiment may authorize:
Work, money, customer exposure, and founder time at risk:
Decision date or completed observation window:

PRECOMMITTED RULES
Belief at risk:
Qualified segment and denominator:
Observable behavior or failure:
Threshold and time window:
Consequence: stop / narrow / change direction / pause
Work forbidden if the rule is triggered:

Repeat the six lines above for each governing risk.

IMMEDIATE PAUSE RULE
Trust, safety, or recovery event that stops exposure:
Evidence required before resuming:

ONE SMALLER TEST, IF EARNED
Uncertainty the current experiment may be unable to resolve:
Why the answer could change the decision:
Maximum participants, spend, founder time, and duration:
Result that closes the path:

AT THE BOUNDARY
What happened, with denominators and unfinished windows:
Which rule was met:
Smallest scope the evidence invalidates:
Decision and work now refused:
Customer, data, support, refund, or recovery obligations:
Date to archive, revisit, or begin the bounded next test:

If the one-smaller-test section cannot be bounded before the result, leave it blank. Its absence is useful. It prevents a post-result objection from automatically purchasing another cycle.

Apply the Rule at the Smallest Honest Scope

At the decision boundary, freeze the experiment record before explaining it. Resolve unfinished observation windows and separate qualified participants from people who never matched the segment. Then compare the observations with the written rules.

If a rule is triggered, identify the smallest claim it invalidates. A failed outbound sequence may kill a message or channel, not the pain. High founder review time may kill the self-serve product claim while preserving a valuable expert service. Failure in one segment may narrow the customer definition rather than erase the product thesis. The decision should be no broader—and no narrower—than the evidence.

Now ask whether the exception boundary was earned. A smaller test is justified only when one unresolved uncertainty could reverse the decision and the test does not require building the disputed product. It must also fit the cap written in advance. If those conditions do not hold, another iteration is a way of declining the result.

Carry out the consequence while the record is open. Remove the planned work, close recruiting, tell affected customers what will happen, preserve required data or refunds, and archive the evidence. A stop decision that leaves the roadmap, landing page, and next outreach batch intact has not released any founder attention.

A Mixed Result Is Not a Voting Contest

Return to the modeled architecture-studio experiment. The success rule required three studios to send a summary to a real client, begin the second weekly cycle without a reminder, and need no more than thirty minutes of founder review per summary by that cycle. The stop rule said to stop summary automation if fewer than two of eight qualified leads sent a summary after onboarding.

Four leads sent a summary. Three returned without a reminder. Only two of those three stayed within the founder-work boundary, and one draft contained an unsupported project claim caught during review.

Should the founder continue or kill the experiment?

Neither broad answer follows from the rules. The stop threshold was not met: four leads used the output in real client work. The success threshold was not met either: only two accounts combined return, real use, and acceptable review work. The evidence supports recurring value for a narrower workflow while refusing a broad self-serve or low-review claim. The unsupported claim also justifies pausing any unreviewed export path.

The next record can therefore narrow the allowed note formats and keep human review visible. It should also make the next consequence harder to evade:

Stop productized summary automation if fewer than three of six qualified studios complete the next weekly cycle without a reminder and within thirty minutes of founder review per summary. Pause exposure immediately if an unsupported consequential claim can reach export without being caught. Do not add formats, archive search, or automated sending during this test.

That rule does not pretend the product has failed. It names the exact path that will lose permission to continue, preserves the evidence already earned, and places a ceiling on the cost of resolving what remains unknown.

End the Experiment in the Work, Not Only on Paper

The record is complete when an unwelcome result has a visible consequence. The founder can say what stopped, what survived, which work is now refused, and whether one bounded uncertainty deserves another test.

Write the criteria before customer exposure. Apply them after the promised observation window, unless a trust or safety pause fires first. Then change the allocation of attention. The point of killing a path is not to declare failure; it is to make room for evidence to govern what happens next.