Solo Founder Product Engineering Handbook
Experiment Results Memo
Turn a closed experiment into an auditable decision without laundering weak signal into certainty.
Close the Experiment Before Explaining It
Experiment notes tend to improve in retrospect. The segment becomes more specific, the success threshold grows more forgiving, and the behavior that looked promising becomes the behavior the test was supposedly designed to measure. Nothing in the record needs to be false for the conclusion to become untrustworthy.
The results memo creates a boundary between the experiment that ran and the story a founder could tell about it. Write it at the declared cutoff or decision date, after every participant has reached the relevant observation window. If a return is not yet due, the experiment is not ready for a retention conclusion. Close the other questions and name that one as pending rather than turning an unfinished clock into a favorable result.
Begin from the document written before exposure: the MVP Experiment Card, test plan, pilot agreement, outreach batch, or launch contract. Copy its question, segment, method, cutoff, success and stop rules, intended decision, and stated limits. Do not rewrite them. When no prewritten rule exists, the memo should say so and limit its conclusion to what the observations can support.
Preserve the Experiment That Actually Ran
Record departures from the plan before interpreting the result. A new message, different price, repaired import path, founder reminder, expanded segment, or extended cutoff may be sensible operating work, but it changes the experiment. Identify who encountered each version. Do not pool unlike experiences merely to obtain a larger number.
Then reconstruct the evidence in the order a customer encountered the test. Keep the denominator attached to every count: qualified participants invited, people who entered, people who reached the promised value, people whose return was due, and people who paid or made another costly commitment. Preserve segment, source, version, and account-level records beneath the totals so that one credible pocket is not averaged together with wrong-fit attention.
Founder work belongs in the same record. Separate intended delivery from rescue. A concierge test may deliberately include expert review; a self-serve test cannot quietly include private data repair, repeated persuasion, or unrecorded reminders and still make the same claim. Include support time, severe exceptions, corrections, recovery, refunds, and trust failures even when the customer eventually succeeded.
Write observations before explanations. “Four project leads sent a summary to a client” is an observation. “The workflow solves a recurring reporting problem” is an interpretation. “Narrow the next test to studios with consistent meeting notes” is a decision. Keeping those sentences distinct lets a careful reader challenge the inference without disputing the record.
Make Confidence Belong to a Claim
Avoid giving the whole experiment one confidence label. The same test may provide strong evidence that a workflow matters, moderate evidence that one segment repeats it, and no evidence about acquisition or price. State each belief at the resolution the experiment earned.
Confidence rises when the observed behavior matches the original question, the participants match the intended segment, the relevant cycle has completed, the result survives counterexamples, and the path can be observed without large gaps. It falls when the founder changed several variables, selected only successful accounts, relied on stated enthusiasm, rescued the outcome, lost the denominator, or stopped before failure had time to appear. Small samples are often appropriate for an early decision, but they support a bounded next test, not a market-wide claim.
Name the strongest counterevidence beside the strongest supporting evidence. If the decision survives both, explain why. If it depends on ignoring one of them, the memo has found the next uncertainty rather than a conclusion.
Copy the Memo
Use only the fields that govern the decision, but do not omit an inconvenient one. Link to raw records instead of pasting a diary of the experiment into the memo.
EXPERIMENT RESULTS MEMO
DECISION AT HAND
Decision this experiment was meant to change:
Original question:
Target segment and qualification rule:
Method, exposure limit, cutoff, and observation window:
Success, change, and stop rules written before the result:
What this experiment could not prove:
WHAT ACTUALLY RAN
Dates, participants, source, offer, price, and product or message version:
Departures from plan and who encountered each one:
Missing records or unfinished observation windows:
OBSERVED EVIDENCE
Qualified participants invited / entered:
Activated, with the exact event and denominator:
Reached experienced value, with account references:
Return due / not yet due:
Returned without reminder / after reminder / did not return:
Paid, renewed, expanded, declined, refunded, or not asked:
Customer language attached to observed behavior:
BURDEN, FAILURE, AND TRUST
Intended founder work per participant or value event:
Explanation, setup, delivery, review, reminder, repair, and rescue work:
Typical load, peak load, and severe exceptions:
Product failures, recovery work, and trust concerns:
READING
Result against each prewritten rule:
Strongest evidence supporting the original belief:
Strongest counterevidence:
Claims this evidence supports, with confidence and reason:
Claims this evidence does not support:
Most plausible competing explanation:
What remains unknown:
DECISION
Choose one primary move: continue / narrow / change direction / harden / stop
Decision and the evidence that permits it:
What changes next:
What stays fixed so the next result remains interpretable:
Work explicitly refused or removed:
Next test, owner, exposure limit, pause rule, cutoff, and review date:
Customers owed follow-up, correction, recovery, refund, or a candid no:
Let a Mixed Result Stay Mixed
Consider a modeled closeout of the architecture-studio experiment from the MVP card. Eight qualified studios entered a two-cycle test. The prewritten success rule required three studios to send a summary to a real client, begin the second cycle without a reminder, and need no more than thirty minutes of founder review per summary by that cycle.
Six studios uploaded meeting notes. Five approved a draft, four sent it to a client, and three began the second cycle without a reminder. Those three required eighteen, twenty-four, and fifty-two minutes of founder review in the second cycle. The longest review caught a plausible but unsupported project claim. Two other studios never uploaded notes after asking how raw client records would be retained.
“Three studios returned without a reminder” is real evidence of recurring value for part of the target segment. It does not satisfy the written success rule: only two of those studios reached the value moment within the founder-work boundary. The third depended on review effort the proposed product was not meant to require. The two non-starters also prevent the memo from treating data trust as a minor onboarding detail.
The result supports a narrower belief: studios with consistent meeting-note formats can use the reviewed summary in recurring client work. It does not yet support self-serve use, acceptable review cost across the segment, or a safe raw-notes promise. The next move is therefore to narrow, not widen. Keep the segment, client-send value event, and two-cycle window fixed; admit two supported note formats, explain retention before upload, expose every claim that needs review, and test whether three studios can return within the work boundary. Refuse archive search, new summary formats, and automated sending until that uncertainty is resolved.
This is not a failed memo rescued by a hopeful interpretation, nor a successful memo spoiled by caveats. It is a mixed result that has earned one smaller, better-defined experiment.
End With a Consequence
A memo is complete when the next allocation of founder attention follows from the record. One primary move should be visible, along with what will stay fixed, what work will not enter the next cycle, and what would reverse the decision. If the evidence has met a prewritten stopping condition and no cheap test can resolve the remaining uncertainty, use the Kill Criteria Checklist rather than inventing another iteration.
Archive the memo with the original plan and account-level evidence. A later reader should be able to recover what was known at the time, disagree with the interpretation, and still see why the decision was reasonable. That is the memo’s standard: not certainty, but an honest consequence from bounded evidence.
Continue reading
Full table of contents