Solo Founder Product Engineering Handbook / Chapter 29
Prototype to Answer One Question
Choose prototype fidelity from the assumption being tested instead of building something polished by default.
Preparing audio…
Audio edition
Prototype to Answer One Question
A Dashboard Would Answer the Wrong Question
A solo founder wants to help boutique recruiting agencies prepare Friday updates for their clients. Recruiter notes are scattered across an applicant-tracking system, email, and private documents. Account managers spend Friday morning reconstructing what changed: which candidates moved, which searches stalled, and which client needs a difficult conversation.
The product in the founder’s head is a polished dashboard. It imports notes, identifies stale candidates, drafts updates, assigns follow-ups, and gives every client a tidy status page. The founder can build it. That is part of the danger.
The dashboard would bundle several unanswered questions. Will agencies share the source data? Are Friday updates painful enough to change behavior? Can a useful digest be produced from inconsistent notes? Will an account manager trust it enough to send it? Will anyone ask for it a second time?
A prototype cannot answer all of those questions at once. It should not try. A prototype is a temporary instrument for reducing one uncertainty enough to make a decision.
The founder begins with the decision:
If an account manager uses a weekly client-status digest in real work and asks for it again, I will offer a paid manual pilot. If the digest is ignored or requires the manager to rewrite it, I will change the promise before building a dashboard.
That sentence does more work than a mockup. It names a consequence. It makes a disappointing result useful. It also reveals what must be prototyped: not the dashboard, but the experience of receiving, reviewing, and sending one trustworthy digest.
Let the Risk Choose the Artifact
Founders often begin with an artifact: a demo, a landing page, a clickable flow, a technical spike. Beginning there quietly assumes that the artifact can expose the risk. Usually the first task is to make the risk precise enough to choose the artifact at all.
“People want better reporting” is too loose to test. “Agency account managers will reuse a weekly client-change summary if it reduces Friday coordination without creating trust problems” names an audience, a situation, an outcome, and behavior that could support or weaken the belief.
Ask what could kill the direction fastest if it were false. The answer usually belongs to one of a few families:
- Value: Will someone change behavior, commit time or data, repeat the task, accept a pilot, or pay for the outcome?
- Meaning: Does the right person understand the promise, objects, language, and sequence?
- Workflow fit: Can the idea survive the real cadence, handoffs, approvals, tools, and exceptions around the work?
- Trust: Will the user review, rely on, correct, send, or act on the result?
- Feasibility: Can the essential promise be delivered with acceptable quality, latency, cost, permissions, and recovery at this stage?
- Reach: Can the founder get a specific audience to notice the promise and take a credible next step?
These are not labels to collect on a canvas. They prevent category mistakes. A clickable flow can expose confusing language, but it cannot show that an agency will pay. A technical spike can reveal whether notes can be parsed, but not whether the resulting digest deserves a place in Friday work. A landing page can test whether a promise attracts qualified interest, but signups do not demonstrate trust or repeated use.
The recruiting founder chooses value and trust as the primary risk. That choice rules out the satisfying dashboard and points toward a concierge test: deliver the outcome by hand, using real inputs, close enough to the account manager to see what earns or loses trust.
Build Only the Experience the Question Requires
The prototype has a simple front and a deliberately manual back.
One agency exports a week’s recruiter notes. The founder uses a spreadsheet to normalize candidate and role names, a prompt template to draft possible changes, and a review checklist to catch unsupported claims. The account manager receives a client-status digest in the same form they would plausibly use on Friday. They edit it, decide whether to send it, and explain any hesitation. Nothing is integrated into the agency’s systems yet.
This is enough realism for the chosen question. The manager works with genuine notes, faces the reputational consequence of a client update, and must decide whether the output saves effort. The founder can observe the edits, missing context, data-access friction, preparation time, and request—or lack of one—for next week’s digest.
It is also no more realistic than necessary. There is no authentication system, shared workspace, billing flow, general-purpose parser, or automated delivery. Those things would make the test look more like a product while adding little evidence about whether the digest belongs in the workflow.
Fidelity means the amount of reality needed for honest behavior. It does not mean visual finish.
When the question is meaning, a sketch may be enough to discover whether users think in clients, searches, roles, or projects. A landing page may be enough to test whether a narrowly worded promise earns a qualified reply. A clickable flow may be necessary when order, labels, setup, or the first value moment cannot be judged from a static page.
When the question is work, the prototype needs contact with the work. Spreadsheets are useful for exposing data shape, statuses, calculations, and exceptions. A manual backend can place a simple form or email in front of a user while the founder operates scripts and checklists behind it. A concierge prototype keeps that human work visible to both sides. A Wizard-of-Oz prototype hides some manual operation so the interaction can be observed as if it were immediate, but it must not make false claims about safety, compliance, human review, or capability. Customer data and consequential decisions still deserve honest handling even when the system is temporary.
When the question is capability, the user may not need an interface at all. An API mock can test a contract and its failure cases. A prompt test set can show whether an AI-assisted step survives representative, ambiguous, and poor-quality inputs under a defined review process. A technical spike can isolate file parsing, latency, matching, permissions, cost, or another feasibility constraint. A spike should end in a decision, not drift into production because the editor is already open.
In each case, the artifact earns its fidelity from a behavior the founder needs to observe. If the participant must imagine the important part, the prototype is too thin. If the founder is building parts that cannot affect the decision, it is too thick.
Decide What the Result Will Mean Before the Test
The founder wants the prototype to succeed. That makes an ambiguous response dangerous. “Interesting” can become demand; a polite follow-up can become urgency; one corrected digest can become proof of automation.
Before requesting the agency’s data, the founder writes a short decision note:
Decision
Offer a paid four-week manual pilot, change the promise, or stop this direction.
Primary assumption
Agency account managers will trust and reuse a weekly client-status digest
because it reduces Friday coordination.
Credible audience
Account managers who prepare recurring client updates from recruiter notes.
Prototype
One manually produced digest from real notes, reviewed in the existing Friday flow.
Support signal
The manager sends the digest with limited factual edits and requests another run.
Change signal
The digest helps, but the manager rewrites its structure or needs a different
source, audience, or cadence.
Stop signal
The agency will not provide the inputs, the output is not used, or the work
creates no reason to repeat or pay for the service.
Boundary
This test cannot prove automation quality, integration feasibility, margins,
retention beyond the test period, or demand across the market.
The numbers and thresholds should fit the decision. A single real use may justify another manual test; it does not justify a market claim. For a consequential build, the founder may require repeated use across several participants or a paid commitment. False precision is not discipline, but a result with no declared consequence is only observation.
The boundary is as important as the support signal. Every prototype lies by omission. A smooth clickable flow omits production errors. A concierge service omits software economics. A technical spike omits customer pull. A strong prompt result omits the distribution of future inputs unless the test set represents them. Naming the omission keeps the evidence in its proper place.
Watch the Work Without Rescuing It
On Thursday, the agency sends the export two hours later than promised. Several candidates have different names in different notes. One client update depends on context that exists only in an account manager’s memory. The prompt draft confidently joins two unrelated comments.
These are not annoyances around the prototype. They are some of its best findings. The product’s input is not “recruiter notes” in the abstract. It is late, inconsistent, incomplete notes plus human context. A dashboard mock would have hidden that.
The founder corrects the digest before delivery and records each intervention. On Friday, the account manager removes one speculative sentence, rewrites two candidate updates, and sends the rest. The digest saves time, but the manager says the most useful section is not the founder’s proposed task list. It is the short list of searches whose status cannot be explained from the notes.
The founder does not defend the task list or explain how a future version would improve it. The test is allowed to push back. The next questions are plain:
- What did you do with this in your real process?
- Which parts did you distrust or change, and why?
- What information was missing at the moment you needed it?
- Would you use another digest next Friday?
- If I ran this manually for a month, would you pay for the pilot?
The final ask must match the assumption. A value test asks for real data, repetition, workflow access, or money. A comprehension test asks the user to explain or complete something without coaching. A feasibility test records a technical result and the product choices it permits or blocks.
Observe what happens, not only what is said. Refusals, delays, corrections, skipped steps, workarounds, forwarded outputs, and requests for another run often carry more weight than praise. Do not rescue every confusion. The moment the founder teaches the participant how to pass the test, the evidence becomes harder to interpret.
Let Ambiguity Narrow the Next Question
Suppose the account manager asks for another digest but refuses a paid pilot. That is neither clean success nor failure. The prototype has still done useful work if the founder resists averaging the signals into a hopeful yes.
Perhaps the digest saves time but the buyer is an agency owner rather than an account manager. Perhaps Friday is the wrong moment to charge because the workflow belongs inside a larger service. Perhaps the missing-status list is valuable while the drafted client update is not. Each interpretation leads to a smaller next test.
The wrong response is to add polish. Ambiguity about value is not repaired by a nicer review screen. Ambiguity about the buyer is not repaired by an integration. Change the audience, promise, output, or commitment being requested, then run the cheapest test that separates the remaining explanations.
The same restraint applies to a positive result. If the agency accepts a paid manual pilot, the founder has earned a paid manual pilot—not a complete platform. The next prototype might be a prompt test set for messy notes, an API mock for one export path, or a clickable review screen. Each should inherit one uncertainty from what the previous test revealed.
Stop When the Prototype Has Answered
Prototypes become products accidentally. A spreadsheet acquires permanent customers. A technical spike gains a background job. A manual process becomes an obligation no one has priced. The founder keeps it alive because it works, even though its original question has been answered.
At the end of a test, write one sentence:
Because of this evidence, I will…
If the sentence does not change what happens next, the prototype was probably detached from a decision. If it does change the work, act on it: offer the pilot, narrow the promise, choose the next uncertainty, or stop.
A prototype is worth making when one consequential belief is exposed to credible behavior, the artifact is only as real as that behavior requires, the signals and limits are declared beforehand, and the result changes a decision. When any of those conditions is missing, improve the question before improving the prototype.
Field Exercise
Choose one feature, workflow, integration, or AI capability you are tempted to build. Write the decision you expect it to support, the belief that could make the decision wrong, whose behavior would count, and the lowest-fidelity experience that could reveal that behavior. Add one support signal, one change signal, one stop signal, and one sentence about what the test cannot prove.
Run the test before touching production code. Then complete: “Because of this evidence, I will…”
The dashboard can wait. The evidence has earned only the next experiment.
Continue reading
Full table of contents