Validation
Put the prototype in front of real users and gather evidence strong enough to change a decision.
Purpose
Validation is where the prototype earns the right to be handed off — or the right to be killed. The point is evidence strong enough to change a decision. A demo isn't validation. Positive vibes aren't validation. A sponsor saying "this is great" isn't validation.
- Entry
- Prototype exists with a defined success signal.
- Artifact
- Evidence pack: what we tested, who we tested with, what we saw, what it means.
- Gate
- Validated bet, iterate once more, or graceful stop.
What we're actually testing
Most Accelerator work is greenfield — new capability, new audience, or a workflow that doesn't exist yet. That means you're not A/B testing a button color. You're testing whether the idea holds up when a real person meets it. Four different questions, run in this order:
Activities
- Write the success signal before the test. One sentence, behavioral, measurable. Example: "4 of 6 target users complete the intake flow without prompting and 3 of 6 say they'd use it this week." If you write it after, you're rationalizing.
- Pick the smallest audience that answers the question. 5–8 real users is enough to see the pattern. Not executives. Not your partners. People who have the problem today.
- Run two testing modes in parallel. Moderated 1:1 sessions (3–5 people, 30 min, screen share) to hear the why. Unmoderated tests in Maze (5–10 people, tasks + follow-up questions) to see the what at scale.
- Instrument the prototype passively. Drop Microsoft Clarity into the Lovable prototype. Every markup.io view becomes a session recording. You'll see confusion the interview never surfaces.
- Capture evidence, not opinions. What did they do, not what did they say. Where did they get stuck. What did they try that the prototype didn't support. Direct quotes beat paraphrases — record with consent.
- Iterate once, at most twice. If two rounds don't move the signal, that is the evidence. Graceful stop, not a longer engagement.
- Assemble the evidence pack. Success signal (met / missed / mixed), what the data showed, three representative clips or quotes, the recommendation, and the assumption that would flip it.
Recommended session shape
Questions worth asking (and ones to avoid)
- "Walk me through the last time you dealt with [problem]."
- "What were you expecting to happen when you clicked that?"
- "If this disappeared tomorrow, what would you do instead?"
- "Who else on your team would need this?"
- "Do you like it?" — everyone says yes to your face.
- "Would you use this?" — hypothetical, not behavioral.
- "Would you pay for it?" — cheap talk unless money moves.
- Any question that starts with "Wouldn't it be cool if…"
Common failure modes
Prompt library
Paste these into a Lovable prototype chat to generate a first draft of research questions, tasks, or a screener. Fill in the bracketed context first, then edit the output down — these are starting points, not finished scripts.
You are helping an Innovation Accelerator design a 30-minute moderated interview to test the desirability of an early prototype. Context: - Problem we think we're solving: [problem] - Target user (role, context, current workflow): [user] - What the prototype does today: [prototype summary] - The riskiest desirability assumption: [assumption] Produce: 1. Three behavioral warm-up questions about how they handle [problem] today — no hypotheticals. 2. Five open questions that surface pain, workarounds, frequency, and cost of the current approach. 3. Three "reaction" questions to ask AFTER they've used the prototype, focused on what they'd change in their workflow, not whether they "liked" it. 4. Two commitment probes (would they pilot it, share it, give up something they use today). Avoid: leading questions, "would you use this", "wouldn't it be cool if", and anything that asks the user to predict their own future behavior. Flag any question that violates this and rewrite it.
Help me design an unmoderated usability test in Maze for the prototype below. Target 8-12 participants, ~15 minutes total.
Prototype: [link or short description of the flows available]
Primary user goal: [goal]
Success criteria I care about: [e.g. completes checkout without help, finds the report in <60s]
Produce:
1. A 2-sentence intro to show participants (no product marketing, no leading language).
2. 4-6 task prompts written as user goals ("You need to…"), not instructions ("Click the blue button"). Each task should map to a success criterion.
3. For each task: one post-task follow-up (open text) and one 1-5 confidence rating.
4. Three post-test questions: one about what was confusing, one about what was missing, one behavioral (what would they do next in real life).
5. A screener with 3 questions to filter for [audience].
Call out any task where success is ambiguous or where the prototype likely can't support the path.You are a pragmatic staff engineer reviewing an early prototype for feasibility risk before we invest in a production build. Prototype summary: [what it does] Data it touches: [systems, PII, volumes] External dependencies: [APIs, models, vendors] Constraints we already know: [SLA, latency, compliance] Produce: 1. Top 5 feasibility risks, each with: risk, why it matters, cheapest experiment to de-risk it in <2 weeks. 2. Any assumption in the prototype that would break at 10x or 100x today's volume. 3. Data / integration questions we should answer BEFORE handoff, phrased so a delivery team could actually answer them. 4. A short "won't work as-is" list — parts of the prototype that fake something the real system can't do that way. Be blunt. Do not hedge. If something is fine, say fine.
Help me pressure-test the viability of this opportunity before we recommend it move to delivery. Opportunity: [1-2 sentence description] Who benefits and how: [users, business unit] Rough cost to build & run (best guess): [cost] What we'd have to stop doing to fund it: [tradeoff] Produce: 1. Five questions I should be able to answer for a sponsor before asking for funding — organized as: value, cost, risk, timing, alternative. 2. The two most likely reasons a CFO or portfolio owner would kill this, and what evidence would defuse each. 3. A "smallest defensible pilot" — the narrowest slice we could run to prove the business case is real, with a rough success metric. 4. One question I'm probably avoiding because I don't want to hear the answer.
Supporting tools
- Maze for structured unmoderated prototype tests — tasks, success rates, follow-up questions.
- Microsoft Clarity dropped into every shared prototype for session replay and rage-click detection.
- markup.io for external users who need to comment without an account.
- NPS panel on Metrics to capture willingness signals in the moment.
