Innovation Accelerator Hub
All playbooks
Operating model stages

Validation

Put the prototype in front of real users and gather evidence strong enough to change a decision.

Published
Owner · Director, InnovationUpdated · 2026-07-19

Purpose

Validation is where the prototype earns the right to be handed off — or the right to be killed. The point is evidence strong enough to change a decision. A demo isn't validation. Positive vibes aren't validation. A sponsor saying "this is great" isn't validation.

Entry
Prototype exists with a defined success signal.
Artifact
Evidence pack: what we tested, who we tested with, what we saw, what it means.
Gate
Validated bet, iterate once more, or graceful stop.

What we're actually testing

Most Accelerator work is greenfield — new capability, new audience, or a workflow that doesn't exist yet. That means you're not A/B testing a button color. You're testing whether the idea holds up when a real person meets it. Four different questions, run in this order:

Test first — cheapest to kill
Desirability
Do users actually want this? Would they stop what they're doing today to use it? If no one leans in, nothing downstream matters.
Test second
Usability
Can they get through it without you narrating? Where do they hesitate, misclick, ask 'what does this do?' The prototype should teach itself.
Test with engineering partners
Feasibility
Can we actually build this at enterprise scale? Rolls into Enterprise Readiness — flag showstoppers now, don't wait.
Test with the sponsor
Viability
Does the business case survive the evidence? If usage is 30% of what the ROI assumed, the number moves — don't hide it.
Greenfield vs iterative — pick the right bar
Greenfield work (new capability, new audience) validates on directional signal: did 4 of 6 users complete the core task, did 3 of 5 say "when can I have this." Iterative work on an existing product needs quantitative lift over the current experience. Don't hold a greenfield bet to an iterative bar — you'll kill things that deserved another round.

Activities

  • Write the success signal before the test. One sentence, behavioral, measurable. Example: "4 of 6 target users complete the intake flow without prompting and 3 of 6 say they'd use it this week." If you write it after, you're rationalizing.
  • Pick the smallest audience that answers the question. 5–8 real users is enough to see the pattern. Not executives. Not your partners. People who have the problem today.
  • Run two testing modes in parallel. Moderated 1:1 sessions (3–5 people, 30 min, screen share) to hear the why. Unmoderated tests in Maze (5–10 people, tasks + follow-up questions) to see the what at scale.
  • Instrument the prototype passively. Drop Microsoft Clarity into the Lovable prototype. Every markup.io view becomes a session recording. You'll see confusion the interview never surfaces.
  • Capture evidence, not opinions. What did they do, not what did they say. Where did they get stuck. What did they try that the prototype didn't support. Direct quotes beat paraphrases — record with consent.
  • Iterate once, at most twice. If two rounds don't move the signal, that is the evidence. Graceful stop, not a longer engagement.
  • Assemble the evidence pack. Success signal (met / missed / mixed), what the data showed, three representative clips or quotes, the recommendation, and the assumption that would flip it.

Recommended session shape

3–5 users · 30 min
Moderated 1:1
Give one real task. Stay silent. Ask 'what are you thinking?' when they pause. End with: 'What would you do next if I wasn't here?'
5–10 users · async
Unmoderated (Maze)
2–3 tasks, one open question per task, one closing 'would you use this weekly / monthly / never'. Read the misclick heatmap before the responses.
Everyone who touches it
Passive (Clarity)
Session recordings + rage-click detection. Watch three at 2x, look for the same hesitation across users — that's your next iteration.

Questions worth asking (and ones to avoid)

Ask
  • "Walk me through the last time you dealt with [problem]."
  • "What were you expecting to happen when you clicked that?"
  • "If this disappeared tomorrow, what would you do instead?"
  • "Who else on your team would need this?"
Avoid
  • "Do you like it?" — everyone says yes to your face.
  • "Would you use this?" — hypothetical, not behavioral.
  • "Would you pay for it?" — cheap talk unless money moves.
  • Any question that starts with "Wouldn't it be cool if…"

Common failure modes

Validating with the wrong audience
Sponsors and executives are not validation. They're funders. Their reaction tells you whether the story lands, not whether the solution works. If real users haven't touched it, you don't have evidence yet.
Confusing enthusiasm with commitment
"This is amazing" from someone who won't change their workflow next Tuesday is noise. Weight behavioral signal (they used it, they shared it, they asked when they could have it) far above verbal enthusiasm.
Iterating past the signal
If two rounds haven't moved the metric, more rounds won't. The graceful stop is the deliverable — a documented "we tested this, here's what we learned, here's why we're not moving forward" is worth more than a polished prototype no one asked for.

Prompt library

Paste these into a Lovable prototype chat to generate a first draft of research questions, tasks, or a screener. Fill in the bracketed context first, then edit the output down — these are starting points, not finished scripts.

Desirability
Draft desirability interview questions
Use in: ChatGPT / Ramble Refiner
You are helping an Innovation Accelerator design a 30-minute moderated interview to test the desirability of an early prototype.

Context:
- Problem we think we're solving: [problem]
- Target user (role, context, current workflow): [user]
- What the prototype does today: [prototype summary]
- The riskiest desirability assumption: [assumption]

Produce:
1. Three behavioral warm-up questions about how they handle [problem] today — no hypotheticals.
2. Five open questions that surface pain, workarounds, frequency, and cost of the current approach.
3. Three "reaction" questions to ask AFTER they've used the prototype, focused on what they'd change in their workflow, not whether they "liked" it.
4. Two commitment probes (would they pilot it, share it, give up something they use today).

Avoid: leading questions, "would you use this", "wouldn't it be cool if", and anything that asks the user to predict their own future behavior. Flag any question that violates this and rewrite it.
Usability
Draft an unmoderated Maze test plan
Use in: ChatGPT, then paste tasks into Maze
Help me design an unmoderated usability test in Maze for the prototype below. Target 8-12 participants, ~15 minutes total.

Prototype: [link or short description of the flows available]
Primary user goal: [goal]
Success criteria I care about: [e.g. completes checkout without help, finds the report in <60s]

Produce:
1. A 2-sentence intro to show participants (no product marketing, no leading language).
2. 4-6 task prompts written as user goals ("You need to…"), not instructions ("Click the blue button"). Each task should map to a success criterion.
3. For each task: one post-task follow-up (open text) and one 1-5 confidence rating.
4. Three post-test questions: one about what was confusing, one about what was missing, one behavioral (what would they do next in real life).
5. A screener with 3 questions to filter for [audience].

Call out any task where success is ambiguous or where the prototype likely can't support the path.
Feasibility
Surface feasibility risks with engineering
Use in: ChatGPT / Lovable chat
You are a pragmatic staff engineer reviewing an early prototype for feasibility risk before we invest in a production build.

Prototype summary: [what it does]
Data it touches: [systems, PII, volumes]
External dependencies: [APIs, models, vendors]
Constraints we already know: [SLA, latency, compliance]

Produce:
1. Top 5 feasibility risks, each with: risk, why it matters, cheapest experiment to de-risk it in <2 weeks.
2. Any assumption in the prototype that would break at 10x or 100x today's volume.
3. Data / integration questions we should answer BEFORE handoff, phrased so a delivery team could actually answer them.
4. A short "won't work as-is" list — parts of the prototype that fake something the real system can't do that way.

Be blunt. Do not hedge. If something is fine, say fine.
Viability
Draft a viability / business-case probe
Use in: ChatGPT
Help me pressure-test the viability of this opportunity before we recommend it move to delivery.

Opportunity: [1-2 sentence description]
Who benefits and how: [users, business unit]
Rough cost to build & run (best guess): [cost]
What we'd have to stop doing to fund it: [tradeoff]

Produce:
1. Five questions I should be able to answer for a sponsor before asking for funding — organized as: value, cost, risk, timing, alternative.
2. The two most likely reasons a CFO or portfolio owner would kill this, and what evidence would defuse each.
3. A "smallest defensible pilot" — the narrowest slice we could run to prove the business case is real, with a rough success metric.
4. One question I'm probably avoiding because I don't want to hear the answer.
Treat prompts like code
When a prompt produces a good question set, save the final version back into this playbook (or the Prompt Library in the Toolkit) with a note on what worked. The prompts get better every engagement — that's the compounding asset, not the individual questions.

Supporting tools

  • Maze for structured unmoderated prototype tests — tasks, success rates, follow-up questions.
  • Microsoft Clarity dropped into every shared prototype for session replay and rage-click detection.
  • markup.io for external users who need to comment without an account.
  • NPS panel on Metrics to capture willingness signals in the moment.