AI Product Discovery Prompts a Product Owner Actually Uses
A prompt is only worth saving if it teaches you something. Each entry below is tied to a real moment in product discovery, names the prompt-engineering technique it relies on, and carries a frank note on where it will mislead you. Copy what helps, ignore the rest, and never let a confident-sounding draft replace the customer conversation it was meant to sharpen.
Product Discovery & Validation
Turn fuzzy ideas into tested bets. Prompts for opportunity assessment, interview synthesis, and validation workshops.
Related reading: Workshop activities to validate ideas
When to use: After a stakeholder pitches a feature, before it enters discovery.
Act as three skeptics reviewing this product opportunity: a budget-conscious CFO, a frustrated end user, and a delivery lead worried about scope. Opportunity: [DESCRIBE THE OPPORTUNITY IN 3-4 SENTENCES]. From each persona, give the single strongest objection and the evidence that would have to be true to overcome it. End with the one assumption that, if false, kills the idea. Do not propose solutions.
Why it works: Distinct skeptical personas surface distinct risks that a single “review” averages into mush. Forcing one killer assumption converts vague doubt into a hypothesis you can actually test.
Adapt it: Swap the CFO for a compliance officer when the opportunity touches regulated data.
Facilitator's note: The personas are stereotypes, not your actual stakeholders—use the objections to design real interviews, never as a substitute for talking to the CFO and users you actually have. A junior PM who treats this as validated research will end up defending assumptions the model invented.
Tested with: Claude Opus, GPT-5
When to use: After 3–5 discovery interviews, before writing up findings.
Here are raw notes from [N] product discovery interviews: [PASTE NOTES]. Work step by step: first list each distinct pain point with the quote that evidences it; then cluster pains into themes; then for each theme rate how strong the evidence is (strong / weak / single-source); finally name the themes you'd be wrong to act on yet. Show your reasoning at each step.
Why it works: Forcing the model to expose evidence before clustering stops confident themes being built on a single offhand comment. The strength rating keeps weak signals from masquerading as patterns.
Adapt it: Add “and propose the riskiest-assumption test for the strongest theme” to bridge into experiment design.
Facilitator's note: The model pattern-matches themes that feel coherent even from thin data—treat any “strong” rating as a claim to verify against your transcripts, not a finding. It cannot hear tone, hesitation, or what users didn't say, which is often where the real insight hides.
Tested with: Claude Opus, GPT-5
When to use: Planning a discovery workshop with cross-functional stakeholders.
Design a 90-minute product discovery workshop to validate [IDEA / PROBLEM] with [ROLES PRESENT]. Output as a table with columns: segment, minutes, activity, facilitator prompt, artifact produced, decision it informs. Keep the total to 90 minutes. Include one divergent and one convergent activity. End with the explicit go/no-go question the group must answer by the end.
Why it works: The strict table contract forces a runnable agenda with timeboxes and outputs instead of a vague list of “activities.” Tying each segment to a decision keeps the workshop from becoming a talking shop.
Adapt it: Change 90 minutes to a two-day format to generate a full discovery-sprint plan.
Facilitator's note: A generated agenda assumes a room that engages on cue—it can't read group dynamics, dominant voices, or silence. You still own facilitation; treat this as a starting structure to adapt live, not a script to read out. Confirm the go/no-go question matches your actual decision rights.
Tested with: Claude Opus, GPT-5
When to use: Before running any discovery experiment.
For this product bet—[DESCRIBE BET]—list the assumptions that must hold for it to succeed, grouped into desirability, viability, and feasibility. Write each as a falsifiable statement with a threshold (e.g. “at least X% of [segment] will…”). Do NOT include assumptions that cannot be tested within two weeks, and do NOT pad the list—fewer, sharper assumptions are better.
Why it works: The “falsifiable + threshold” requirement and the ban on untestable padding force assumptions you can actually run an experiment against, instead of motherhood statements nobody can disprove.
Adapt it: Add “rank by (impact if wrong × uncertainty)” to get a test sequence.
Facilitator's note: Thresholds the model invents (the “X%”) are placeholders, not benchmarks—set them from your own data and risk tolerance before anyone treats them as targets. The framework is sound; the specific numbers are guesses that happen to look authoritative.
Tested with: Claude Opus, GPT-5
When to use: Setting up a weekly customer-touchpoint habit.
You are a senior product discovery coach in the continuous-discovery tradition. Help me build a 20-minute interview guide to uncover [SEGMENT]'s real behaviour around [PROBLEM SPACE]. Favour questions about specific past instances over hypotheticals and opinions. Flag any question I should cut because it leads the witness. Keep it to six questions plus follow-ups.
Why it works: Priming the model with a named discovery tradition and the past-behaviour heuristic steers it away from the leading, hypothetical questions that produce flattering but useless answers.
Adapt it: Swap the interview frame for an in-product survey to get unmoderated question wording.
Facilitator's note: Good questions don't rescue a bad sample—if you only interview happy power users, the guide just helps them tell you what you want to hear. The model can't choose your participants or catch your sampling bias; that judgment stays yours.
Tested with: Claude Opus, Gemini
Backlog Generation & Prioritization
Get from a validated problem to a defensible, prioritized backlog—including the uncertainty of AI/ML work.
Related reading: Backlog management skills for AI products
When to use: Right after discovery, converting a validated problem into work.
Decompose this validated problem into a draft product backlog: [PROBLEM + TARGET OUTCOME + KNOWN CONSTRAINTS]. Break it into outcome-oriented epics, then candidate stories under each, then for each story note the discovery question it still depends on. Mark which items are assumptions to test versus things to build. Do not assign estimates or sprints.
Why it works: Decomposing outcome → epic → story keeps the backlog anchored to the goal rather than a feature wishlist. Separating “test” from “build” stops unvalidated items sneaking into delivery.
Adapt it: Add “map each epic to a measurable product metric” to connect backlog to outcomes.
Facilitator's note: A generated backlog is a strawman to refine with your team, not a plan of record—the model doesn't know your tech constraints, dependencies, or what's already half-built. Importing it wholesale creates the illusion of progress while burying the prioritization thinking that actually matters.
Tested with: Claude Opus, GPT-5
When to use: When the backlog is too big and everything feels urgent.
Score these backlog items using RICE (Reach, Impact, Confidence, Effort): [PASTE ITEMS WITH ANY DATA YOU HAVE]. Output a table sorted by RICE score showing each input value and the final score. Where I haven't given you data, mark the input as “assumed” and state the assumption. End by listing the three items whose ranking is most sensitive to a wrong assumption.
Why it works: The contract forces every score to expose its inputs, so a high rank built on an “assumed” confidence is visible rather than hidden. Flagging sensitivity tells you where to get real data first.
Adapt it: Swap RICE for WSJF when you need a cost-of-delay lens for a scaled org.
Facilitator's note: RICE produces a number, not a decision—the model's “impact” and “reach” guesses can be off by an order of magnitude and still look precise. Use the ranking to structure the prioritization conversation with stakeholders, not to outsource it. The false precision is the trap.
Tested with: Claude Opus, GPT-5
When to use: Before committing a backlog to a planning cycle.
Here is my current backlog for [PRODUCT / OUTCOME]: [PASTE]. Critique it as if you'll be blamed when it ships incomplete. What user journeys, edge cases, non-functional needs (security, accessibility, performance), and “unsexy” enabler work are missing? Group gaps by severity. Be specific to what's here—don't list generic best practices I clearly already covered.
Why it works: Framing the model as accountable for omissions pushes it past generic checklists toward the specific gaps in your list. The “don't repeat what I covered” guard keeps the critique high-signal.
Adapt it: Re-run with “critique purely for accessibility and inclusive design” for a focused audit.
Facilitator's note: The model flags common gaps, not gaps specific to your domain or regulatory context—a fintech or health product has obligations it won't reliably know. Treat its list as a prompt to think harder, and pull in your security and compliance people for anything it can't be trusted to know.
Tested with: Claude Opus, Gemini
When to use: Planning work for a feature whose core is a model, not deterministic code.
I'm building [AI / ML FEATURE]. Help me structure the backlog around its uncertainty. List the work items, and for each label whether its effort/outcome is predictable or research-dependent (a spike). For the research-dependent ones, state what we'd need to learn before we can size them, and be explicit about your confidence in each call. Flag where a “story” is really an open research question in disguise.
Why it works: Asking for explicit confidence and a spike/build split stops ML work being planned as if it were ordinary CRUD. Surfacing disguised research questions prevents committing to deadlines on things nobody can yet estimate.
Adapt it: Add “and propose a kill criterion for each research spike” to timebox exploration.
Facilitator's note: AI-product effort is genuinely hard to predict and the model's own confidence labels are themselves uncertain—don't convert its “predictable” tag into a committed estimate. Data quality, eval results, and model behaviour will move the work in ways no upfront plan captures. Keep buffers and revisit often.
Tested with: Claude Opus, GPT-5
When to use: Refining a story before it enters a sprint.
Here are two examples of acceptance criteria my team considers good: [EXAMPLE 1] [EXAMPLE 2] Now write acceptance criteria in the same style and rigour for this story: [STORY]. Include the unhappy paths and at least one measurable / non-functional criterion. Match the format of my examples exactly.
Why it works: Two concrete examples teach the model your team's bar far more reliably than an adjective like “detailed.” Matching your format keeps the output drop-in usable.
Adapt it: Provide examples of bad criteria too, asking it to avoid those failure modes.
Facilitator's note: AC quality is capped by your examples—feed it mediocre ones and you'll scale mediocrity. The model also can't know your real system's edge cases, so its unhappy paths are plausible guesses your engineers must confirm against the actual implementation.
Tested with: Claude Opus, GPT-5
Epic Breakdown & Story Slicing
Split big work into thin, user-valuable slices instead of horizontal layers nobody can ship.
Related reading: User story slicing techniques
When to use: An epic is too big for a sprint and you're tempted to split by layer.
I need to slice this epic into stories that each deliver end-to-end user value: [EPIC + USER + OUTCOME]. Reason step by step: identify the core user journey; find the thinnest end-to-end slice that's still useful; then list further slices in priority order. For each slice, name the value it delivers and what it deliberately leaves out. Avoid slicing by technical layer (frontend / backend).
Why it works: Walking the user journey before slicing keeps each story shippable and valuable. The explicit ban on layer-slicing prevents the classic anti-pattern of stories nobody can release alone.
Adapt it: Add “using the SPIDR pattern, label which technique each slice uses” to teach the team a vocabulary.
Facilitator's note: The “thinnest valuable slice” depends on your users and architecture, which the model only knows from your brief—its slices are a proposal to challenge in refinement, not a mandate. A junior PO who ships them unexamined can deliver technically correct stories that don't actually let a user finish a task.
Tested with: Claude Opus, GPT-5
When to use: You want options for how to split, not just one answer.
Take this epic—[EPIC]—and produce a table applying five slicing techniques: by workflow step, by business rule, by happy/unhappy path, by data variation, and by effort/spike. Columns: technique, resulting stories, smallest releasable story, when this technique is the right choice. Keep stories vertical (user-valuable), not horizontal layers.
Why it works: The table contract forces breadth—five lenses instead of the first split that comes to mind—and the “when right” column teaches judgment about which technique fits which epic.
Adapt it: Ask it to recommend the single best technique for your epic and justify it against the others.
Facilitator's note: More slicing options can become analysis paralysis—the goal is shipping, not a perfect taxonomy. The model can't tell you which split your team can actually deliver given current skills and dependencies, so use this to widen options, then decide fast with your engineers.
Tested with: Claude Opus, Gemini
When to use: A stakeholder asks for “a dashboard” or some solution-shaped feature.
A stakeholder asked for [SOLUTION-SHAPED REQUEST]. Before we treat it as an epic, interview me to uncover the real problem. Ask one question at a time—about who, the job-to-be-done, the current workaround, and how we'd know it worked—wait for my answer, then go deeper. After five questions, restate the underlying problem and whether the original request is even the right solution.
Why it works: One-question-at-a-time probing forces you to articulate the problem behind the feature request, often revealing the asked-for solution is wrong. It rehearses the conversation you should be having with the stakeholder.
Adapt it: Point it at your own draft epic to stress-test whether you've assumed the solution.
Facilitator's note: This rehearses your thinking; it does not replace the real conversation with the stakeholder, who holds context the model never will. Don't return to them with the model's restated problem as fact—return with better questions. The risk is mistaking a tidy AI reframing for actual alignment.
Tested with: Claude Opus, GPT-5
When to use: Planning the shape of a release across a user journey.
Build a user story map for [PRODUCT / RELEASE GOAL]. Decompose the primary user's journey into sequential activities (the backbone), then under each activity list the tasks/stories from essential to nice-to-have. Then mark a horizontal “walking skeleton” line for the minimum viable release. Output as an indented outline. Note any activity where you're guessing the user's steps.
Why it works: Story mapping by decomposition keeps the release organised around the user's flow rather than a flat backlog. The walking-skeleton line forces an explicit MVP cut.
Adapt it: Ask for two release slices (MVP and fast-follow) instead of one line.
Facilitator's note: A map built from your brief reflects the journey you described, which may not be the journey users actually take—validate the backbone against real behaviour before trusting the cut line. The model's “essential vs nice-to-have” is opinion dressed as structure; your evidence decides.
Tested with: Claude Opus, GPT-5
When to use: During refinement, sanity-checking story size.
Review these stories for hidden size and ambiguity: [PASTE STORIES]. For each, flag if it's likely larger than a few days, hides multiple user outcomes, or has acceptance criteria vague enough to cause mid-sprint debate. Suggest a split only where warranted. Critique what's written—don't assume requirements I didn't state.
Why it works: A focused critique pass catches the oversized, multi-outcome stories that derail sprints, while the “don't assume” guard keeps it from inventing scope you never intended.
Adapt it: Add “estimate relative size in t-shirt sizes with your confidence” for a quick sizing gut-check.
Facilitator's note: Size is relative to your team's velocity and skills, which the model can't see—its “too big” is a heuristic, not your team's reality. Use it to spark the splitting conversation in refinement; the people doing the work are the only valid estimators.
Tested with: Claude Opus, Gemini
PRD & Requirements Drafting
Draft specs whose gaps are visible, then stress-test them before engineering ever sees them.
Related reading: Product Owner online training
When to use: Turning a validated opportunity into a written spec.
Draft a PRD for [FEATURE] using these sections only: Problem & evidence, Target user & job, Goals & non-goals, Success metrics, Scope (in/out), Key user flows, Open questions, Risks. Under Open questions, list everything you had to assume because I didn't tell you. Keep each section tight. Do not invent metrics or user research—mark those as “TBD: needs data.”
Why it works: A fixed section contract plus a forced “assumptions” list produces a PRD whose gaps are visible, so review focuses on real open questions instead of polished-looking fiction.
Adapt it: Add a “Decision log” section to capture why scope was cut.
Facilitator's note: The most dangerous PRD is the one that reads complete but rests on invented evidence—scan the Open-questions and TBD markers first, because that's where the model papered over what it didn't know. A confident draft can lull you into skipping the validation the PRD is supposed to force.
Tested with: Claude Opus, GPT-5
When to use: Before circulating a PRD draft.
Two skeptical reviewers critique my PRD: a staff engineer and a designer. PRD: [PASTE]. As the engineer: where is it ambiguous, what edge cases are unhandled, what forces a mid-sprint clarification? As the designer: where does it assume UX that doesn't exist, where is the user problem under-specified? End with the three questions most likely to derail kickoff. Critique only—do not rewrite the PRD.
Why it works: Two named adversarial personas surface engineering and UX failure modes a generic review blends together. “Critique, don't rewrite” keeps the output diagnostic and the document yours.
Adapt it: Add a “skeptical security reviewer” persona for anything touching user data.
Facilitator's note: The critique reflects common patterns, not your specific users or codebase—and the real danger is treating AI-surfaced edge cases as exhaustive, then stopping. It's a prompt for your judgment and your engineers' review, never a sign-off. The PRD's correctness is still your responsibility.
Tested with: Claude Opus, GPT-5
When to use: Handing a PRD to engineering and QA.
Decompose this PRD into testable requirements: [PASTE PRD]. For each user flow, produce requirements as Given / When / Then scenarios covering the happy path, edge cases, and failure states. Mark any requirement that depends on an external system or an unconfirmed assumption. Don't gold-plate—only requirements traceable to the PRD.
Why it works: Decomposing into Given/When/Then makes requirements verifiable by QA. Flagging external dependencies and unconfirmed assumptions stops the team building against guesses.
Adapt it: Ask it to additionally output a coverage-gap list—PRD statements with no requirement yet.
Facilitator's note: Generated scenarios are a first pass your QA and engineers must validate against the real system—the model invents plausible edge cases but misses the ones unique to your architecture and data. Treating its scenarios as a complete test plan is how production bugs slip through.
Tested with: Claude Opus, Gemini
When to use: At the very start, before anyone proposes features.
Help me write a one-paragraph problem statement for [SITUATION]. It must name the user, their context, the job they're trying to do, and the cost of the status quo—using only evidence I provide: [EVIDENCE]. Do NOT propose any solution, feature, or technology. Do NOT add evidence I didn't give. If the evidence is too thin to support a claim, say so.
Why it works: Banning solutions and unsupported evidence forces a disciplined problem statement. The “say if evidence is thin” instruction surfaces when you're not actually ready to proceed.
Adapt it: Follow up with “now list what we'd need to learn to strengthen this” to plan discovery.
Facilitator's note: A well-written problem statement can feel like progress while resting on weak evidence—the polish is the model's, the evidence is yours, and only you can judge if it's real. Don't let a clean paragraph substitute for the customer conversations you haven't had.
Tested with: Claude Opus, GPT-5
Upskilling & the AI-Era PO
Stay relevant as AI reshapes the role—skills planning, certifications, and the move toward agentic PM.
Related reading: ICAgile AI for Product Discovery training
When to use: A PO deciding how to stay relevant as AI reshapes the role.
I'm a Product Owner with [EXPERIENCE / CONTEXT]. Build a realistic 90-day plan to become effective at AI-assisted discovery and delivery. Reason from my current gaps to priorities: first infer my likely skill gaps from my context, then sequence what to learn, then suggest one practical artifact to produce each fortnight. Be honest about what AI won't do for me. Keep it to what fits [HOURS/WEEK].
Why it works: Reasoning from stated context → gaps → sequence produces a plan tailored to you rather than a generic syllabus. The time constraint keeps it achievable.
Adapt it: Aim it at a whole team by swapping individual context for team maturity.
Facilitator's note: The model infers your gaps from a short description and will guess wrong about your real strengths and blind spots—treat the plan as a draft to validate with a mentor or your manager. It also can't know your org's actual tooling and constraints, which shape what's worth learning.
Tested with: Claude Opus, GPT-5
When to use: Weighing a training such as ICAgile's ICP-AIPD against alternatives.
I'm considering [CERTIFICATION / TRAINING] to strengthen my AI product skills. Given my goal—[GOAL]—assess it honestly: what it credibly teaches, what it won't, who it's genuinely for, and cheaper or better alternatives. Be explicit about where you're uncertain or your information may be dated, and tell me what to verify on the official source before deciding.
Why it works: Asking for explicit uncertainty and a “verify this” list stops the model stating cert details as fact when its training may be stale, turning it into a research aid rather than an oracle.
Adapt it: Have it draft three questions to ask a course provider before paying.
Facilitator's note: Course curricula, prices, and exam details change and the model may be working from outdated information—never make a purchase decision on its description alone. Confirm everything against the official provider page; treat this as a way to frame your evaluation, not the evaluation itself.
Tested with: Claude Opus, GPT-5
When to use: Stuck on a hard call—cut scope, push back on a stakeholder, and so on.
Act as a product coach. I'm facing this decision: [DECISION + CONSTRAINTS]. Don't give advice yet. Ask me one sharp question at a time to expose my assumptions, the real stakes, and what I'm avoiding. After about six questions, summarise the decision as you now understand it and the two or three options I actually have—then let me choose.
Why it works: Socratic questioning surfaces the avoided issue and untested assumptions behind a hard decision far better than instant advice. Withholding the recommendation keeps ownership with you.
Adapt it: End with “name the option my future self would regret not trying” for a gut-check.
Facilitator's note: The coach has no stake in your career and no knowledge of your politics or relationships—its summary can sound wise while missing what actually matters in your org. Use it to think more clearly, then carry the decision to people who know your context.
Tested with: Claude Opus, Gemini
When to use: A PM exploring how autonomous agents change the operating model.
Help me think through moving parts of my product workflow toward AI agents. Decompose my role into recurring tasks: [LIST, OR ASK ME]. For each, classify it as: safe to delegate to an agent now, augment-only (human in the loop), or keep fully human—and why. Flag the accountability that must stay with me no matter what. Be concrete, not hypey.
Why it works: Decomposing the role task-by-task replaces vague “AI will change everything” anxiety with a concrete delegation map. Forcing an accountability column reinforces that responsibility doesn't transfer to the agent.
Adapt it: Add “and the risk if this delegation fails silently” to each row.
Facilitator's note: “Safe to delegate now” is the model's optimism, not a guarantee—agents fail in ways that are easy to miss, and accountability for outcomes stays with you regardless of what you automated. Pilot anything before trusting it, and never delegate the judgment a customer would hold you responsible for.
Tested with: Claude Opus, GPT-5
Responsible & Governed Discovery
Surface bias, harm, and governance questions during discovery—before launch, not after.
Related reading: The 48-hour AI discovery framework (pillar)
When to use: During discovery of any feature that runs a model on user data.
Three reviewers examine this AI feature for harm and bias: an affected end user from a marginalised group, a responsible-AI specialist, and a regulator. Feature: [DESCRIBE INCLUDING THE DATA AND THE DECISIONS IT INFLUENCES]. From each, give the most serious concern, who could be harmed and how, and one thing we should test or change before building. Be concrete about this feature, not generic AI ethics.
Why it works: Distinct stakeholder lenses surface concrete harms a generic “is this ethical?” prompt glosses over. Naming who is harmed and a testable change makes the risk actionable in discovery rather than after launch.
Adapt it: Add a “data protection officer” persona when personal data is processed.
Facilitator's note: This raises candidate risks; it is not a fairness audit or a legal review. The model misses harms specific to your users and jurisdiction, and can also raise false alarms—route serious concerns to real responsible-AI and legal experts. Treating its output as clearance is a genuine liability.
Tested with: Claude Opus, GPT-5
When to use: Planning governance for an AI product in a standards-conscious org.
We're shaping [AI FEATURE] and want to align with ISO/IEC 42001 thinking on responsible AI management. At a discovery level, what governance questions should we be able to answer about this feature—around risk, data, human oversight, and monitoring? Be explicit about where you're uncertain on the standard's specifics and what we must confirm with a qualified source. Don't quote clause numbers you're unsure of.
Why it works: Asking for governance questions plus explicit uncertainty keeps the output a discovery checklist rather than false compliance guidance. The “don't quote uncertain clauses” guard prevents fabricated standard references.
Adapt it: Reframe for the EU AI Act risk tiers if that's your binding regime.
Facilitator's note: The model is not a substitute for the actual standard or a certified auditor and may misstate or invent specifics—use this only to prepare better questions for qualified governance people. Acting on its interpretation as if it were compliance advice is exactly the risk ISO 42001 exists to manage.
Tested with: Claude Opus, Gemini
When to use: Documenting risk before greenlighting an AI feature build.
Draft a one-page model-risk note for [AI FEATURE] using these headings: Intended use, Who's affected, Failure modes & impact, Data risks (quality / bias / privacy), Human oversight plan, What we'll monitor post-launch, Open risks. Keep claims grounded in what I tell you: [CONTEXT]. Mark anything speculative. Do not overstate our safeguards.
Why it works: A fixed-heading contract produces a reviewable risk artifact instead of prose, and the “don't overstate safeguards” instruction counters the model's tendency to write reassuring boilerplate.
Adapt it: Add a “rollback / kill switch” heading for higher-stakes deployments.
Facilitator's note: A tidy risk note can create false comfort—the document is only as honest as the context you give and the safeguards you actually implement. The model will happily describe oversight you haven't built; verify every claimed control exists before this note informs a go decision.
Tested with: Claude Opus, GPT-5
Frequently Asked Questions
Are these AI product discovery prompts free to use?
Yes. Copy any prompt, adapt the bracketed variables to your product, and use it in your own AI tool. They are shared as a teaching resource for the product community.
Which AI models do these prompts work with?
They are model-agnostic and were drafted and checked against current frontier chat models such as Claude and GPT-class models. Output quality still depends on the context you provide in the brackets.
Can AI replace product discovery?
No. Every prompt here assumes a human product owner stays accountable. AI accelerates synthesis and drafting; it cannot talk to your customers, own the decision, or be accountable for the outcome.
How should I adapt a prompt to my own product?
Replace the [BRACKETED VARIABLES] with your real context, keep the technique and the constraints intact, and always treat the output as a first draft to validate—never as a finished artifact.