Every list of AI tools for product managers makes the same promise: paste your interview transcripts into an LLM and get themes in three minutes. It's true. It's also where most PMs stop, and that's the problem.
On a recent discovery project I had two interviews, one with a customer and one with the person who serves her. One call was recorded by two transcription tools. Both versions were wrong, in different ways. One got the speakers right and mangled the words. The other got the words right and handed the customer's single most important sentence, her answer to "what's your biggest problem?", to me, the interviewer.
If I had summarized that file, the AI would have reported the customer's biggest problem as my opinion, with full confidence. Her own words would have disappeared before the analysis even started.
Summarizing is the easy part
A summary averages people together. The decisions that shape an MVP live in the places a summary smooths over:
Who said what. A misattributed quote, or your own leading question folded into the customer's answer, turns your hypothesis into "customer demand."
What people say versus what they do. "The quality is good" sits two minutes away from "I stopped chasing refunds, I just eat it." The second sentence is the finding.
In B2B the customer and the person serving her often describe the same event in two different ways. One says "no answer means yes." The other was asleep when the question arrived.
Where two sides disagree.
What nobody mentioned. A step in the workflow that nobody owns rarely shows up in either interview. It shows up when you map both together.
None of this is hard for an LLM. It just has to be asked, step by step, and someone has to make the call at each step.
Seven steps, one decision each
I split the work into seven steps. Each one produces a single file that feeds the next, and each one ends with a question only the PM can answer.
Step
What the AI does
What you decide
1. Reconcile
Merges two transcripts of the same call, fixes speaker labels and mangled terms, numbers every turn
Which of its corrections are right
2. Extract
Lists each persona's pains with quotes, workarounds, numbers, and where what they say differs from what they describe doing
Whether the top pain matches your read of the call
3. Collide
Maps the workflow with an owner for every step and finds where the two sides clash
Which side v1 favors
4. Frame
Writes the problem in under 80 words, with no solution words allowed
Which problem you are actually solving
5. Scope
Proposes v1 and a cut list, each cut with a reason, a cost and a trigger to revisit
What you push back on
6. Pick the risk
Ranks assumptions by impact, uncertainty and whether a prototype can test them
Which single assumption to test
7. Build and critique
Builds one flow, then reviews it as a skeptical head of product in a fresh context
What to fix before anyone sees it
Two rules run through every step. Every claim about what someone said carries a turn ID back to the transcript, so any quote can be checked. And no solution is allowed before step 5: ideas that come up early get parked, not built on.
The mistake in my first version
The first version had a "run everything" mode. If nobody answered a question, it took its own recommended default and moved on. On paper, that was efficient.
In practice it meant the AI was choosing the problem framing, deciding which side of a conflict v1 should favor, and accepting its own cut list. Those are exactly the calls I'm supposed to make. A pipeline that makes them for you just produces a confident document faster.
So I removed the autopilot. Now every step ends the same way:
It shows the step's key results in the chat, not just a file path.
It asks one to three decisions, with its recommendation clearly marked as a recommendation.
It asks how to continue: next step, redo with changes, edit together, or stop.
Then it waits. Even if you say "run the whole pipeline," it runs one step and stops. Every answer goes into a decision log with who made it.
The AI does each step in minutes. It doesn't get to make the decisions.
Try it yourself
The nine skills (the seven steps, an interview guide for before the calls, and an orchestrator) are free and open source on GitHub. In Claude Code:
Each skill folder also works on its own as a Claude skill.
The repo includes a fictional sample case to practice on. Harvest Line is a produce supplier. Tamar, a head chef, orders by WhatsApp voice note at midnight. Avi, her sales rep, types those orders into the warehouse system until 1 a.m. Leadership wants an ordering app. The three raw transcripts are messy on purpose, and the interviews point somewhere else entirely.