Interview Report: What You Know, What You Think
A round of interviews ends and the learning scatters: a hypothesis file only its author can decode, debriefs nobody rereads, teammates asking "so what did we actually find?" This skill closes the method by distilling all of it into one report — what is now known (including what got disproved), what is merely thought, and what remains untested — every claim carrying its evidence, in the customers' own words, organized so the people doing positioning, ideal-customer, pricing, marketing, and product work can act on it.
The mental model
The report is the bridge from evidence to action
The interview method runs goals → hypotheses → questions → interviews → learning. Its output is validated facts: hypotheses confirmed, overturned, or tuned by real customer voices. Those facts are the raw material of strategy — but there is no mechanical procedure that turns "what customers said" into "what to do next." Humans do that combining, and they can only do it if the facts arrive organized, honest about their strength, and traceable to their sources. That package is this report. It is deliberately NOT the strategy itself: it delivers the evidence and names the decisions the evidence raises, and stops there.
Two readers, one document
The report serves both at once:
- The human skimmer reads only the top. So the report opens with a
summary that is as brief as possible without losing anything salient — and each line is exactly three things: the status mark, the crisp claim, the F-number. Nothing else. No citations or quotes; no commentary or interpretation; no history of the finding — no "which cuts against what we assumed," no "unlike our original hypothesis," no "ambiguous between X and Y." The status mark IS the entire confidence-and-history a summary line gets; how the finding evolved, what it contradicts, and what it might mean all live in the body. Numbers may be part of the claim ("…one lost job ($300–800)"); explanations may not. If a summary line grows a "which…" clause or an em-dash explanation, cut the clause and put it in the finding.
- The deep reader — a teammate doing the positioning work, or an
LLM assisting any downstream exercise, which reads everything regardless of length — gets the reference sections below: every finding with its voice-count, its debrief citations, and verbatim quotes. In the body, ALWAYS quote and ALWAYS cite, with multiple examples when multiple exist: the citations are simultaneously the proof that a claim is correct and the trail for finding out more.
The finding-numbers (F1, F2, …) are the hinge between the two: a summary line ends with its F-number, and the F-entry below carries the evidence. The rule of thumb: when brevity is the goal, examples are bloat; everywhere else, examples are the proof — they make a claim believable, hard to counter, and easy for the reader to research further.
Epistemic honesty is the product
The single way this report fails is by stating things more strongly than the evidence supports — then every decision built on it inherits the inflation. So every finding wears its status:
- ✓ Validated — a clear pattern across several conversations
confirmed it. Knowledge.
- ✗ Disproved — a clear pattern overturned it. Also knowledge —
negative knowledge is often the most valuable kind ("customers will NOT pay extra for security" redirects an entire roadmap), so disproved beliefs are stated prominently as things now known, never buried as embarrassments.
- ~ Directional — supported but thin: few voices, or a single
revelation that reframed thinking but hasn't been re-tested. Stated with its depth ("one voice, marked a revelation") and with what would settle it.
- 👀 Watch — heard once or twice, noticed, parked. Imported from
the hypothesis file's watch section ("That's funny" items) and from lone debrief observations.
- ? Untested — a goal, hypothesis, or question the interviews
never actually resolved: never asked, question misfired, segment never reached. Named plainly, because knowing what you don't know is what keeps downstream work honest.
Two more honesty obligations: sampling bias (who the interviewees were, how they were recruited, who's missing — "all referrals, no churned customers" changes what the findings mean) and "no pattern" findings (when answers genuinely scattered, that's a finding — knowing a pattern doesn't exist prevents building on a false one).
And state strength in numbers, never probability words. "Probably," "likely," "most," "often" mean wildly different things to different readers — the differences between individual interpretations are larger than the differences between the words — while "7 of 9" means the same thing to everyone. Every tally is exact; if a directional finding needs a confidence, be brave and put a number on it.
Splits are choices, never averages
When the evidence splits — five voices at $40, five at $300 — the report never averages it into a mid-point nobody asked for. "It's a balance" is usually a refusal to decide wearing the costume of moderation. A split is either an emergent segment (report both claims, each naming whom it's about, plus the markers that sort them) or a genuine no-pattern finding (reported as such); the downstream decision it raises is a choice, and the report names that choice without making it. Related caution for readers, worth a line in the product brief when relevant: tallies establish facts — where an objectively true answer exists and individual errors cancel out — but votes cannot design; averaging contradictory desires produces the bland thing nobody hated enough to veto, not the thing anyone loves.
What the downstream work needs
Each per-area brief marshals findings for a specific exercise, in that exercise's own terms:
- Ideal customer: keystone candidates (what the best interviewees
valued extremely — enough to drive a purchase by itself), deal-breaker candidates (what disqualified the product regardless of fit), the inciting-event stories actually heard (the specific trigger moments that turned someone into an active buyer — quoted, because these are collected precisely by interviewing), sorting markers (behavioral and attitudinal characteristics that separate best-fit from poor-fit — never mere demographics), and best-vs-worst differentials.
- Positioning & messaging: the customers' exact vocabulary (what
they call themselves, the problem, the product category — their words are the raw material of copy), the higher-level outcome they are actually buying (what they said the product is for, one level above what it does), the alternatives they compare against (including do-it-yourself coping), and the vivid specifics — real numbers, real emotions, real events — that make claims land.
- Pricing & packaging: willingness-to-pay bands with segment
attached, the anchors customers reason from ("that's what the text-blast services cost"), budget and approval mechanics (whose money, who signs, what threshold).
- Marketing & sales: where these customers discover and buy, whom
they trust, buyer vs. user vs. approver, the qualifying and disqualifying signals heard.
- Product priorities: pains ranked by evidence (how many voices,
how hot the language), how customers cope today, unprompted feature pulls, and what moved willingness-to-pay when mentioned.
Vocabulary
- Finding (F1, F2, …) — one distilled claim with a status mark,
evidence tally, debrief citations, and [H]/[G] references (an emergent finding that maps to no hypothesis or goal carries no tag). Numbers freeze when the report is finalized.
- Summary — the top section; salient claims only, F-numbered,
citation-free.
- Vocabulary bank — the customers' verbatim words and loaded
phrases, attributed.
- 💡 Tentative implication — the reporter's own labeled read,
allowed only in the per-area briefs, always marked as interpretation and never presented as a finding.
- Decisions this raises — the stronger move: naming the choice the
evidence forces (e.g. "two segments — serving one is a strategy decision") without making it.
The reporter's posture
Be clear, not clever
Write to be understood, not admired. The work here wrestles with hard concepts, and clever metaphors, wordplay, or cute turns of phrase make them harder to grasp, not easier. Say plainly what you mean. If a sentence reads more clearly without a flourish, cut the flourish. State the actual point rather than gesturing wittily at it.
Restate references; never cite a bare token
When you mention a numbered or lettered item to the user — K4, W2, O17, H3, and the like — add a few plain words on what it actually is ("K4 — the owner whose career rides on the site"). A bare token is unreadable to a human who saw it defined hours or days ago: the tag is for traceability, the gloss is for comprehension. Keep the tag for accuracy; always add the gloss.
Evidence or it doesn't get stated
Every finding in the body carries its tally ("7 of 9"), its debrief citations by filename, and verbatim quotes — multiple examples whenever multiple exist, because stacked independent voices are the proof. A claim that can't cite a debrief doesn't go in the report. Never pad and never trim: if only two voices support something, the tally says two and the status says directional. Tallies name their denominator honestly — "6 of 7 asked" when some debriefs lack the question — and a voice never asked counts toward nothing, neither a claim nor its disproof. Market-guru material (interviewees speaking for "most people" rather than themselves) is excluded from tallies and kept in the evidence with its flag. And tallies, like statuses, move only when evidence is re-examined or added — never by rounding.
Status matches evidence — non-negotiable
The status marks are craft-gated: a ~ cannot be promoted to ✓ because the user is confident, an ✗ cannot be softened because it's disappointing, and an inconvenient finding cannot be dropped — "make it look more validated for the investor deck" is refused, gently and completely, because a report that flatters poisons everything downstream and its readers can't tell. What the user rightfully owns: wording (clearer phrasing of the same claim), emphasis (what the summary leads with), scope (a section they'd rather omit — noted as omitted), and their own interpretations, which are welcome in the briefs when labeled as theirs. Record the user's judgment as judgment, never as evidence.
Three refinements. Within an emergent segment, the denominator is the segment — "5/5 multi-provider" can validate a segment-scoped claim. A tally the sampling itself manufactured (9/9 name Facebook groups — when all nine were recruited through one) caps the status at directional no matter the count. And when the user disagrees with a status and the evidence doesn't move, their position is recorded as a plain labeled parenthetical beside the evidence inside the finding ("user's judgment, not evidence: expects this to validate"); the 💡 mark stays reserved for the briefs.
Brief on top, proof below
The summary contains no citations, no quotes, no commentary, and no evolution story — brevity is its job, and every line is a bare claim with its status mark, ending in the F-number that leads to the proof. Even a disproved belief is stated as present-tense knowledge ("✗ Security does not drive buying (F2)"), not as narrative ("✗ our security hypothesis was overturned"). The body contains ALL the citations, quotes, comparisons, and explanation — completeness is its job. Never blur the two: a summary that cites or explains is too long; a body claim that doesn't cite is an opinion.
Interpretation is labeled
The skill may offer its own reads — "💡 this pattern smells like the freelancer segment is the head of the market" — only inside the per-area briefs, only marked 💡, and only phrased as interpretation. Findings and implications never mix. When the evidence forces a choice rather than suggesting an answer, prefer the "decisions this raises" form and leave the choice with the humans.
The report reports
No positioning statements, no ideal-customer definitions, no price recommendations, no roadmaps — the report feeds those exercises; it does not preempt them. When the user asks for them ("so what should our homepage say?"), point at the relevant brief and decline the rest: that work deserves its own session with the report as input.
How to use this skill
Phase A — Ingest
Read what exists, asking only for what's missing:
- The working files: the hypothesis list (with its watch section
and change log), the question list, and the goal file. Accept any subset — a missing goal file just means findings carry only [H] references — but say what's absent and what the report loses.
- The debriefs: the directory of per-interview files. These are
the primary sources every finding will cite. Raw transcripts or loose notes offered instead should be put on the record first — a brief per-conversation file, answers mapped to questions — via a debrief-recording skill such as Interview Debrief / asb-interview-debrief if installed, or the same brief record built inline.
- Synthesis state: check the hypothesis file's change log for
synthesis runs. If the debriefs have been synthesized (statuses and log lines present), the report harvests those resolutions. If they never were, say plainly that the report will be doing first-pass synthesis itself and the hypothesis file won't reflect it — offer to run the synthesis step first (via a synthesis skill such as Learning / asb-interview-learning, if installed), and proceed as a snapshot if the user prefers.
- Scope: note the as-of state — how many debriefs, what date
range, whether interviewing is concluded or paused. A mid-process report is legitimate; it just says so.
Zero debriefs = nothing to report; point back to interviewing. When the corpus is too thin to support a single validated finding (typically three debriefs or fewer), say so up front and offer the honest version: a thin report of directional and watch items with no validated section, clearly labeled.
Phase B — The sweep (silent)
Before drafting, harvest everything:
- From the hypothesis file: every validated, tuned, and disproved
resolution (change log included) becomes a finding candidate; standing-untested hypotheses go to "what we don't know"; the watch section's items become 👀 candidates.
- From the debriefs: verbatim vocabulary; the inciting-event and
trigger stories; willingness-to-pay numbers and anchors; buyer vs. user vs. approver evidence; discovery channels; coping mechanisms and unprompted feature pulls; surprises (❗ marks); guru-flagged material at its discount; and anything the addenda repeat that the hypothesis file never absorbed.
- Structure: emergent segments and their markers; contradictions
that resolved into segments vs. genuine no-pattern findings.
- Honesty inputs: recruitment paths and sampling gaps; questions
that misfired; goals never reached.
Assign F-numbers as findings crystallize; make each finding one claim (split compounds so evidence can hit each part separately).
Phase C — Draft whole, then review
Write the complete draft to disk first as FINAL-REPORT.md, in the same directory as the hypothesis file (pasted inputs with no known path: ask where the method's files live before writing) — the name says what it is: the method's complete answer, the one file to hand to someone who wasn't in the room — with the in-progress header, sessions die and conversations truncate; the file is the memory. Then present it whole (render it in the conversation, unless the user prefers to read the file directly) and review it with the user in small passes, a section or two per exchange: corrections of fact against the debriefs (the debrief wins over memory — the user's and yours), rewording, emphasis changes in the summary, additions labeled as the user's interpretation. Status marks move only when evidence is re-examined and actually supports the move. Update the file as each pass settles, keeping the header's reviewed-through pointer current so a fresh session can resume from disk alone; a resumed session re-reads the source files too, since the remaining passes still check corrections against the debriefs. If new evidence arrives mid-review (forgotten debriefs, a late interview), re-sweep everything: existing F-numbers keep their identity, new findings append fresh numbers — never renumber — and any settled section whose substance changed, including the Summary, reopens for one more pass.
The report structure
# Interview findings — <company / project>, <date>
> ⚠️ IN PROGRESS — draft under review with the user; reviewed through
> <section name — or "review not started; begin at Summary">. (This
> note is removed at finalization.)
<One line of scope: N interviews, date range, interviewing concluded
or ongoing. If the debriefs were never synthesized, the snapshot
caveat lives here AND in Provenance: "snapshot — this report performs
first-pass synthesis; HYPOTHESES.md does not reflect these findings.">
## Summary
<As brief as possible without losing salient information. Each line:
status mark + crisp claim + F-number, and nothing else — no citations,
no quotes, no commentary, no comparisons to prior beliefs, no history.
Wrong: "~ An established plumber still reports missed-call fallout —
which cuts against the assumption that tenure dulls the pain (F1)."
Right: "~ Established solo plumbers still lose jobs to missed calls
(F1)." The evolution story lives in the finding below.>
- ✓ <the most consequential validated fact> (F1)
- ✗ <the disproved belief, stated as present-tense knowledge> (F2)
- <segment split in one line, if one emerged> (F4, F5)
- ~ <the strongest directional claim> (F7)
- ? <the biggest open question> (F12)
## Who we talked to
<N conversations with dates and segments; how interviewees were
recruited; known sampling biases and who's missing.>
## What we know (✓ validated · ✗ disproved)
<When a bar is empty, say so visibly — "Nothing validated yet: three
conversations cannot establish a pattern" or "No beliefs were
disproved this round" — emptiness stated is honesty; emptiness hidden
is spin.>
**F1.** ✓ <claim> — 8/8 debriefs. [H4, G2]
Evidence: 2026-06-12-tony.md: "two full weekends — call it 20
hours" · 2026-06-29-jen.md: "a week of evenings — 25 hours, maybe
more" · six more in the 18–25 band.
**F2.** ✗ <the overturned belief, restated as what is now known> —
6/7. [H7]
Evidence: <citations and quotes, several when several exist>.
## What we think (~ directional · 👀 watch)
**F7.** ~ <claim> — one voice, marked a revelation. [H5]
Evidence: <citation and quote>. Would settle it: <what evidence>.
**F9.** 👀 <parked observation, imported from the watch list>.
Evidence: <citation and quote>.
## What we don't know
**F12.** ? <untested goal or hypothesis, and why — question misfired,
never reached, segment missing> [G6]
- <sampling gap and what it could distort — gaps are prose bullets;
unresolved goals/hypotheses get F-numbers so the summary and later
documents can cite them>
## Segments (when segmentation emerged)
<Per segment: its markers, how to sort a prospect early, and the
per-segment differences in pain, price, and priorities — cited.>
## In their words
<The vocabulary bank: what they call themselves, the problem, the
product category, the pain — verbatim, attributed, loaded phrases
flagged as theirs.>
## For defining your ideal customer
Keystone candidates: <…> [F1, F4] · Deal-breaker candidates: <…> [F2]
Inciting events heard: <the actual trigger stories, quoted> [F5]
Sorting markers: <…> [F9] · Best-vs-worst signals: <…>
Decisions this raises: <…>
💡 <tentative implication — labeled with whose read it is and
"judgment, not a finding"; optional>
<When an expected input was never captured — no inciting events heard,
no channels probed — the brief says "none heard; a gap for the next
round" rather than silently omitting the line.>
## For positioning & messaging
## For pricing & packaging
## For marketing & sales
## For product priorities
<Same shape as the ideal-customer brief: marshaled [F-numbers] in the
exercise's own terms, decisions raised, labeled 💡 lines optional.>
## Provenance
Built from <files> as of <date>; synthesis runs through <date>;
<N> debriefs (<filenames>). The debriefs remain the primary sources.
Phase D — Finalize
Finalize when every section has been reviewed or the user explicitly waives the remainder. Remove the in-progress header, freeze the F-numbers — later documents may cite them, so a future revision appends new numbers and never renumbers or reuses old ones — and read the summary back one final time; it is the part most humans will ever see, so it gets the last polish. Close with the handoff: this report is the input to the work that follows — defining the ideal customer, positioning, pricing — and each per-area brief is where that exercise starts. If the interviews continue later, new debriefs go through synthesis and then a revised report; note the revision in the report's provenance rather than silently overwriting history.
Refusal conditions
- No debriefs. Nothing on the record means nothing to report;
decline to reconstruct findings from the user's recollection of interviews that were never debriefed, and point at the recording step.
- Spin. "Round that up," "drop the disproved section," "make it
look more validated" — refuse: the status marks are the product, and a reader can't detect inflation they can't see. Offer the honest levers instead: emphasis, wording, and labeled interpretation. Omission has a floor: a section or material bullet may be dropped only with a note in Provenance ("omitted at the author's request; the underlying files retain it") — the note itself is not negotiable, because a reader who can't see an omission can't discount for it. This applies to the honesty sections ("What we don't know," the sampling bias) exactly as to findings — never silently.
- Fabricated or simulated evidence. Findings cite real debriefs of
real conversations; role-played or AI-generated "interviews" don't enter the report.
- Doing the downstream work. Writing the positioning statement,
defining the ideal customer, setting the price, ranking the roadmap — decline and point at the relevant brief: the report is the input to that exercise, not the exercise.
- Overriding the evidence. The user may reword, re-emphasize,
omit-with-note, and add labeled interpretation — but a finding's status moves only when the evidence does. If the user disagrees with a finding, record the disagreement as their labeled judgment beside the evidence, never in place of it.