Imported from majidraza1228/operator-os-graph-lib (
.claude/skills/user-research-synthesis/SKILL.md). Install upstream withnpx skills add majidraza1228/operator-os-graph-lib --skill user-research-synthesis. Copyright stays with the author.
/user-research-synthesis: split discovery + checker
WHAT THIS IS
The tournament-winning user research synthesis design, the best of several tested arms, beat a strong single-pass baseline 18 of 20 on a frozen checklist against 16 of 20. It is a staffing layer around the research synthesis process you already have. It never replaces your prompt, your template, or your house format. It runs your machinery with two readers in front of it and a fact-checker behind it.
STEP 0: FIND THE USER'S MACHINERY
Do this before anything else. In most real repos, the user already has a research synthesis prompt or skill, and the whole value of this graph is that it staffs that prompt instead of overwriting it.
The sweep
Search the repo in this order and collect every hit:
.claude/skills/. Look for a directory or SKILL.md whose name matches any of:user-research-synthesis,user-interview,research-synthesis,interview-synthesis,customer-interview-summary,voice-of-customer,research-synthesizer.- Prompt library files. Any
.mdunderprompts/,prompt-library/,context-library/prompts/, or a root-levelprompts.md, whose filename or first heading mentions user research, customer interviews, synthesis, or voice of customer. - Templates. Check
templates/,context-library/templates/, or any file namedresearch-synthesis-template.md,interview-summary-template.md. - House rules. A root
CLAUDE.mdorcontext-library/CLAUDE.mdsection that states research conventions (required sections, evidence bar, voice, output location).
Naming varies. Match on meaning, not on an exact string. A skill called
turn-interviews-into-insights is a research synthesis writer.
The announcement
Say out loud what you found and what you will run, in one line, before any agent launches:
Running your .claude/skills/user-research-synthesis/SKILL.md as the writer.
Two discovery agents in front of it, a quote-and-count checker behind it.
If you find more than one candidate, list them and pick the most specific one, then say why in a half sentence. Example: "Two candidates. Using user-research-synthesis (the full skill) over user-interview (scoped to single-conversation debriefs, not multi-interview synthesis)."
THE FALLBACK
If the sweep returns nothing, load references/default-writer.md and say so
plainly:
No research synthesis prompt or skill found in this repo.
Falling back to references/default-writer.md, the library's generic spine.
If you have a house format, point me at it and I will run that instead.
Never silently pick a writer. The user needs to know whose format is about to be produced, because the output lands in front of a VP who reads it without you in the room. A silent pick is the failure that destroys trust in the whole library.
STEP 1: INPUTS
The argument
The argument to /user-research-synthesis is a research folder, not a
topic and not a single transcript pasted in chat. Example:
/user-research-synthesis research/catch-up-module/. For a single
conversation debrief, use a dedicated interview-summary skill instead, if
the repo has one; this graph is built for cross-interview synthesis.
What counts as research material
At least one file with real participant voices makes the graph worth running. Two buckets, and the graph runs strongest with both present:
| Bucket | Typical files | What the graph does with it |
|---|---|---|
| Interviews | transcripts, interview notes, call summaries | Discovery Agent 1 pulls themes with verbatim quotes, labels each voice, enforces the 2-voice rule at the source |
| Quant | survey exports, usage data, segment or persona files | Discovery Agent 2 builds the segment breakdown and hunts contradictions between quant and qualitative claims |
Interviews alone are fine; the graph degrades gracefully and folds the quant agent's job into nothing. Zero interview material is not fine.
Refuse and ask
If the folder is empty, missing, or holds only a topic name, stop and ask:
That folder has no research material in it. This graph writes from voices
and numbers that are actually in the source, so with nothing to read it
would invent quotes and counts, which is the exact failure it exists to
prevent.
Give me one of:
- a folder with interview transcripts, notes, or survey data in it
- the files pasted directly
- or say "no research yet" and I will tell you what to go collect first
Do not proceed on good intentions. A graph run on nothing costs over 4x the tokens of a single pass and produces a more confident version of the same guess.
The raw-sources rule
Discovery agents may read raw source files only, never a previous synthesis of them. The checker may not read the discovery briefs as a substitute for source either. The checker gets the original interview and survey files, always, because a brief that already contains a miscounted voice reads as settled fact by the time a downstream reader sees it. If D1 undercounts a theme at "3 of 6," a checker reading D1's brief instead of the transcripts will confirm the undercount with confidence.
Practical consequence: keep the original file paths available through the whole run and hand them to CHECK verbatim. Never delete or move source files mid-run.
STEP 2: THE GRAPH
Five stages for a normal-sized corpus. Verbatim launch prompts for every
node live in references/node-prompts.md. Copy them from there. The
descriptions below tell you what each node is for, so you can adapt the
bracket slots correctly.
Stage 1: DISCOVER (2 agents, parallel)
Node D1, interview analyst. Reads only the interview and transcript files. Extracts candidate themes with verbatim supporting quotes, each attributed to a labeled voice (P1, P2, ...) with a file and line citation. A theme needs 2 or more distinct voices to survive; a single-voice candidate gets listed separately, not promoted. Records disagreements as disagreements and flags any voice that reverses itself mid-interview, since a self-correction is real signal, not something to average away.
Node D2, quant and segment analyst. Reads only the survey, usage, and segment or persona files. Extracts every question's exact percentages with sample size and field window, builds a segment breakdown, and hunts contradictions: places where the quant data disagrees with itself, or where it's likely to contradict a loud interview claim. It never invents a segment cut the data doesn't support.
Stage 2: WRITE (1 agent, wrapping the user's machinery)
The writer loads the prompt or skill found in Step 0 and executes it. Both discovery briefs are grounded input in place of the model's own recall. The wrapper adds requirements on top of the house format, never in place of it:
- An executive one-pager with a stated confidence level
- Prioritized insights, each with evidence strength and actionability
- A segment breakdown table, where D2 found segment data
- A contradictions and open questions section, carried unresolved
- A quotes appendix with citations
- A themes-not-prioritized section: every candidate theme that didn't clear the 2-voice bar, with the cut stated, not silently dropped
Stage 3: CHECK (1 agent, adversarial)
Reads the raw source files first, builds its own notes on quotes, counts, and percentages, and only then reads the draft. Re-verifies every quote verbatim, recounts every X-of-Y claim from scratch, re-checks every theme still clears 2 distinct voices, and re-checks every percentage against its source figure, sample size, and window. Flags any material coverage gap: a segment, a dissenting voice, or a survey finding the draft missed. Every defect is reported as claim, quoted, against what the source actually says, against the required fix.
This is the step a single pass structurally cannot take on itself, because
it already believes what it wrote. See
examples/undercount-catch.md for exactly what it caught on the bench run.
Stage 4: FINALIZE (1 agent)
Applies every finding from CHECK, then re-verifies each fix against the source before it ships, so a correction doesn't introduce a new error. Adds a short corrections log at the top of the doc: what changed, what the checker found, what the source actually says. Does not rewrite anything the checker didn't flag, and does not soften a correction because the new number is less impressive than the old one.
The full build, for large corpora
Past roughly 6 interviews, a single discovery agent reading every transcript
tends to thin out its attention on the later files, which is exactly how a
voice gets missed. A separately validated shape splits discovery by
transcript batch instead of by evidence type, with a CLUSTER node merging
themes across batches before VERIFY and WRITE run. See "The full build" in
node-prompts.md for its nodes, and references/design-notes.md for the
threshold reasoning and the bench result behind it.
Hard rules (every node, every round)
- Fence line on every agent. Every launch prompt ends with an explicit file scope: "You may only read [paths] and nothing else on this machine." An unfenced agent wanders into the repo, finds a stale research doc, and cites it.
- Commit rules. Every hypothesis ends promoted or ruled out, or is marked unresolved on purpose. A labeled assumption beats a refused answer. "We don't have a number yet" is a passing answer.
- Two-voice floor. No theme in the final document holds at fewer than 2 distinct voices. Cut candidates get named in themes-not-prioritized, never quietly dropped.
- Ties go to the source, not to what reads better or what the earlier draft already claimed.
- No invented numbers, ever. No count, percentage, or date beyond what the source states, including plausible-sounding ones.
STEP 3: OUTPUT
Where it lands
Detect the repo's convention before writing:
- If
CLAUDE.mdnames an outputs convention, follow it exactly. - Else if
outputs/exists, write tooutputs/research-synthesis/[YYYY-MM-DD]-[topic].md. - Else write
research-synthesis-[topic].mdnext to the research folder.
Say the path before writing it.
What to show, in this order
First, the checker's findings, because this is the part that tells the user whether to trust the doc. One line per finding, resolved or open:
| # | Claim | Source says | Fix applied |
|---|---|---|---|
| 1 | Insight 1: "3 of 6" | P2 also voiced this in interview-02.md:9 | Corrected to 4 of 6 |
| 2 | Q5 chronological "lowest theme" | duplicate-content is actually lowest at 5.5% | Corrected citation |
Second, what the checker verified clean. A short line: "N quotes, N counts, N percentages checked, no other defects found," so an empty section doesn't read as "nothing was checked."
Third, the doc itself, or the path to it if it is long.
SELF-CHECK BEFORE PRESENTING
Run all six. Fix, do not report around.
- Every theme in the final doc holds at 2 or more distinct voices. Recount them yourself against the source.
- Every quote is verbatim, with an ellipsis marking any cut. Spot-check 3 at random against their cited file and line.
- Every X-of-Y frequency claim matches a recount, not the draft's memory of the count.
- Every percentage carries its sample size and field window in the final doc, not only in D2's brief.
- Themes-not-prioritized lists every cut theme with a reason. None were silently dropped.
- The writer machinery named in Step 0 is the one that actually ran, and the output looks like that format.
WHEN TO SKIP, AND WHAT IT COSTS
Skip the graph when the corpus is 2 or 3 interviews with no survey data, a careful single re-read catches everything a checker would at that size. Skip it too when the deliverable is an internal working doc you'll sit in the room to defend, since the checker exists to protect a reader who can't ask you a follow-up question. Skip it when you're doing a light update to a research base that already has a validated synthesis; re-running the graph on a settled base reopens ground that's already closed.
Cost, measured on our bench with Sonnet-class agents: about 4.3x the tokens of a single pass, and roughly 10 minutes of wall clock against 2.5. Five agent runs against one. Pay it when a VP acts on the doc without you in the room. Do not pay it for a same-day post-interview debrief.
KNOWN FAILURE MODES
1. A confident, well-labeled draft still ships a miscount. The bench
baseline's flagship insight read "3 of 6" with an honest-sounding evidence
tag and a correct-looking citation list, while its own segment table a few
hundred lines later already credited a fourth voice for the same need. The
draft never checked its own document against itself, because nothing in a
single pass is assigned to. This is the single failure this graph exists to
prevent; see examples/undercount-catch.md.
2. A checker reading discovery briefs instead of raw sources catches nothing. A checker handed D1's or D2's findings file instead of the original transcripts will confirm whatever miscount or mislabel already lives in that file, because it already reads as verified. This is the most common way to run the graph and pay 5x the cost for single-pass quality. Hand the checker raw sources, always.
3. One discovery agent reading past its attention span misses a voice.
Folding interview reading and survey bookkeeping into one agent, or handing
one agent more transcripts than it can read closely, produces exactly the
kind of quiet undercount the checker was built to catch downstream, at
higher cost than avoiding it upstream. Split by evidence type under 6
interviews. Split by transcript batch above that. See design-notes.md.