Mandar Bhoyar UXR & Strategy

AI-Assisted Qualitative Research Workflow

Defining the trust architecture for AI in qualitative research, then shipping it.

Sample

11research practitioners, from first-year to seven-plus years in the craft

Capabilities weighed

6scored on what researchers wanted against what they were ready to adopt

Analysis time reduced

60%after the recommendation was built and put into my own workflow

What it produced

An evidence dossier, a capability roadmap, and a four-phase pilot with pass / fail gates

The strategic question

Every research team was under pressure to adopt AI. Nobody had asked researchers what they would actually trust. Can AI synthesize research? Which parts of synthesis can AI accelerate without weakening context, evidence, or researcher accountability?

Why this method

A trust question cannot be answered with a features survey. People have to reveal where they would actually let go of control, and that shows up in what they do, not in what they rate.

So I paired stated preference against demonstrated burden. Participants rated six AI capabilities, but they also mapped their own workflow by friction type and ranked which tasks were most challenging, most time-consuming, and most frustrating. Where the rating and the mapping disagreed, the mapping won.

  • Pre-interview survey
  • 11 semi-structured interviews
  • Workflow mapping activity
  • Ranked burden exercise
  • Six-capability evaluation
  • Quantitative triangulation

Analysis and synthesis was the single largest phase of a researcher’s week. That is what made it the credible place to automate. Mean self-reported share of research time, n = 11. No other single phase came close.

Where a researcher’s week actually goes

Mean self-reported share of research time. n = 11.

  • 30.9%Analysis & synthesis — one phase, and the largest single one
  • 24.1%Data collection — one phase
  • 45.0%Planning, reporting & stakeholder alignment — three phases added together
The hatched share is the largest number on the chart, but it is three jobs rather than one. Analysis is the biggest thing a researcher does in a single stretch, and that is what made it the defensible place to automate first rather than the loudest complaint.
Where the friction actually concentrates The workflow mapping activity · 6 stages × 3 friction types

Each cell is the number of participants who attached that kind of friction to that stage.

Blank cells were not measured rather than measured at zero. Tool limitation was the fourth friction type in the activity, but it was never resolved to stage-level counts, so it is not shown. n = 11.
Friction type Planning Coding & tagging Affinity mapping Insight writing Report creation Stakeholder alignment
Cognitive overload 6 5 7 5
Tedious work 8 5
Stakeholder delay 7 7 8
Three different problems, and they do not sit in the same place. Tedium concentrates almost entirely in tagging, cognitive load in insight writing, and delay in alignment and reporting. That separation is the reason the recommendation targets classification rather than synthesis: they are not the same bottleneck.

What the evidence showed

  1. Manual classification is the clearest automation target

    8/11 called tagging tedious 7/11 would remove it entirely

    Coding and tagging was necessary but almost never valued as a human contribution. Researchers defended what the step protects and resented how it is executed: repetitive, slow, and vulnerable to fatigue at exactly the moment attention matters most.

    “It’s just the most mind numbing thing that a person can ever do in their life. If you’re a researcher, honestly, that’s it for me.”

    P08

    Why it matters When repetitive classification consumes a researcher’s attention, the process gets slower without getting more insightful.

  2. Researchers want a co-pilot, not an autonomous analyst

    9/11 required human review or ownership

    Current AI use ranged from never to daily, and the split was not about familiarity. Participants were evaluating whether a tool could accelerate a task while preserving their ability to review, correct, and shape the result.

    “I don’t want to put data in and it synthesize it for me and tell me what they said. I wanted it to aid me in my synthesis. That’s a difference to me.”

    P10

    Why it matters Faster work has no value when the cost of verifying it equals the cost of doing it by hand.

  3. Traceability is the mechanism that makes adoption possible

    9/11 rated evidence traceability medium or high value 4.27/5 adoptability

    Participants wanted every generated theme, count, and insight to stay connected to its underlying evidence. Traceability lowered the cost of validation and made stakeholder questions answerable without reconstructing the analysis.

    “Evidence traceability, super valuable. Nobody ever believes anything we say. So being able to prove it is great.”

    P02

    Why it matters Traceability is not a feature alongside automation. It is the precondition that lets automation be accepted at all.

From transcript to principle

The same traceability the study recommends, applied to the study itself. Each principle below can be walked back to the participant who produced it.

From findings to strategy

Pair automated tagging with evidence traceability. Ship neither alone. Six capabilities were scored on how much researchers wanted them and how ready they were to adopt them. Only two cleared both bars, and they clear them together. One saves the time; the other buys the trust that makes the saving acceptable.

Desirability against adoptability

Mean participant rating on a 1–5 scale, for the two capabilities scored on both axes. Gridlines mark each scale point.

  • Desirability — how much they wanted it
  • Adoptability — how ready they were to accept it
  • Automated tagging

    Wanted 4.73 · would adopt 4.82 · gap +0.09

  • Evidence traceability

    Wanted 3.73 · would adopt 4.27 · gap +0.54

Tagging sits near the top of both scales, so it is the easy call. Traceability is the interesting one: its gap is six times wider, because researchers never described it as exciting. They described it as the thing that would let them say yes.
The full build order, and what I would not build All six capabilities ranked · the bottom four are the useful part
  1. Automated tagging with evidence linksFastest path to measurable time saved, and the highest stated adoption readiness in the study.4.82 / 5
  2. Evidence traceability as shared infrastructureLowers verification cost for tagging, clustering, drafting, and reporting all at once.4.27 / 5
  3. Researcher-guided clusteringUseful as a starting point only when researchers control granularity and can reorganize the output.
  4. Insight draftingHelpful for first drafts and language, but requires context, attached evidence, and review.
  5. Storytelling and report supportUse for editing, compression, and structure. Not for owning the narrative.2.64 / 5
  6. Stakeholder rework automationLowest priority, because the underlying problem is changing inputs and organizational alignment, not document production.
Nine conditions a credible AI research workflow has to meet Participant requirements rewritten as constraints a tooling decision can fail
  • Review-first automationAI proposes; researchers approve, edit, or reject before anything propagates.
  • Evidence attached by defaultEvery theme, count, and statement carries source excerpts and coverage.
  • Researcher-controlled granularityUsers define the codebook, abstraction level, and what counts as sufficient support.
  • Context preservationSpeaker, sequence, study question, and tone cues survive the pipeline.
  • Correction as learningResearcher edits improve later suggestions without silently rewriting prior work.
  • Transparent uncertaintyLow-confidence classifications and thin findings are visibly flagged.
  • InteroperabilityOutput moves cleanly between transcript, repository, analysis board, and report.
  • Privacy and governanceSensitive data stays in approved environments with clear retention and audit.
  • Variable autonomyAutomation rises for repetitive work and falls as the task approaches interpretation.
The validation pilot, and the four gates it has to clear Four phases on one bounded study · any single failure sends it back
  • 01 Baseline Tag a representative subset by hand. Record time, agreement, and confidence.
  • 02 Parallel run Generate suggestions without showing them to the human coder. Compare after.
  • 03 Assisted run Let the researcher accept, edit, reject, and add while measuring correction effort.
  • 04 Downstream test Cluster and retrieve from approved tags. Verify whether errors propagate.

Continue only if all four hold. Any one of them failing sends the capability back rather than into production.

  • ≥ 30% Net time saved Measured after human review and correction, not raw model processing time.
  • ≥ 80% Tag agreement Against the researcher’s final code set, counting missed context and false positives.
  • < 10% High-severity misses Coverage of evidence, counterexamples, and participant counts.
  • ≥ 4 / 5 Researcher trust Confidence, willingness to reuse, and perceived loss of control.

What I would bring to your team

This is the trust architecture I apply to any AI research tooling decision: proposal before automation, evidence before acceptance, review before scale. I do not evaluate an AI capability on whether it is impressive. I evaluate it the way I would evaluate any product concept, on desirability and adoptability, and I insist on knowing which one is the constraint.