
Testing AI-Assisted research synthesis
Context:
User Testing a concept design for an Internal Safety, Security & Wellbeing Platform—tested for comprehension, usability, and readability
Method:
- Method: Moderated usability testing
- Duration: 1 hour per participant
- Participants: 8, across Retail, Contact Centres, Home/Office, Corporate Units, Contractors, and In-Field/Infrastructure—a mix of users with and without experience using Risk Management features
My Role:
Lead UX researcher and designer. I conducted manual synthesis first, then used Miro’s AI synthesis feature to cross-check results and explore what AI-assisted synthesis could offer in practice.
Outcomes:
AI synthesis matched my manual read on themes, and beat it on speed and number-crunching. But it missed scale, depth, and emotional nuance—so I built four rules for using it on real projects:
- Tag before you synthesise
- Always ask for the count
- Prompt for details on each insight
- Read the raw notes yourself
Introduction
After completing moderated sessions with all 8 participants, I had a large volume of qualitative data to synthesise—notes from myself and two co-facilitators across 35 test scenarios.
I synthesised the data manually first, then ran the same data through Miro’s AI synthesis tool to compare results. The goal wasn’t to hand off the work to AI—it was to understand where AI adds value, and where it doesn’t.
What the AI did well
The AI’s insights were closely aligned with mine. It identified the same key themes and pain points, and it was particularly strong with numerical analysis — averaging ratings, calculating frequencies, and surfacing patterns across the dataset accurately and quickly.
For a dataset of this size, that kind of calculation would have taken considerably more time to do manually.
The AI’s insights were closely aligned with mine. It identified the same key themes and pain points, and it was particularly strong with numerical analysis — averaging ratings, calculating frequencies, and surfacing patterns across the dataset accurately and quickly.
For a dataset of this size, that kind of calculation would have taken considerably more time to do manually.
Correct Themes
The AI identified the same key pain points and patterns as manual synthesis — with no prompting on what to look for.
Correct Calculations
Averaging ratings and calculating frequencies across 35 scenarios accurately and fast—the kind of task that eats time in manual synthesis.
Where AI needed help
Accurate on themes and numbers, but the AI had clear limits. It stayed at surface level, missed context that was sitting right in the notes, and gave no indication of how many people were behind each finding. Left unchecked, those gaps can quietly skew your research.
Stayed on the surface
The AI identified themes but rarely dug into the detail behind them, even when that detail was in the data.
No sense of scale
It reported what users said, but not how many. One person’s comment looked the same as eight people’s.
Missed the emotional register
Patterns and pain points came through — but the texture of how participants felt didn’t.
Four things I learned
Those gaps aren’t reasons to avoid AI synthesis—they’re reasons to know how to work with it. Running manual and AI synthesis in parallel surfaced four practical lessons for getting better output and knowing where to stay hands-on.
AI can stay at surface level—prompt it to go deeper
The AI correctly identified that participants wanted “Next Steps” information, but it didn’t surface what those next steps were, even though that detail was in the notes. I had to go back and prompt for more. Treat AI output as a starting point, not a finished synthesis.
Insufficient data categorisation can give you blunt output—tag intentionally before you direct AI
I tagged every note by role, team, scenario, and participant ID before synthesis. This made it possible to ask the AI to compare Manager feedback against Worker feedback, or isolate findings from a specific scenario. Without structured tagging, you get undifferentiated output. With it, you get genuinely useful cuts of the data.
AI can skip sample size—always ask for percentages
The AI reported what users said without indicating how many users said it. One person’s comment can look like a pattern if it’s not qualified. Always ask the AI to include the percentage or count behind each insight—it changes how you weight them.
AI can’t build your empathy—read the raw notes yourself
Because I had two co-facilitators running some sessions, I didn’t hear every participant firsthand. Manually reading through their notes felt necessary — I needed to feel what participants were experiencing, not just read a summary of it.
My working hypothesis: if you conduct all the interviews yourself, you’ll build empathy through the sessions directly, and AI synthesis probably won’t cost you much. But if others are doing the research, make sure you’re reading the raw notes before you rely on AI output. This is something worth testing further.
Reflection
AI synthesis is fast, accurate on numbers, and good at pattern recognition across large datasets. It’s a legitimate tool for research — but it works best as a collaborator, not a replacement.
The human work that remains: knowing what questions to ask, reading for emotional nuance, and making sure the numbers tell the whole story.
