We built the AI that runs the interview.
Maze wanted user research without a researcher in the room. Diffco delivered the AI Moderator — the interview intelligence and the live participant experience it speaks through — inside Maze’s own product, on Maze’s stack.

Briefing
The Project At A Glance
Client Profile
Maze
User-research leader, $40M Series B led by Felicis, $60M total funding raised.
Active Scale
6M+ Participants
Trusted and deployed by enterprise teams at Atlassian, Volvo, Cisco, Lenovo, and Revolut.
Core Delivery
The AI Moderator
Adaptive autonomous interviews paired with real-time reactive participant WebGL canvas.
AI Engine Stack
Multimodal Core
OpenAI GPT · Claude Sonnet · Deepgram ASR · Google Gemini Evaluator.
Product Architecture
React & WebGL
Seamless Next.js container environment fully customized within the Maze local design system.
International Readiness
20 Languages
Active globally today, managing automated translation, transcription, and dialect analysis.
The bottleneck
Where they started
Maze sells user research to product teams. The bottleneck in that business is a human being: someone has to sit in the session, ask the next question, and know when to push. In Maze’s own research, 63% of product teams named time and bandwidth as their top research challenge.
Maze wanted to remove that constraint without giving up research rigor — interviews that adapt to what a participant actually says, stay anchored to the study’s goals, run in any language at any hour, and come out the other side as evidence a product team can act on.
That is a harder AI problem than a chat window. An interview is a real-time, turn-taking, goal-directed conversation with a stranger who can go off-topic, misunderstand, or stop talking. The intelligence has to be good, and the participant has to feel it working — otherwise they fill the silence, talk over it, or quit.

Human-moderated vs. Autonomous scaling
63% of product teams named time and bandwidth limitations as their single greatest research challenge. Diffco eliminated that constraint without compromising research rigor.
- 63%Time/Bandwidth challenge
- 6M+Participants in Maze’s research panel

Adaptive Logic
AI-moderated interviews that adapt in real time.
The moderator probes deeper when an answer is thin and moves on when the research goal is met. Every question maps back to what the study is trying to learn.

Goal MappingSystem Check
Stay on context
Every follow-up stays tied to what the study needs to learn — tangents get acknowledged and steered back.
Neutrality ShieldZero Bias
Check for bias
Questions avoid steering the participant toward a preferred answer — bias is checked on every turn.
Traceable SourceVerified
Source attribution
Themes and conclusions connect back to the participant’s own words as verifiable video and text evidence.
Resilient Infrastructure
A model layer built to be swapped, not bet on
OpenAI GPT and Claude Sonnet, via AWS Bedrock: two frontier vendors behind one interface. Interview quality varies by model and by workload; multi-model means the feature improves as the field does, without a rewrite.

Audio Processing
Speech that survives a real conversation
Deepgram ASR handles high-fidelity voice-to-text for live interviews. Real participants interrupt, trail off, use jargon, and switch languages mid-sentence — absolute transcription quality is what the entire downstream analysis rests on.


Evaluation
An evaluation harness, not a demo
Every single conversation is scored across 25 strict behavioral metrics — goal alignment, neutrality, optimal sequencing. We run continuous stress-testing against thousands of synthetic sessions.
25
Behavioral Quality Metrics
Evaluating interviewer neutrality, goal alignment, and conversation sequencing systematically.
2.5d
Median Time to Insight
Median, from study creation to insight.
35
Beta Customers
Ran the moderator in beta ahead of general availability.
Interaction Canvas
A live conversational interface, rebuilt in WebGL
With no human on the other side of the desk, the interface is the protocol. Diffco engineered a customized WebGL visualizer rendering real-time states (listening, thinking, responding) as fluid canvas graphics operating off the DOM main thread.

Infrastructure
The AI stack
| Reasoning Core | OpenAI GPT and Claude Sonnet, via AWS Bedrock | Two frontier vendors behind one interface. Interview quality varies by model and by workload; multi-model means the feature improves as the field does, without a rewrite. |
|---|---|---|
| Model Brokerage | AWS Bedrock | Enterprise data terms and regional control on the same cloud the product already runs on — a procurement requirement for Maze’s enterprise buyers, not an implementation detail. |
| Acoustics & Voice | Deepgram ASR | Purpose-built transcription for messy, real-time, multi-language human speech. Everything downstream — themes, quotes, reports — inherits its accuracy. |
| Evaluation Control | 25 quality metrics per conversation, plus Google Gemini in the monitoring loop | AI research tooling has to be defensible to the researcher using it. Scoring every live conversation is what makes “research-grade” a claim rather than a slogan. |
| Client Interface | WebGL | A continuously animating multi-state conversational canvas belongs off the DOM. It has to stay smooth while an interview is in flight. |
| Platform Delivery | Maze’s React and Next.js monorepo | AI features die in integration, not in the demo. We built inside their product — their conventions, review process, CI and preview environments — so it shipped rather than sat in a branch. |

Results
More conversations. Less research overhead. Faster answers to "why".
The AI Moderator is live and is now one of the features Maze leads with — sold on adaptive questioning, traceable quotes, synthesized themes and auto-generated reports.
20
Languages spoken fluently
35
Customers ran the moderator in beta ahead of general availability
Integrated Experience
The moderation environment plugs natively into the existing participant screen flow, ensuring ZERO friction for onboarding.
“It’s hard to think that something went wrong when we were able to get the 10 interviews without using any people hours.”









