Moderated usability testing, run by an AI moderator.
Guided tasks, think-aloud prompts, and screen share. The moderator watches, probes when an answer is thin, and never hints at the answer. It is not a human researcher — and this page is specific about where that matters.
You want to cancel the subscription on this account. Start from the homepage and do what you would normally do.
You paused there — what were you expecting to see?
Always available. Never going to redesign your study mid-session.
Unmoderated tests scale, then go quiet at the exact moment something surprising happens. Moderated tests catch the surprise, but each one costs an hour of a researcher’s day, so you get as many as fit in the week.
This sits between the two, and we would rather be precise about it than sell you a robot researcher. The AI moderator works through your task list, watches the shared screen, prompts people to think aloud, and probes a thin answer. What it will not do is notice the study is asking the wrong question, or that participant five is the wrong person entirely. That part is still yours.
“It will ask the follow-up you wrote down. It will not rewrite the study for you.”
Tasks in. Evidence out.
A usability test is a question type, not a separate product — so it sits in the same study as your screeners and your post-test questions.
Write the tasks
Scenarios, not instructions.
Give each task a goal and a starting URL. Add success criteria for your own review, plus a warm-up task if the flow needs one.
Run the session
Share screen, think aloud.
Respondents share their screen and talk through the task. The moderator observes, prompts, and probes without giving hints.
Review the evidence
Clips, not raw footage.
Each task returns its own clip, transcript, and ease rating, with synthesized pain points sitting on top of them.
Start from a task list someone already wrote.
Navigation, checkout, onboarding, account settings — editable scenarios with post-task questions and a final debrief.
What it actually does during a session.
The moderator runs your scenarios in order — counterbalanced across participants if you ask it to — and stays out of the way while someone works. When a participant goes quiet, it prompts them to narrate. When an answer is thin, it probes once and moves on.
It is briefed never to hint. A participant who cannot find the button is data; a participant who was told where the button is, is not.
- Reads your task scenarios in order, with optional counterbalancing
- Prompts for think-aloud when someone falls silent
- Probes a thin answer instead of accepting it
- Collects a post-task ease rating
- Never hints at the answer or the next step
Where screen share works, precisely.
On desktop, respondents share through the browser’s own screen capture — nothing to install. On Android, the QuestionPunk app uses native screen sharing. On iPhone, the session hands off to the QuestionPunk iOS app, which does the same thing natively.
If your study has to run on a particular surface, set the respondent surface policy in Settings — web only, app only, or both — before you distribute the link.
- Desktop: the browser’s screen capture, no install
- Android: native screen sharing in the QuestionPunk app
- iPhone: hands off to the QuestionPunk iOS app
- Surface policy in Settings decides where a study can be taken
Where you still want a human in the room.
Use the AI moderator for the parts that repeat: the same five tasks, run thirty times, transcribed and rated identically. That is the work that makes human moderation expensive and is the least interesting part of a researcher’s day.
Keep a human for the sessions where the next question depends on judgment you have not written down yet — and be honest with yourself about which of your studies those really are.
- Exploratory studies where the research question is still moving
- Sessions a stakeholder is watching live and will want to redirect
- Participants who need accommodation a script cannot anticipate
- Anything where the next probe depends on an unwritten judgment call
Evidence at the level of a single task.
Per-task video clips
Each task is cut into its own clip automatically, so the twelve seconds that prove the problem are already isolated.
Pain points with evidence
Synthesized problems across participants, each one traceable back to the sessions it came from.
Navigation analytics and ease ratings
Where people went, and how hard each task felt afterwards — the number that makes a pattern arguable.
Built for the people who watch the recordings.
Watch someone actually use it.Your first test is free.
Write the tasks, send the link, read the clips. Twenty free responses every month, no credit card, no sales call unless you ask for one.