Voice AI interviews

Conduct voice-based AI interviews with real-time transcription.

AI audio interviews let respondents speak their answers instead of typing. The AI conducts a real-time voice conversation with adaptive follow-ups and automatic transcription.

aiaudiovoiceinterviews10-20 minutes to set upIntermediateResearchersUX teams

Steps

  1. Add a voice AI interview or Usability Test
    In the survey editor, click + Add, choose AI interview, and select Voice under How should it be conducted?. Choose Usability Test instead for a guided task-based session with screen sharing.
  2. Configure voice settings
    Set the interview topic, AI personality, and prompt. By default a voice interview chains speech-to-text, an AI model, and text-to-speech, which lets you pick each provider. Turn on Ultra-low latency to use a single realtime model instead (fewer languages, no voice choice). New voice interviews use Claude Sonnet 5 as the conversation model by default; existing surveys keep their saved model. Voice interviews run for a set time: choose a duration preset (1, 2, 3, 5, 10, 15, 20, 30, 45, or 60 minutes) or enter a custom length from 1 to 240 minutes.
  3. Review transcripts
    After respondents complete the voice interview, transcripts are automatically generated and available in the Results tab alongside the audio recordings.

AI audio interviews remove the friction of typing, making it easier for respondents to share detailed, nuanced feedback through natural conversation.

Two setups are available: the default chain (separate speech-to-text, AI model, and text-to-speech for maximum flexibility) and Ultra-low latency (a single realtime model for the fastest replies).

Text-to-speech providers: Cartesia Sonic-3/3.5 (recommended, ~40ms latency, multilingual), ElevenLabs (v3, Flash v2.5, Turbo, Multilingual), Deepgram Aura-2, Rime Arcana (fast English with expressive voices), xAI TTS (20 languages, inline speech tags), Inworld TTS (emotional range for interactive contexts), and more. With Ultra-low latency on, the realtime model handles speech itself.

Speech-to-text providers: Deepgram Flux (recommended for voice agents, integrated turn detection), Cartesia Ink Whisper (fastest, 80+ languages), AssemblyAI Universal-3 Pro (promptable for domain-specific terms), Deepgram Nova-3 (best accuracy at 6.84% WER, 36 languages), and ElevenLabs Scribe V2 Realtime (99+ languages with word-level timestamps).

Avatar support: Add a visual AI avatar to voice interviews using Tavus or LemonSlice. Avatars give respondents a face to talk to, increasing engagement for longer interviews. Choose from the preset avatars available in the voice settings.

Guided usability testing: AI audio interviews support a guided usability testing mode where respondents complete step-by-step tasks on a website or app while the AI moderator observes via screen share. Configure task scenarios, time budgets, warmup tasks, think-aloud prompts, and counterbalanced task ordering. See the usability testing guide for full setup details.

Pricing: The default voice setup costs about 54 AI credits (roughly 5 cents) a minute in provider fees, plus QuestionPunk's margin, billed per minute with a 30-second minimum. Other voices and realtime models cost more.

Built-in audio enhancement (ai-coustics) applies voice focus and speaker isolation to improve transcription accuracy in noisy environments. Short noises and one-word interjections like coughs or "mm-hm" don't cut the interviewer off.

Transcripts are automatically generated for every voice interview, making analysis straightforward even with audio data.