# AI audio interviews

Canonical page: https://www.questionpunk.com/support/ai-audio-interviews

Conduct voice-based AI interviews with real-time transcription.

Time: 10-20 minutes to set up

Level: Intermediate

Audience: Researchers, UX teams

## Steps

### 1. Add a voice AI interview or Usability Test

In the survey editor, click **+ Add**, choose **AI interview**, and select **Voice** under **How should it be conducted?**. Choose **Usability Test** instead for a guided task-based session with screen sharing.

### 2. Configure voice settings

Set the interview topic, AI personality, and system prompt. Choose a session mode: **Pipeline** (separate STT, LLM, and TTS components for flexibility and cost control) or **Realtime** (native multimodal models from OpenAI or Google Gemini for the lowest latency). New Pipeline voice interviews use **Claude Sonnet 5** as the conversation model by default; existing surveys keep their saved model. Select your preferred TTS and STT providers. For time-based voice interviews, choose from **duration presets** (1, 2, 3, 5, 10, 15, 20, 30, 45, or 60 minutes) or enter a custom duration.

### 3. Review transcripts

After respondents complete the voice interview, transcripts are automatically generated and available in the Results tab alongside the audio recordings.

## Details

AI audio interviews remove the friction of typing, making it easier for respondents to share detailed, nuanced feedback through natural conversation.

Two session modes are available: **Pipeline** (separate STT + LLM + TTS components for maximum flexibility) and **Realtime** (native multimodal models from OpenAI or Google for lowest latency).

**Text-to-speech providers:** Cartesia Sonic-3/3.5 (recommended, ~40ms latency, multilingual), ElevenLabs (v3, Flash v2.5, Turbo, Multilingual), Deepgram Aura-2, Rime Arcana (fast English with expressive voices), xAI TTS (20 languages, inline speech tags), Inworld TTS (emotional range for interactive contexts), and more. In Realtime mode, OpenAI and Google Gemini handle speech natively.

**Speech-to-text providers:** Deepgram Flux (recommended for voice agents, integrated turn detection), Cartesia Ink Whisper (fastest, 80+ languages), AssemblyAI Universal-3 Pro (promptable for domain-specific terms), Deepgram Nova-3 (best accuracy at 6.84% WER, 36 languages), and ElevenLabs Scribe V2 Realtime (99+ languages with word-level timestamps).

**Avatar support:** Add a visual AI avatar to voice interviews using Tavus or LemonSlice. Avatars give respondents a face to talk to, increasing engagement for longer interviews. Choose from the preset avatars available in the voice settings.

**Guided usability testing:** AI audio interviews support a [guided usability testing mode](https://www.questionpunk.com/support/usability-testing) where respondents complete step-by-step tasks on a website or app while the AI moderator observes via screen share. Configure task scenarios, time budgets, warmup tasks, think-aloud prompts, and counterbalanced task ordering. See the [usability testing guide](https://www.questionpunk.com/support/usability-testing) for full setup details.

**Pricing:** Pipeline mode costs approximately $0.05/minute. OpenAI Realtime runs at approximately $0.30/minute, and Google Gemini Realtime at approximately $0.08/minute.

Built-in audio enhancement (ai-coustics) applies voice focus and speaker isolation to improve transcription accuracy in noisy environments. Adaptive interruption handling uses ML to reject false barge-ins like coughs and back-channeling.

Transcripts are automatically generated for every voice interview, making analysis straightforward even with audio data.

## Related articles

- [How AI interviews work](https://www.questionpunk.com/support/ai-how-it-works.md)
- [Choosing a model](https://www.questionpunk.com/support/ai-choosing-a-model.md)
- [Usability testing with screen share](https://www.questionpunk.com/support/usability-testing.md)
