Choosing a model
How to pick the right AI model for your study.
QuestionPunk offers models from multiple AI providers including Anthropic Claude, OpenAI GPT, Google Gemini, and more. This guide explains how to choose the right model for your study.
Steps
- Understand the optionsQuestionPunk supports 150+ models across 17+ providers. New text AI interviews default to OpenAI GPT-5.4 Mini, while Claude Haiku 4.5 remains the generic and voice-interview default. Existing surveys keep their saved model.
Anthropic Claude: Claude Haiku 4.5 (fast and economical), Sonnet 4.6 and Sonnet 5 (balanced speed and depth), Opus 4.6 through Opus 4.8 (richest follow-ups for deep qualitative work), and Fable 5. Thinking-enabled variants are also available for extended reasoning.
OpenAI: GPT-5.5/Pro (latest flagship), GPT-5.4/Pro/Mini/Nano, plus GPT-5.3, 5.2, 5.1, and 5 series models, GPT-4.1 family, GPT-4o, o3, o3 Mini, o3 Pro, o4 Mini, gpt-oss open-weight models, and specialized Codex and Chat variants.
Google: Gemini 3.5 Flash, Gemini 3.1 Pro/Flash Preview, Gemini 3 Flash Preview, Gemini 2.5 Flash/Pro, plus Gemma 4 and Gemma 3 open models.
Meta: Llama 4 Maverick, Llama 4 Scout, Llama 3.3, and 3.2 models.
Mistral: Medium 3.5, Large 3, Medium 3.1, Small 4, Nemo, Ministral 3, Devstral, Codestral, and Saba models.
DeepSeek: V4 Pro/Flash, V3.2, V3.1, R1 series.
xAI: Grok 4 series.
Amazon: Nova Pro, Lite, and Micro.
Moonshot: Kimi K2 series.
Cohere: Command A.
Qwen, Nvidia, StepFun, Arcee, LiquidAI, Upstage, Zhipu, and others are also available, including free-tier and zero-cost options. - Select in survey settingsChoose the model in the AI interviewer settings for each interview question. You can set different models for different interview questions within the same survey.
- Pilot and compareRun a short A/B pilot with different model settings to compare output quality before sending to your full sample.
QuestionPunk supports 150+ models across 17+ providers including Anthropic, OpenAI, Google (Gemini), Meta (Llama), Mistral, DeepSeek, xAI (Grok), Amazon (Nova), Moonshot (Kimi), Cohere, Qwen, Nvidia, StepFun, Arcee, LiquidAI, Upstage, and Zhipu. New text AI interviews default to OpenAI GPT-5.4 Mini. Claude Haiku 4.5 remains the generic and voice-interview default, and existing surveys keep their saved model.
For most studies, Claude Haiku 4.5 or Sonnet 4.6 provides the best balance of quality and speed. Opus 4.6 through 4.8 deliver the richest, most nuanced follow-ups for complex qualitative research. Thinking-enabled model variants are available for tasks that benefit from extended reasoning.
Zero-cost models include Google Gemma 3/4 variants, DeepSeek R1, OpenAI gpt-oss (20B/120B), and Llama 3.3 70B. Many other open-weight models — including Llama 4, Mistral Small, and Nvidia Nemotron — are available at very low cost. Premium models from Anthropic and OpenAI offer the highest quality.
Run a short pilot with 5-10 respondents using different model settings to compare output quality before committing to your full sample.