Experimentation & A/B Testing Maturity Assessment
Assesses experimentation program maturity across culture, process, tooling, governance, and outcomes. Designed for product, growth, and data teams to benchmark capabilities and identify improvement priorities.
Sample questions
A preview of what’s in the template. Every question is editable before you launch.
Which function best describes your primary role?
- Product management
- Growth / performance marketing
- Lifecycle / CRM
- Brand / creative marketing
- Data / analytics
- Engineering
- Design / UX
- Other (please specify)
Over the past 6 months, how would you rate the overall rigor of your team's experiment hypotheses?
Which experimentation tools or platforms does your team currently use? (Select all that apply)
- Optimizely
- VWO
- AB Tasty
- Statsig
- Eppo
- Amplitude Experiment
- LaunchDarkly or Flagsmith
- Google Optimize (legacy)
- In-house / custom platform
- None currently
- Other (please specify)
When deciding whether to ship a winning variant, rank these factors by importance to your team (most important first).
- Effect size vs. baseline
- Statistical significance or credible interval
- Impact on guardrail metrics
- Estimated business value
- Implementation cost / complexity
- Qualitative feedback / UX signals
Overall, how would you rate the maturity of experimentation in your organization today?
What are the biggest blockers or challenges to effective experimentation in your organization right now?
What is your seniority level?
- Individual contributor
- Manager
- Director
- VP
- C-level
- Other
Thank you for completing the Experimentation Maturity Assessment! Your responses will be analyzed in aggregate to produce benchmarking insights. If you opted in, results will be shared with participants once the analysis is complete. If you have any questions, please contact the research team at the email provided in your invitation.
Approximately how many people on your team are directly involved in experimentation?
- 1
- 2–5
- 6–10
- 11–20
- 21–50
- 51+
Our team documents a clear hypothesis for every experiment before launch.
How are experiment datasets integrated with your analytics and data warehouse?
- Fully integrated with analytics and warehouse
- Partial integration; some manual pulls required
- Isolated within the experimentation tool only
- I don't know
Which risk controls does your team typically apply to experiments? (Select all that apply)
- Guardrail metrics monitored
- Kill switches / instant rollback
- Ethics / privacy review when needed
- Traffic allocation caps
- Country / segment exclusions
- QA and instrumentation checklist
- None of the above
- Other (please specify)
Typically, how many business days elapse between a test ending and a final decision being made?
- Same day
- 1–2 days
- 3–5 days
- 6–10 days
- 11–20 days
- Over 20 days
- We don't track this
Based on your survey responses, we'd like to explore your experimentation challenges and aspirations in a bit more depth.
Approximately how many employees are in your company?
- 1–10
- 11–50
- 51–200
- 201–1,000
- 1,001–5,000
- 5,001–10,000
- 10,001+
In the last 90 days, approximately how many experiments did your team launch?
- 0
- 1–2
- 3–5
- 6–10
- 11–20
- 21+
We have a clear prioritization framework for deciding which experiments to run.
Do you have a defined and versioned metrics catalog for experiments?
- Yes, centrally defined and versioned
- Yes, team-specific only
- In progress
- No
Is there an experimentation council or governance body at your organization?
- Yes, org-wide
- Yes, within my business unit
- No, but being considered
- No
In the last 6 months, approximately what share of completed experiments led to a production rollout?
- 0–10%
- 11–25%
- 26–40%
- 41–60%
- 61–80%
- 81–100%
- We don't track this
Which industry best describes your organization?
- Consumer software
- B2B / SaaS
- E-commerce / retail
- Financial services / fintech
- Media / entertainment
- Healthcare / life sciences
- Gaming
- Telecom
- Travel / hospitality
- Other (please specify)
What are the primary objectives your experiments target? (Select up to 5)
- Conversion rate
- Retention / churn
- Engagement
- Monetization / revenue
- Activation / onboarding
- Acquisition / traffic
- Feature adoption
- Pricing / packaging
- Brand / creative effectiveness
- Learning about user behavior
- Other (please specify)
Experiment designs and analysis plans are peer-reviewed before launch.
How does your team typically determine sample size and test duration?
- Fixed-horizon power analysis
- Sequential testing / alpha spending
- Heuristics or benchmarks
- Vendor tool auto-calculates
- We usually don't calculate this
- I don't know
- Other (please specify)
Where are experiment plans and results typically documented? (Select all that apply)
- Central system of record
- Team wiki or docs
- Within the testing tool
- Spreadsheets
- Not consistently documented
- Other (please specify)
Where are you primarily based?
- North America
- Latin America
- Europe
- Middle East
- Africa
- Asia
- Oceania
Learnings from experiments are shared broadly and inform future decisions across teams.
How many years have you worked with experimentation or A/B testing?
- Less than 1
- 1–3
- 4–6
- 7–10
- 11+
Which test or study types does your team run regularly? (Select all that apply)
- A/B or split tests
- Multivariate tests (MVT)
- Holdout / control tests
- Quasi-experiments / observational studies
- Multi-armed bandits
- Sequential tests
- UX / usability studies
- Surveys / concept tests
- Feature-flag rollouts / experiments
- Other (please specify)
What is the typical runtime for a single experiment, from launch to decision?
- Same day
- 1–3 days
- 4–7 days
- 1–2 weeks
- 3–4 weeks
- Over 4 weeks
- Varies widely
Rank the following phases by where your team spends the most effort in a typical experiment (most effort first).
- Ideation / prioritization
- Design, UX, and copy
- Instrumentation and data quality
- Implementation / engineering
- QA and launch
- Monitoring during run
- Analysis and interpretation
- Documentation and sharing
- Rollout and follow-up
What’s included
AI follow-ups
Adaptive probes on open-ended answers that pull out detail a static form would miss.
Attention checks
Built-in safeguards against rushed answers and low-quality respondents.
AI-drafted copy
Wording, ordering, and branching written by the AI — tuned to your research goal.
Auto report
Themes, quotes, and a plain-English summary write themselves once responses come in.
Ready to launch?
Open this template in the editor. Every part is yours to change before the first respondent sees it.
Related templates
More studies from the same category.
Academic Researcher Experience & Support Survey
Measures how graduate students, postdocs, and faculty actually spend their research time, and how well funding, mentorship, protected time, and collaboration support their work. An AI follow-up interview digs into the single biggest obstacle each respondent names, reconstructing a concrete recent example instead of a vague complaint. Built for research offices, deans, and PIs benchmarking research support.
View templateWorkplace Climate & Psychological Safety Survey
Measures how safe, fair, and supported people feel in their day-to-day work — trust in leadership, psychological safety, workload fairness, and manager impact — for HR and people teams tracking team health between engagement cycles. An AI follow-up reconstructs the specific moment behind each person's rating instead of settling for a number on a dashboard.
View templateCompetitive Landscape & Switching Drivers Survey
Benchmarks how customers perceive your brand against named competitors — consideration, head-to-head ratings, and switching likelihood — then uses an AI follow-up interview to surface the real story behind a preference or switch. Built for product marketing, competitive intelligence, and strategy teams tracking share of consideration.
View templateResearch Informed Consent & Enrollment
A complete consent flow for research studies: plain-language study information, granular consent statements with a comprehension check, signature capture, and contact enrollment. Built for IRB-style requirements — participants confirm they understand, not just that they scrolled.
View templatePhoto Diary Study: Product Use in Context
A repeatable diary entry participants complete each time they use your product in real life: a photo of the moment, what they were trying to do, and what got in the way — with an AI probe on the day's friction. Run it daily or weekly to see usage where it actually happens, not in a lab.
View templateAI Moderator Effectiveness by Demographic Group
A post-experience research survey evaluating the effectiveness of AI interview moderators across demographic groups. Measures overall experience quality, perceived empathy and understanding, cultural sensitivity, accessibility, and direct comparisons with human moderators. Designed to detect demographic moderation effects.
View template