All templates
Developer & Engineering

SRE/DevOps Toil Measurement & Automation Gap Analysis

Quantifies toil sources, automation maturity, and incident-resolution quality for SRE, platform, and DevOps teams over a 30-day period. Use to benchmark reliability operations and prioritize tooling investments.

Sample questions

A preview of what’s in the template. Every question is editable before you launch.

27 questions · ~12 min
Q01
Message

Welcome to the SRE/DevOps Toil & Automation Survey. This survey asks about your experience with operational toil, automation, and reliability tooling over the last 30 days. It should take approximately 12 minutes to complete. Your participation is voluntary and you may stop at any time. There are no right or wrong answers—we are interested in your honest experience. All responses are confidential and will be reported only in aggregate. Please click next to begin.

Q02
Multiple Choice

What is your primary role?

  • SRE / Production Engineer
  • Platform / Infrastructure Engineer
  • Software Engineer
  • DevOps Engineer
  • Engineering Manager
  • Other (please specify)
Q03
Multiple Choice

In the last 30 days, which activities consumed the most of your working time? Select up to 3.

  • Project / feature work
  • Incident response / on-call
  • Maintenance / operations changes
  • CI/CD and deployments
  • Troubleshooting / bug fixing
  • Meetings / coordination
  • Documentation / runbooks
  • Repetitive manual tasks
Q04
Multiple Choice

Which tooling do you actively use to manage reliability and reduce toil? Select all that apply.

  • Alerting / Monitoring (e.g., Prometheus, Datadog)
  • Incident management (e.g., PagerDuty, Opsgenie)
  • Infrastructure as Code (e.g., Terraform, Pulumi)
  • Configuration management (e.g., Ansible, Chef)
  • CI/CD orchestration (e.g., Jenkins, GitHub Actions)
  • Feature flags / progressive delivery
  • SLO / Error budget tooling
  • Runbooks / ChatOps automation
  • Change management (e.g., ServiceNow)
  • Internal developer portal (e.g., Backstage)
  • Chaos / Resilience testing
  • None of the above
  • Other (please specify)
Q05
Dropdown

Roughly how many incidents with user impact did your team experience in the last 30 days?

  • 0
  • 1–2
  • 3–5
  • 6–10
  • 11–20
  • 21+
  • Not sure / Don't track
Q06
Long Text

What single tooling change would most reduce toil for your team?

Q07
Multiple Choice

How many years have you worked in this type of role?

  • 0–1
  • 2–4
  • 5–7
  • 8–10
  • 11+
Q08
Message

Thank you for completing this survey. Your input helps us track toil patterns and prioritize the right reliability tooling investments. Results will be shared in aggregate only.

Q09
Multiple Choice

How often do you take on-call rotations?

  • Never
  • Ad hoc / occasionally
  • Weekly
  • Every 2 weeks
  • Monthly
  • Less often than monthly
Q10
Dropdown

In the last 30 days, approximately how many hours per week did you spend on repetitive manual tasks?

  • 0 hours
  • 1–3 hours
  • 4–7 hours
  • 8–12 hours
  • 13–20 hours
  • More than 20 hours
Q11
Opinion Scale

Overall, how automated are your common operations tasks today?

Scale: 17
Min:Not at all automatedMax:Fully automated
Q12
Multiple Choice

Compared to 3 months ago, how has your median time to resolve incidents changed?

  • Improved (decreased)
  • About the same
  • Worsened (increased)
  • Not sure / Don't track
Q13
AI Interview

What are the biggest blockers to automating more of your operations work next quarter?

Q14
Multiple Choice

Approximately how large is your organization?

  • 1–49 employees
  • 50–249
  • 250–999
  • 1,000–4,999
  • 5,000–19,999
  • 20,000+
Q15
Multiple Choice

In the last 30 days, which were your main sources of toil? Select up to 5.

  • Noisy or flaky alerts
  • Manual deployments
  • Brittle CI/CD pipelines
  • Environment drift or config mismatch
  • Access or permissions requests
  • Manual change approvals
  • Capacity management chores
  • Ticket handoffs or coordination
  • Limited observability or telemetry gaps
  • Flaky tests
  • Rollback or roll-forward complexity
  • Data migrations or backfills
  • Tooling integrations or gaps
  • Other (please specify)
Q16
Opinion Scale

How effective are your current tools for monitoring and alerting?

Scale: 17
Min:Not at all effectiveMax:Extremely effective
Q17
Multiple Choice

During your most significant incident in the last 30 days, what added the most toil?

  • Paging noise or alert confusion
  • Manual runbook steps
  • Access or permissions delays
  • Coordination or hand-off overhead
  • Rollback or roll-forward complexity
  • Limited data or observability gaps
  • Change approvals or governance delays
  • No significant incidents in the last 30 days
Q18
Long Text

Based on your responses in this survey, please share any additional thoughts or feelings about toil, reliability, or tooling that we didn't cover.

Q19
Multiple Choice

Approximately how large is your SRE/Platform team?

  • 1
  • 2–5
  • 6–10
  • 11–20
  • 21+
Q20
Ranking

Rank the following by how disruptive they are to your focused engineering time (1 = most disruptive).

  1. Noisy alerts / pages
  2. Manual deployments
  3. Access / permissions requests
  4. Environment setup / configuration
  5. Manual change approvals
  6. Capacity / infrastructure changes
Drag to rank
Q21
Opinion Scale

How effective are your current tools for deployment and CI/CD?

Scale: 17
Min:Not at all effectiveMax:Extremely effective
Q22
Multiple Choice

Which region best describes your primary working time zone?

  • Americas
  • EMEA
  • APAC
  • Other / Multiple
Q23
Opinion Scale

How effective are your current tools for incident management and response?

Scale: 17
Min:Not at all effectiveMax:Extremely effective
Q24
Multiple Choice

What is your work location model?

  • Remote
  • Hybrid
  • Onsite
Q25
Opinion Scale

How effective are your current tools for infrastructure provisioning and configuration?

Scale: 17
Min:Not at all effectiveMax:Extremely effective
Q26
Opinion Scale

How effective are your current tools for change management and approvals?

Scale: 17
Min:Not at all effectiveMax:Extremely effective
Q27
Dropdown

Approximately how many manual steps did you automate or remove from runbooks in the last 30 days?

  • 0
  • 1–5
  • 6–15
  • 16–30
  • 31+

What’s included

  • AI follow-ups

    Adaptive probes on open-ended answers that pull out detail a static form would miss.

  • Attention checks

    Built-in safeguards against rushed answers and low-quality respondents.

  • AI-drafted copy

    Wording, ordering, and branching written by the AI — tuned to your research goal.

  • Auto report

    Themes, quotes, and a plain-English summary write themselves once responses come in.

Why this template

What this template is built to do — we found no directly comparable template from other survey tools to review.

What sets it apart

  • Includes an AI follow-up interview question that adaptively probes on operational blockers, something no static form can replicate
  • Covers toil sources, automation maturity across six distinct tool categories (monitoring, CI/CD, incident response, provisioning, change management), and incident trend data in one structured flow
  • Uses ranking and multi-select toil-source questions plus open-text fields to capture both quantifiable patterns and qualitative nuance
  • Closes with role, org size, team size, region, and work-model demographics so results can be benchmarked and segmented

Ready to launch?

Open this template in the editor. Every part is yours to change before the first respondent sees it.

Related templates

More studies from the same category.

See all
Developer & Engineering

API/SDK Migration Readiness & Blockers Assessment

Assess developer teams' API and SDK migration status, identify top blockers, and surface support needs to plan lower-risk, faster upgrades.

View template
Developer & Engineering

LLM Prompt Injection Awareness & Mitigation Practices Survey

Measures developer awareness of prompt injection threats, captures current security mitigation practices, and identifies gaps in LLM application defense. Designed for engineering teams building or evaluating LLM-integrated features.

View template
Developer & Engineering

OpenTelemetry Adoption & Readiness Assessment

Measures developer familiarity, adoption stage, blockers, and rollout priorities for OpenTelemetry across engineering teams to inform instrumentation strategy and resource planning.

View template
Developer & Engineering

Developer Open-Source License Compliance Experience Survey

Measures how developers navigate open-source license compliance, including confidence levels, tooling satisfaction, workflow clarity, and key barriers. Designed for engineering teams and developer-tool organizations seeking to improve compliance processes and SBOM adoption.

View template
Developer & Engineering

Developer Latency Sensitivity & SLO Benchmarking Survey

Measures developer-perceived latency thresholds, tail-latency tolerance, and performance trade-off priorities by use case. Use it to benchmark acceptable response times, set data-informed SLOs and SLAs, and prioritize performance investments that align with what developers actually care about.

View template
Developer & Engineering

Developer Documentation Findability & Navigation UX Survey

Evaluates how easily developers can find, navigate, and understand technical documentation. Measures discoverability, search quality, information architecture fit, and terminology clarity to prioritize documentation UX improvements.

View template