SRE/DevOps Toil Measurement & Automation Gap Analysis
Quantifies toil sources, automation maturity, and incident-resolution quality for SRE, platform, and DevOps teams over a 30-day period. Use to benchmark reliability operations and prioritize tooling investments.
Sample questions
A preview of what’s in the template. Every question is editable before you launch.
What is your primary role?
- SRE / Production Engineer
- Platform / Infrastructure Engineer
- Software Engineer
- DevOps Engineer
- Engineering Manager
- Other (please specify)
In the last 30 days, which activities consumed the most of your working time? Select up to 3.
- Project / feature work
- Incident response / on-call
- Maintenance / operations changes
- CI/CD and deployments
- Troubleshooting / bug fixing
- Meetings / coordination
- Documentation / runbooks
- Repetitive manual tasks
Which tooling do you actively use to manage reliability and reduce toil? Select all that apply.
- Alerting / Monitoring (e.g., Prometheus, Datadog)
- Incident management (e.g., PagerDuty, Opsgenie)
- Infrastructure as Code (e.g., Terraform, Pulumi)
- Configuration management (e.g., Ansible, Chef)
- CI/CD orchestration (e.g., Jenkins, GitHub Actions)
- Feature flags / progressive delivery
- SLO / Error budget tooling
- Runbooks / ChatOps automation
- Change management (e.g., ServiceNow)
- Internal developer portal (e.g., Backstage)
- Chaos / Resilience testing
- None of the above
- Other (please specify)
Roughly how many incidents with user impact did your team experience in the last 30 days?
- 0
- 1–2
- 3–5
- 6–10
- 11–20
- 21+
- Not sure / Don't track
What single tooling change would most reduce toil for your team?
How many years have you worked in this type of role?
- 0–1
- 2–4
- 5–7
- 8–10
- 11+
Thank you for completing this survey. Your input helps us track toil patterns and prioritize the right reliability tooling investments. Results will be shared in aggregate only.
How often do you take on-call rotations?
- Never
- Ad hoc / occasionally
- Weekly
- Every 2 weeks
- Monthly
- Less often than monthly
In the last 30 days, approximately how many hours per week did you spend on repetitive manual tasks?
- 0 hours
- 1–3 hours
- 4–7 hours
- 8–12 hours
- 13–20 hours
- More than 20 hours
Overall, how automated are your common operations tasks today?
Compared to 3 months ago, how has your median time to resolve incidents changed?
- Improved (decreased)
- About the same
- Worsened (increased)
- Not sure / Don't track
What are the biggest blockers to automating more of your operations work next quarter?
Approximately how large is your organization?
- 1–49 employees
- 50–249
- 250–999
- 1,000–4,999
- 5,000–19,999
- 20,000+
In the last 30 days, which were your main sources of toil? Select up to 5.
- Noisy or flaky alerts
- Manual deployments
- Brittle CI/CD pipelines
- Environment drift or config mismatch
- Access or permissions requests
- Manual change approvals
- Capacity management chores
- Ticket handoffs or coordination
- Limited observability or telemetry gaps
- Flaky tests
- Rollback or roll-forward complexity
- Data migrations or backfills
- Tooling integrations or gaps
- Other (please specify)
How effective are your current tools for monitoring and alerting?
During your most significant incident in the last 30 days, what added the most toil?
- Paging noise or alert confusion
- Manual runbook steps
- Access or permissions delays
- Coordination or hand-off overhead
- Rollback or roll-forward complexity
- Limited data or observability gaps
- Change approvals or governance delays
- No significant incidents in the last 30 days
Based on your responses in this survey, please share any additional thoughts or feelings about toil, reliability, or tooling that we didn't cover.
Approximately how large is your SRE/Platform team?
- 1
- 2–5
- 6–10
- 11–20
- 21+
Rank the following by how disruptive they are to your focused engineering time (1 = most disruptive).
- Noisy alerts / pages
- Manual deployments
- Access / permissions requests
- Environment setup / configuration
- Manual change approvals
- Capacity / infrastructure changes
How effective are your current tools for deployment and CI/CD?
Which region best describes your primary working time zone?
- Americas
- EMEA
- APAC
- Other / Multiple
How effective are your current tools for incident management and response?
What is your work location model?
- Remote
- Hybrid
- Onsite
How effective are your current tools for infrastructure provisioning and configuration?
How effective are your current tools for change management and approvals?
Approximately how many manual steps did you automate or remove from runbooks in the last 30 days?
- 0
- 1–5
- 6–15
- 16–30
- 31+
What’s included
AI follow-ups
Adaptive probes on open-ended answers that pull out detail a static form would miss.
Attention checks
Built-in safeguards against rushed answers and low-quality respondents.
AI-drafted copy
Wording, ordering, and branching written by the AI — tuned to your research goal.
Auto report
Themes, quotes, and a plain-English summary write themselves once responses come in.
Why this template
What this template is built to do — we found no directly comparable template from other survey tools to review.
What sets it apart
- Includes an AI follow-up interview question that adaptively probes on operational blockers, something no static form can replicate
- Covers toil sources, automation maturity across six distinct tool categories (monitoring, CI/CD, incident response, provisioning, change management), and incident trend data in one structured flow
- Uses ranking and multi-select toil-source questions plus open-text fields to capture both quantifiable patterns and qualitative nuance
- Closes with role, org size, team size, region, and work-model demographics so results can be benchmarked and segmented
Ready to launch?
Open this template in the editor. Every part is yours to change before the first respondent sees it.
Related templates
More studies from the same category.
API/SDK Migration Readiness & Blockers Assessment
Assess developer teams' API and SDK migration status, identify top blockers, and surface support needs to plan lower-risk, faster upgrades.
View templateLLM Prompt Injection Awareness & Mitigation Practices Survey
Measures developer awareness of prompt injection threats, captures current security mitigation practices, and identifies gaps in LLM application defense. Designed for engineering teams building or evaluating LLM-integrated features.
View templateOpenTelemetry Adoption & Readiness Assessment
Measures developer familiarity, adoption stage, blockers, and rollout priorities for OpenTelemetry across engineering teams to inform instrumentation strategy and resource planning.
View templateDeveloper Open-Source License Compliance Experience Survey
Measures how developers navigate open-source license compliance, including confidence levels, tooling satisfaction, workflow clarity, and key barriers. Designed for engineering teams and developer-tool organizations seeking to improve compliance processes and SBOM adoption.
View templateDeveloper Latency Sensitivity & SLO Benchmarking Survey
Measures developer-perceived latency thresholds, tail-latency tolerance, and performance trade-off priorities by use case. Use it to benchmark acceptable response times, set data-informed SLOs and SLAs, and prioritize performance investments that align with what developers actually care about.
View templateDeveloper Documentation Findability & Navigation UX Survey
Evaluates how easily developers can find, navigate, and understand technical documentation. Measures discoverability, search quality, information architecture fit, and terminology clarity to prioritize documentation UX improvements.
View template