Developer Synthetic Data Adoption & Ethics Survey
Measures developer experience, tooling preferences, risk perceptions, and adoption intent for synthetic data. Designed for engineering and data science teams evaluating synthetic data readiness and ethical boundaries.
Sample questions
A preview of what’s in the template. Every question is editable before you launch.
In the last 12 months, have you worked with synthetic data?
- Yes, regularly (monthly or more)
- Yes, occasionally
- No, but I am familiar with the concept
- No, and I am not familiar with it
Which tools or approaches have you used to generate synthetic data? (Select all that apply)
- In-house generation scripts
- Open-source libraries (e.g., SDV, SynthCity, ydata-synthetic)
- Vendor platform (e.g., Gretel, Mostly AI, Tonic)
- Data augmentation utilities not aimed at privacy
- Haven't used tools directly (consumed output from others)
- Other (please specify)
How much demonstrated fidelity and utility do you require before using synthetic data in production?
How likely are you to increase your use of synthetic data in the next 6 months?
Based on your responses in this survey, please share any additional thoughts about your limits, ideal use cases, or expectations for synthetic data.
Which best describes your primary role?
- Software engineer
- ML/AI engineer
- Data scientist
- Data/ML platform engineer
- Security/privacy engineer
- Product or engineering manager
- Researcher/academic
- Other (please specify)
Thank you for participating! Your input helps us understand practical needs and considerations around synthetic data. If you have any questions about this research, please contact [research team email].
<p>For this survey, <strong>synthetic data</strong> refers to artificially generated data (e.g., via simulations or generative models) intended to mimic real data's statistical properties while protecting sensitive information or filling gaps.</p>
Which use cases for synthetic data are most relevant to you or your team? (Select all that apply)
- Prototyping or training ML models
- Class imbalance augmentation
- Privacy-preserving sharing or compliance
- Testing and QA (e.g., edge cases, rare events)
- Synthetic logs or telemetry for load testing
- Analytics demos or sandboxing
- Education or training
- Other (please specify)
Approximately what percentage of data in your projects over the last 12 months was synthetic?
- 0% (none)
- 1–10%
- 11–25%
- 26–50%
- 51–75%
- 76–100%
- Not sure
<p>How appropriate is synthetic data for <strong>model training and development</strong> in your context?</p>
Which factors most limit your use of synthetic data today? (Select all that apply)
- Hard to evaluate quality or metrics
- Limited domain coverage
- Tooling or integration gaps
- Compute or cost constraints
- Stakeholder skepticism or buy-in
- Policy or legal uncertainty
- No clear need
- Other (please specify)
We'd like to explore a few of your responses in more depth. An AI moderator will ask you up to 2 brief follow-up questions based on what you've shared so far.
How many years of professional experience do you have in software or data roles?
- 0–1 years
- 2–4 years
- 5–9 years
- 10–14 years
- 15+ years
<p>How appropriate is synthetic data for <strong>production decision-making</strong> in your context?</p>
Rank the following improvements by how much they would accelerate synthetic data adoption in your organization, from most to least impactful.
- Better quality and validation metrics
- Broader domain and data type coverage
- Easier integration with existing pipelines
- Lower cost or compute requirements
- Clear policy and legal guidance or templates
- Independent benchmarks and case studies
- Training and best-practice playbooks
What is your primary domain or industry?
- Technology
- Finance/FinTech
- Healthcare/Life sciences
- Retail/Consumer
- Telecom/Media
- Manufacturing/Industrial
- Government/Public sector
- Education
- Other
<p>How appropriate is synthetic data for <strong>testing and QA</strong> in your context?</p>
What is your organization's approximate size (global headcount)?
- 1–9
- 10–49
- 50–249
- 250–999
- 1,000–4,999
- 5,000–19,999
- 20,000+
<p>How appropriate is synthetic data for <strong>external reporting or compliance submissions</strong> in your context?</p>
In which region do you primarily work?
- North America
- Latin America
- Europe
- Middle East
- Africa
- South Asia
- East Asia
- Southeast Asia
- Oceania
Rank your top concerns about synthetic data from most to least concerning.
- Privacy leakage or re-identification
- Bias amplification or fairness issues
- Poor realism or utility
- Regulatory or compliance risk
- Lack of transparency or traceability
- Leakage of secrets or intellectual property
If compliance or privacy is a priority in your work, briefly describe the data types or regulations you must satisfy (e.g., HIPAA, GDPR, PCI-DSS). If not applicable, you may skip this question.
What’s included
AI follow-ups
Adaptive probes on open-ended answers that pull out detail a static form would miss.
Attention checks
Built-in safeguards against rushed answers and low-quality respondents.
AI-drafted copy
Wording, ordering, and branching written by the AI — tuned to your research goal.
Auto report
Themes, quotes, and a plain-English summary write themselves once responses come in.
Ready to launch?
Open this template in the editor. Every part is yours to change before the first respondent sees it.
Related templates
More studies from the same category.
Workplace AI Adoption & Compliance Assessment
Measures employee AI tool usage patterns, shadow AI risks, policy awareness, and training needs to inform governance and safe-adoption strategies across the organization.
View templateAI Refusal Message Clarity & Tone Evaluation
A stimulus-comparison survey for UX researchers and AI product teams to evaluate the clarity, tone, and helpfulness of AI safety and refusal messages. Produces actionable data on user preferences and improvement priorities.
View templateShared Prompt Library: Discovery & Reuse Experience Survey
Assesses how users find, customize, and derive value from a shared AI prompt library. Use this to identify discovery friction, reuse patterns, and outcome perceptions to prioritize product improvements.
View templateAI Error Reporting Friction & Trust Impact Survey
Measures how AI users experience error reporting workflows and how unresolved issues affect trust and future reporting intent. Designed for product and UX teams seeking to reduce reporting friction and improve AI reliability perceptions.
View templateOn-Device AI Training: Consumer Trust & Privacy Perceptions
Measures consumer awareness, trust, privacy concerns, and adoption intentions regarding on-device AI training. Designed for product, UX, and privacy teams seeking to understand how users perceive and evaluate local AI learning features.
View templateDeveloper Content Filter False Positive Impact Assessment
Assess how content filter false positives affect developer productivity, workflow disruption, and tool adoption decisions. Designed for developer experience researchers and tooling teams seeking actionable improvement priorities from software practitioners.
View template