Developer Synthetic Data Adoption & Ethics Survey
Measures developer experience, tooling preferences, risk perceptions, and adoption intent for synthetic data. Designed for engineering and data science teams evaluating synthetic data readiness and ethical boundaries.
샘플 질문
템플릿에 포함된 내용을 미리 확인해 보세요. 모든 질문은 설문 공개 전에 자유롭게 수정할 수 있습니다.
In the last 12 months, have you worked with synthetic data?
- Yes, regularly (monthly or more)
- Yes, occasionally
- No, but I am familiar with the concept
- No, and I am not familiar with it
Which tools or approaches have you used to generate synthetic data? (Select all that apply)
- In-house generation scripts
- Open-source libraries (e.g., SDV, SynthCity, ydata-synthetic)
- Vendor platform (e.g., Gretel, Mostly AI, Tonic)
- Data augmentation utilities not aimed at privacy
- Haven't used tools directly (consumed output from others)
- Other (please specify)
How much demonstrated fidelity and utility do you require before using synthetic data in production?
How likely are you to increase your use of synthetic data in the next 6 months?
Based on your responses in this survey, please share any additional thoughts about your limits, ideal use cases, or expectations for synthetic data.
Which best describes your primary role?
- Software engineer
- ML/AI engineer
- Data scientist
- Data/ML platform engineer
- Security/privacy engineer
- Product or engineering manager
- Researcher/academic
- Other (please specify)
Thank you for participating! Your input helps us understand practical needs and considerations around synthetic data. If you have any questions about this research, please contact [research team email].
<p>For this survey, <strong>synthetic data</strong> refers to artificially generated data (e.g., via simulations or generative models) intended to mimic real data's statistical properties while protecting sensitive information or filling gaps.</p>
Which use cases for synthetic data are most relevant to you or your team? (Select all that apply)
- Prototyping or training ML models
- Class imbalance augmentation
- Privacy-preserving sharing or compliance
- Testing and QA (e.g., edge cases, rare events)
- Synthetic logs or telemetry for load testing
- Analytics demos or sandboxing
- Education or training
- Other (please specify)
Approximately what percentage of data in your projects over the last 12 months was synthetic?
- 0% (none)
- 1–10%
- 11–25%
- 26–50%
- 51–75%
- 76–100%
- Not sure
<p>How appropriate is synthetic data for <strong>model training and development</strong> in your context?</p>
Which factors most limit your use of synthetic data today? (Select all that apply)
- Hard to evaluate quality or metrics
- Limited domain coverage
- Tooling or integration gaps
- Compute or cost constraints
- Stakeholder skepticism or buy-in
- Policy or legal uncertainty
- No clear need
- Other (please specify)
We'd like to explore a few of your responses in more depth. An AI moderator will ask you up to 2 brief follow-up questions based on what you've shared so far.
How many years of professional experience do you have in software or data roles?
- 0–1 years
- 2–4 years
- 5–9 years
- 10–14 years
- 15+ years
<p>How appropriate is synthetic data for <strong>production decision-making</strong> in your context?</p>
Rank the following improvements by how much they would accelerate synthetic data adoption in your organization, from most to least impactful.
- Better quality and validation metrics
- Broader domain and data type coverage
- Easier integration with existing pipelines
- Lower cost or compute requirements
- Clear policy and legal guidance or templates
- Independent benchmarks and case studies
- Training and best-practice playbooks
What is your primary domain or industry?
- Technology
- Finance/FinTech
- Healthcare/Life sciences
- Retail/Consumer
- Telecom/Media
- Manufacturing/Industrial
- Government/Public sector
- Education
- Other
<p>How appropriate is synthetic data for <strong>testing and QA</strong> in your context?</p>
What is your organization's approximate size (global headcount)?
- 1–9
- 10–49
- 50–249
- 250–999
- 1,000–4,999
- 5,000–19,999
- 20,000+
<p>How appropriate is synthetic data for <strong>external reporting or compliance submissions</strong> in your context?</p>
In which region do you primarily work?
- North America
- Latin America
- Europe
- Middle East
- Africa
- South Asia
- East Asia
- Southeast Asia
- Oceania
Rank your top concerns about synthetic data from most to least concerning.
- Privacy leakage or re-identification
- Bias amplification or fairness issues
- Poor realism or utility
- Regulatory or compliance risk
- Lack of transparency or traceability
- Leakage of secrets or intellectual property
If compliance or privacy is a priority in your work, briefly describe the data types or regulations you must satisfy (e.g., HIPAA, GDPR, PCI-DSS). If not applicable, you may skip this question.
포함된 기능
AI 후속 질문
정형화된 설문이 놓치는 세부 내용을, 주관식 답변에 맞춰 AI가 심층 질문으로 끌어냅니다.
주의력 확인 장치
성의 없는 답변과 저품질 응답자를 걸러내는 내장 안전장치입니다.
AI가 작성한 문안
문구, 질문 순서, 분기 로직까지 AI가 연구 목표에 맞춰 작성합니다.
자동 리포트
응답이 모이면 주요 주제, 인용문, 이해하기 쉬운 요약이 자동으로 작성됩니다.
설문을 공개할 준비가 되셨나요?
이 템플릿을 편집기에서 열어 보세요. 첫 응답자가 보기 전에 모든 부분을 원하는 대로 바꿀 수 있습니다.
관련 템플릿
같은 카테고리의 다른 설문을 만나 보세요.
Workplace AI Adoption & Compliance Assessment
Measures employee AI tool usage patterns, shadow AI risks, policy awareness, and training needs to inform governance and safe-adoption strategies across the organization.
템플릿 보기AI Refusal Message Clarity & Tone Evaluation
A stimulus-comparison survey for UX researchers and AI product teams to evaluate the clarity, tone, and helpfulness of AI safety and refusal messages. Produces actionable data on user preferences and improvement priorities.
템플릿 보기Shared Prompt Library: Discovery & Reuse Experience Survey
Assesses how users find, customize, and derive value from a shared AI prompt library. Use this to identify discovery friction, reuse patterns, and outcome perceptions to prioritize product improvements.
템플릿 보기AI Error Reporting Friction & Trust Impact Survey
Measures how AI users experience error reporting workflows and how unresolved issues affect trust and future reporting intent. Designed for product and UX teams seeking to reduce reporting friction and improve AI reliability perceptions.
템플릿 보기온디바이스 AI 학습: 소비자 신뢰 및 프라이버시 인식
온디바이스 AI 학습에 대한 소비자 인지도, 신뢰, 프라이버시 우려, 채택 의향을 측정합니다. 사용자가 로컬 AI 학습 기능을 어떻게 인식하고 평가하는지 이해하고자 하는 제품, UX, 프라이버시 팀을 위해 설계되었습니다.
템플릿 보기Developer Content Filter False Positive Impact Assessment
Assess how content filter false positives affect developer productivity, workflow disruption, and tool adoption decisions. Designed for developer experience researchers and tooling teams seeking actionable improvement priorities from software practitioners.
템플릿 보기