RAG System Quality & Grounding Assessment (Developer Survey)
Evaluates developer experiences with Retrieval-Augmented Generation systems across retrieval quality, grounding accuracy, evaluation practices, and infrastructure. Designed for ML engineers, backend engineers, and data scientists actively building or maintaining RAG pipelines.
샘플 질문
템플릿에 포함된 내용을 미리 확인해 보세요. 모든 질문은 설문 공개 전에 자유롭게 수정할 수 있습니다.
Have you built or maintained a RAG system in the last 6 months?
- Yes
- No
- Not sure
Which role best describes your day-to-day work?
- ML engineer
- Backend engineer
- Data scientist
- MLOps/Platform
- Product engineer
- Researcher
- Architect/Tech lead
- Other
Which content sources feed your retriever today? Select all that apply.
- Proprietary documents
- Code repositories
- Product knowledge base
- Web crawl
- Vendor API docs
- Slack/Chat logs
- Support tickets
- Wiki/Confluence
- Database/warehouse
- Not applicable
- Other
How are model answers grounded or cited in your RAG system? Select all that apply.
- Inline citations with URLs
- Inline citations with document IDs
- Evidence block after the answer
- Tool outputs included verbatim
- Structured JSON evidence list
- No grounding/citations
- Other
Which evaluation tools or libraries do you use for RAG? Select all that apply.
- Ragas
- TruLens
- DeepEval
- Promptfoo
- Custom harness
- LlamaIndex evals
- None
- Other
What is the primary programming language you use for RAG development?
- Python
- JavaScript/TypeScript
- Java
- Go
- C
- Rust
- Other
How critical is retrieval quality to the overall success of your RAG system?
How many years of professional experience do you have in software, data, or ML?
- 0–1
- 2–4
- 5–9
- 10–14
- 15+
- Prefer not to say
Thank you for completing this survey! Your input is valuable and will help improve RAG systems and developer tooling. All results will be reported in aggregate only.
In the last 30 days, how well did retrieved context meet your task requirements?
Over the last 30 days, how much do you trust the correctness of cited evidence in your RAG system's outputs?
Which metrics best reflect your RAG quality today? Select all that apply.
- Precision@k
- Recall@k
- MRR
- nDCG
- Answer faithfulness
- Context precision/recall
- Groundedness score
- Human ratings
- Production usage signals
- Custom internal metrics
What is your primary vector store or retriever backend?
- Pinecone
- Weaviate
- Milvus
- FAISS
- Elasticsearch/OpenSearch
- pgvector
- Chroma
- Vespa
- Not applicable
- Other
How satisfied are you with your RAG system overall today?
In which region do you primarily work?
- Africa
- Asia
- Europe
- North America
- Oceania
- South America
- Prefer not to say
How do you set or tune top-k and related retrieval parameters?
- Manual experimentation
- Grid/Random search
- Bayesian optimization
- Vendor auto-tuning
- Learned retrieval policy
- Not tuned
- Other
Rank your top 3 preferred citation/grounding display styles.
- Inline per sentence
- Numbered endnotes
- Collapsible evidence panel
- Top-k sources with scores
- Link to full passages
- Show only on demand
If you use a custom metric, please briefly describe it and how you compute it.
Which embedding model is your primary choice?
- OpenAI text-embedding-3
- OpenAI small embedding
- Cohere Embed
- VoyageAI
- Jina embeddings
- E5 family
- Instructor
- BGE family
- Local model
- Other
Rank your top 3 priorities for improving your RAG system in the next 3 months.
- Improve retrieval precision/recall
- Better grounding/citations
- Reduce latency
- Lower cost per query
- Scale to more data sources
- Harden evaluation pipeline
- Security/compliance
- Developer ergonomics
What is the approximate size of your organization (number of employees)?
- 1–10
- 11–50
- 51–200
- 201–1,000
- 1,001–5,000
- 5,001+
- Prefer not to say
In the last 30 days, how often did you encounter irrelevant or off-topic passages in retrieval results?
In the last 30 days, how frequently did you observe hallucinations in your RAG outputs despite grounding?
How automated is your evaluation workflow?
- None (manual only)
- Some scripts
- CI-integrated checks
- Continuous eval in production
Do you use a reranker after initial retrieval?
- Yes
- No
- Experimenting
Based on your responses in this survey, please share any additional thoughts or experiences about your RAG retrieval or grounding challenges.
What is the primary industry or domain for your RAG work?
- Technology
- Finance
- Healthcare/Life sciences
- Retail/CPG
- Education
- Government/Public sector
- Manufacturing
- Media/Entertainment
- Other
- Prefer not to say
In the last 30 days, how often did you encounter missing context (key information not retrieved)?
Please describe a recent grounding failure you encountered and its impact on your work.
How often do you run RAG benchmarks?
- Before each release
- Weekly
- Biweekly
- Monthly
- Quarterly
- Ad hoc only
Which reranker do you use most often?
- Cohere Rerank
- Voyage Rerank
- Jina Reranker
- Cross-encoder (e.g., MS MARCO)
- Self-hosted reranker
- Other
In the last 30 days, how often did you encounter stale or outdated content in retrieval results?
We'd like to understand more about your experience with grounding and citation quality. An AI moderator will ask you a couple of follow-up questions.
What is your end-to-end RAG latency target per query?
- < 200 ms
- 200–500 ms
- 500 ms–1 s
- 1–2 s
- 2–5 s
- > 5 s
- No specific target
In the last 30 days, how often did you encounter duplicate or near-duplicate chunks in retrieval results?
포함된 기능
AI 후속 질문
정형화된 설문이 놓치는 세부 내용을, 주관식 답변에 맞춰 AI가 심층 질문으로 끌어냅니다.
주의력 확인 장치
성의 없는 답변과 저품질 응답자를 걸러내는 내장 안전장치입니다.
AI가 작성한 문안
문구, 질문 순서, 분기 로직까지 AI가 연구 목표에 맞춰 작성합니다.
자동 리포트
응답이 모이면 주요 주제, 인용문, 이해하기 쉬운 요약이 자동으로 작성됩니다.
설문을 공개할 준비가 되셨나요?
이 템플릿을 편집기에서 열어 보세요. 첫 응답자가 보기 전에 모든 부분을 원하는 대로 바꿀 수 있습니다.
관련 템플릿
같은 카테고리의 다른 설문을 만나 보세요.
API 속도 제한 공정성 및 생산성 영향 설문조사
API 속도 제한의 공정성, 문서의 명확성, 생산성 영향, 대응 전략에 대한 개발자의 인식을 측정합니다. 스로틀링 정책과 요금제 등급을 개선하기 위한 실행 가능한 데이터를 찾는 API 제품 팀과 개발자 경험 연구자를 위해 설계되었습니다.
템플릿 보기Refund Experience & Policy Satisfaction Survey
Measures customer satisfaction with refund policies, processing speed, communication, and perceived effort. Designed for consumer research teams seeking to identify friction points and prioritize improvements in return and refund workflows.
템플릿 보기Patient Visit Satisfaction: Wait Times, Access & Counseling
Measures patient satisfaction with a recent healthcare visit across three domains—wait times, appointment access, and clinician counseling quality. Suitable for clinics, hospitals, and health systems seeking actionable feedback to improve the patient experience.
템플릿 보기Leisure Travel Behavior & Preferences Survey
Measures leisure travelers' trip frequency, destination decision drivers, transportation preferences, satisfaction, and likelihood to recommend — providing actionable segmentation and experience insights for travel brands, tourism boards, or academic research.
템플릿 보기Subscription Retention & Win-Back Offer Testing Survey
Identifies churn drivers among at-risk subscribers and tests the persuasiveness of specific retention offers, discounts, and feature commitments to inform cancel-flow strategy.
템플릿 보기Micro-Fulfillment Center Neighborhood Impact Survey
Measures resident perceptions of how micro-fulfillment delivery hubs affect traffic, noise, safety, jobs, and quality of life, providing evidence for local planning, zoning, and permitting decisions.
템플릿 보기