모든 템플릿

RAG System Quality & Grounding Assessment (Developer Survey)

Evaluates developer experiences with Retrieval-Augmented Generation systems across retrieval quality, grounding accuracy, evaluation practices, and infrastructure. Designed for ML engineers, backend engineers, and data scientists actively building or maintaining RAG pipelines.

샘플 질문

템플릿에 포함된 내용을 미리 확인해 보세요. 모든 질문은 설문 공개 전에 자유롭게 수정할 수 있습니다.

질문 36개 · 약 15분
Q01
메시지

Welcome! Thank you for participating in this survey about your experience with Retrieval-Augmented Generation (RAG) systems. This survey takes approximately 15 minutes. Your responses are completely anonymous, reported only in aggregate, and used for internal research purposes. There are no right or wrong answers — we are interested in your honest opinions and experiences. Participation is voluntary, and you may stop at any time.

Q02
객관식

Have you built or maintained a RAG system in the last 6 months?

  • Yes
  • No
  • Not sure
Q03
객관식

Which role best describes your day-to-day work?

  • ML engineer
  • Backend engineer
  • Data scientist
  • MLOps/Platform
  • Product engineer
  • Researcher
  • Architect/Tech lead
  • Other
Q04
객관식

Which content sources feed your retriever today? Select all that apply.

  • Proprietary documents
  • Code repositories
  • Product knowledge base
  • Web crawl
  • Vendor API docs
  • Slack/Chat logs
  • Support tickets
  • Wiki/Confluence
  • Database/warehouse
  • Not applicable
  • Other
Q05
객관식

How are model answers grounded or cited in your RAG system? Select all that apply.

  • Inline citations with URLs
  • Inline citations with document IDs
  • Evidence block after the answer
  • Tool outputs included verbatim
  • Structured JSON evidence list
  • No grounding/citations
  • Other
Q06
객관식

Which evaluation tools or libraries do you use for RAG? Select all that apply.

  • Ragas
  • TruLens
  • DeepEval
  • Promptfoo
  • Custom harness
  • LlamaIndex evals
  • None
  • Other
Q07
객관식

What is the primary programming language you use for RAG development?

  • Python
  • JavaScript/TypeScript
  • Java
  • Go
  • C
  • Rust
  • Other
Q08
의견 척도

How critical is retrieval quality to the overall success of your RAG system?

척도: 17
최소:Not at all critical최대:Extremely critical
Q09
객관식

How many years of professional experience do you have in software, data, or ML?

  • 0–1
  • 2–4
  • 5–9
  • 10–14
  • 15+
  • Prefer not to say
Q10
메시지

Thank you for completing this survey! Your input is valuable and will help improve RAG systems and developer tooling. All results will be reported in aggregate only.

Q11
의견 척도

In the last 30 days, how well did retrieved context meet your task requirements?

척도: 17
최소:Far below needs최대:Far above needs
Q12
의견 척도

Over the last 30 days, how much do you trust the correctness of cited evidence in your RAG system's outputs?

척도: 17
최소:No trust at all최대:Complete trust
Q13
객관식

Which metrics best reflect your RAG quality today? Select all that apply.

  • Precision@k
  • Recall@k
  • MRR
  • nDCG
  • Answer faithfulness
  • Context precision/recall
  • Groundedness score
  • Human ratings
  • Production usage signals
  • Custom internal metrics
Q14
객관식

What is your primary vector store or retriever backend?

  • Pinecone
  • Weaviate
  • Milvus
  • FAISS
  • Elasticsearch/OpenSearch
  • pgvector
  • Chroma
  • Vespa
  • Not applicable
  • Other
Q15
의견 척도

How satisfied are you with your RAG system overall today?

척도: 17
최소:Not at all satisfied최대:Extremely satisfied
Q16
객관식

In which region do you primarily work?

  • Africa
  • Asia
  • Europe
  • North America
  • Oceania
  • South America
  • Prefer not to say
Q17
객관식

How do you set or tune top-k and related retrieval parameters?

  • Manual experimentation
  • Grid/Random search
  • Bayesian optimization
  • Vendor auto-tuning
  • Learned retrieval policy
  • Not tuned
  • Other
Q18
순위 매기기

Rank your top 3 preferred citation/grounding display styles.

  1. Inline per sentence
  2. Numbered endnotes
  3. Collapsible evidence panel
  4. Top-k sources with scores
  5. Link to full passages
  6. Show only on demand
드래그하여 순위 지정
Q19
장문형

If you use a custom metric, please briefly describe it and how you compute it.

Q20
객관식

Which embedding model is your primary choice?

  • OpenAI text-embedding-3
  • OpenAI small embedding
  • Cohere Embed
  • VoyageAI
  • Jina embeddings
  • E5 family
  • Instructor
  • BGE family
  • Local model
  • Other
Q21
순위 매기기

Rank your top 3 priorities for improving your RAG system in the next 3 months.

  1. Improve retrieval precision/recall
  2. Better grounding/citations
  3. Reduce latency
  4. Lower cost per query
  5. Scale to more data sources
  6. Harden evaluation pipeline
  7. Security/compliance
  8. Developer ergonomics
드래그하여 순위 지정
Q22
객관식

What is the approximate size of your organization (number of employees)?

  • 1–10
  • 11–50
  • 51–200
  • 201–1,000
  • 1,001–5,000
  • 5,001+
  • Prefer not to say
Q23
의견 척도

In the last 30 days, how often did you encounter irrelevant or off-topic passages in retrieval results?

척도: 15
최소:Never최대:Very often
Q24
의견 척도

In the last 30 days, how frequently did you observe hallucinations in your RAG outputs despite grounding?

척도: 15
최소:Never최대:Very frequently
Q25
드롭다운

How automated is your evaluation workflow?

  • None (manual only)
  • Some scripts
  • CI-integrated checks
  • Continuous eval in production
Q26
객관식

Do you use a reranker after initial retrieval?

  • Yes
  • No
  • Experimenting
Q27
장문형

Based on your responses in this survey, please share any additional thoughts or experiences about your RAG retrieval or grounding challenges.

Q28
객관식

What is the primary industry or domain for your RAG work?

  • Technology
  • Finance
  • Healthcare/Life sciences
  • Retail/CPG
  • Education
  • Government/Public sector
  • Manufacturing
  • Media/Entertainment
  • Other
  • Prefer not to say
Q29
의견 척도

In the last 30 days, how often did you encounter missing context (key information not retrieved)?

척도: 15
최소:Never최대:Very often
Q30
장문형

Please describe a recent grounding failure you encountered and its impact on your work.

Q31
드롭다운

How often do you run RAG benchmarks?

  • Before each release
  • Weekly
  • Biweekly
  • Monthly
  • Quarterly
  • Ad hoc only
Q32
객관식

Which reranker do you use most often?

  • Cohere Rerank
  • Voyage Rerank
  • Jina Reranker
  • Cross-encoder (e.g., MS MARCO)
  • Self-hosted reranker
  • Other
Q33
의견 척도

In the last 30 days, how often did you encounter stale or outdated content in retrieval results?

척도: 15
최소:Never최대:Very often
Q34
AI 인터뷰

We'd like to understand more about your experience with grounding and citation quality. An AI moderator will ask you a couple of follow-up questions.

Q35
드롭다운

What is your end-to-end RAG latency target per query?

  • < 200 ms
  • 200–500 ms
  • 500 ms–1 s
  • 1–2 s
  • 2–5 s
  • > 5 s
  • No specific target
Q36
의견 척도

In the last 30 days, how often did you encounter duplicate or near-duplicate chunks in retrieval results?

척도: 15
최소:Never최대:Very often

포함된 기능

  • AI 후속 질문

    정형화된 설문이 놓치는 세부 내용을, 주관식 답변에 맞춰 AI가 심층 질문으로 끌어냅니다.

  • 주의력 확인 장치

    성의 없는 답변과 저품질 응답자를 걸러내는 내장 안전장치입니다.

  • AI가 작성한 문안

    문구, 질문 순서, 분기 로직까지 AI가 연구 목표에 맞춰 작성합니다.

  • 자동 리포트

    응답이 모이면 주요 주제, 인용문, 이해하기 쉬운 요약이 자동으로 작성됩니다.

설문을 공개할 준비가 되셨나요?

이 템플릿을 편집기에서 열어 보세요. 첫 응답자가 보기 전에 모든 부분을 원하는 대로 바꿀 수 있습니다.