RAG System Quality & Grounding Assessment (Developer Survey)
Evaluates developer experiences with Retrieval-Augmented Generation systems across retrieval quality, grounding accuracy, evaluation practices, and infrastructure. Designed for ML engineers, backend engineers, and data scientists actively building or maintaining RAG pipelines.
設問の例
テンプレートの内容をプレビューできます。すべての設問は公開前に自由に編集できます。
Have you built or maintained a RAG system in the last 6 months?
- Yes
- No
- Not sure
Which role best describes your day-to-day work?
- ML engineer
- Backend engineer
- Data scientist
- MLOps/Platform
- Product engineer
- Researcher
- Architect/Tech lead
- Other
Which content sources feed your retriever today? Select all that apply.
- Proprietary documents
- Code repositories
- Product knowledge base
- Web crawl
- Vendor API docs
- Slack/Chat logs
- Support tickets
- Wiki/Confluence
- Database/warehouse
- Not applicable
- Other
How are model answers grounded or cited in your RAG system? Select all that apply.
- Inline citations with URLs
- Inline citations with document IDs
- Evidence block after the answer
- Tool outputs included verbatim
- Structured JSON evidence list
- No grounding/citations
- Other
Which evaluation tools or libraries do you use for RAG? Select all that apply.
- Ragas
- TruLens
- DeepEval
- Promptfoo
- Custom harness
- LlamaIndex evals
- None
- Other
What is the primary programming language you use for RAG development?
- Python
- JavaScript/TypeScript
- Java
- Go
- C
- Rust
- Other
How critical is retrieval quality to the overall success of your RAG system?
How many years of professional experience do you have in software, data, or ML?
- 0–1
- 2–4
- 5–9
- 10–14
- 15+
- Prefer not to say
Thank you for completing this survey! Your input is valuable and will help improve RAG systems and developer tooling. All results will be reported in aggregate only.
In the last 30 days, how well did retrieved context meet your task requirements?
Over the last 30 days, how much do you trust the correctness of cited evidence in your RAG system's outputs?
Which metrics best reflect your RAG quality today? Select all that apply.
- Precision@k
- Recall@k
- MRR
- nDCG
- Answer faithfulness
- Context precision/recall
- Groundedness score
- Human ratings
- Production usage signals
- Custom internal metrics
What is your primary vector store or retriever backend?
- Pinecone
- Weaviate
- Milvus
- FAISS
- Elasticsearch/OpenSearch
- pgvector
- Chroma
- Vespa
- Not applicable
- Other
How satisfied are you with your RAG system overall today?
In which region do you primarily work?
- Africa
- Asia
- Europe
- North America
- Oceania
- South America
- Prefer not to say
How do you set or tune top-k and related retrieval parameters?
- Manual experimentation
- Grid/Random search
- Bayesian optimization
- Vendor auto-tuning
- Learned retrieval policy
- Not tuned
- Other
Rank your top 3 preferred citation/grounding display styles.
- Inline per sentence
- Numbered endnotes
- Collapsible evidence panel
- Top-k sources with scores
- Link to full passages
- Show only on demand
If you use a custom metric, please briefly describe it and how you compute it.
Which embedding model is your primary choice?
- OpenAI text-embedding-3
- OpenAI small embedding
- Cohere Embed
- VoyageAI
- Jina embeddings
- E5 family
- Instructor
- BGE family
- Local model
- Other
Rank your top 3 priorities for improving your RAG system in the next 3 months.
- Improve retrieval precision/recall
- Better grounding/citations
- Reduce latency
- Lower cost per query
- Scale to more data sources
- Harden evaluation pipeline
- Security/compliance
- Developer ergonomics
What is the approximate size of your organization (number of employees)?
- 1–10
- 11–50
- 51–200
- 201–1,000
- 1,001–5,000
- 5,001+
- Prefer not to say
In the last 30 days, how often did you encounter irrelevant or off-topic passages in retrieval results?
In the last 30 days, how frequently did you observe hallucinations in your RAG outputs despite grounding?
How automated is your evaluation workflow?
- None (manual only)
- Some scripts
- CI-integrated checks
- Continuous eval in production
Do you use a reranker after initial retrieval?
- Yes
- No
- Experimenting
Based on your responses in this survey, please share any additional thoughts or experiences about your RAG retrieval or grounding challenges.
What is the primary industry or domain for your RAG work?
- Technology
- Finance
- Healthcare/Life sciences
- Retail/CPG
- Education
- Government/Public sector
- Manufacturing
- Media/Entertainment
- Other
- Prefer not to say
In the last 30 days, how often did you encounter missing context (key information not retrieved)?
Please describe a recent grounding failure you encountered and its impact on your work.
How often do you run RAG benchmarks?
- Before each release
- Weekly
- Biweekly
- Monthly
- Quarterly
- Ad hoc only
Which reranker do you use most often?
- Cohere Rerank
- Voyage Rerank
- Jina Reranker
- Cross-encoder (e.g., MS MARCO)
- Self-hosted reranker
- Other
In the last 30 days, how often did you encounter stale or outdated content in retrieval results?
We'd like to understand more about your experience with grounding and citation quality. An AI moderator will ask you a couple of follow-up questions.
What is your end-to-end RAG latency target per query?
- < 200 ms
- 200–500 ms
- 500 ms–1 s
- 1–2 s
- 2–5 s
- > 5 s
- No specific target
In the last 30 days, how often did you encounter duplicate or near-duplicate chunks in retrieval results?
含まれる機能
AIによる深掘り
自由回答に合わせてAIが追加で質問し、固定のフォームでは拾えない具体的な内容を引き出します。
注意確認設問
急いだ回答や質の低い回答者を除外する仕組みを標準で備えています。
AIが作成する設問文
文言、設問の順序、条件分岐をAIが調査の目的に合わせて作成します。
自動レポート
回答が集まると、テーマ、引用、わかりやすい要約が自動で作成されます。
よくあるご質問
「RAG System Quality & Grounding Assessment (Developer Survey)」テンプレートにはどのような設問が含まれていますか?
すぐに使える設問が36問含まれており、最初の設問は次のとおりです:「Welcome! Thank you for participating in this survey about your experience with Retrieval-Augmented Generation (RAG) syst…」・「Have you built or maintained a RAG system in the last 6 months?」・「Which role best describes your day-to-day work?」。すべての設問は上でプレビューでき、自由に編集できます。
このアンケートの回答にはどのくらい時間がかかりますか?
回答者は通常、36問を約15分で回答し終えます。
テンプレートは編集できますか?
はい。公開前であれば、すべての設問、選択肢、順序を編集できます。設問の追加や削除のほか、調査の目的に合わせた作り直しをAIエディターに依頼することもできます。
このテンプレートは無料で使えますか?
はい。エディターで開けば、すぐに編集を始められます。お試しにアカウントは不要で、無料プランでアンケートを公開できます。
公開の準備はできましたか?
このテンプレートをエディターで開いてみてください。最初の回答者が目にする前に、すべてを自由に変更できます。
関連テンプレート
似たテーマのほかの調査もご覧ください。