すべてのテンプレート
Sample Templates

RAG System Quality & Grounding Assessment (Developer Survey)

Evaluates developer experiences with Retrieval-Augmented Generation systems across retrieval quality, grounding accuracy, evaluation practices, and infrastructure. Designed for ML engineers, backend engineers, and data scientists actively building or maintaining RAG pipelines.

設問の例

テンプレートの内容をプレビューできます。すべての設問は公開前に自由に編集できます。

全36問・約15分
Q01
メッセージ

Welcome! Thank you for participating in this survey about your experience with Retrieval-Augmented Generation (RAG) systems. This survey takes approximately 15 minutes. Your responses are completely anonymous, reported only in aggregate, and used for internal research purposes. There are no right or wrong answers — we are interested in your honest opinions and experiences. Participation is voluntary, and you may stop at any time.

Q02
選択式

Have you built or maintained a RAG system in the last 6 months?

  • Yes
  • No
  • Not sure
Q03
選択式

Which role best describes your day-to-day work?

  • ML engineer
  • Backend engineer
  • Data scientist
  • MLOps/Platform
  • Product engineer
  • Researcher
  • Architect/Tech lead
  • Other
Q04
選択式

Which content sources feed your retriever today? Select all that apply.

  • Proprietary documents
  • Code repositories
  • Product knowledge base
  • Web crawl
  • Vendor API docs
  • Slack/Chat logs
  • Support tickets
  • Wiki/Confluence
  • Database/warehouse
  • Not applicable
  • Other
Q05
選択式

How are model answers grounded or cited in your RAG system? Select all that apply.

  • Inline citations with URLs
  • Inline citations with document IDs
  • Evidence block after the answer
  • Tool outputs included verbatim
  • Structured JSON evidence list
  • No grounding/citations
  • Other
Q06
選択式

Which evaluation tools or libraries do you use for RAG? Select all that apply.

  • Ragas
  • TruLens
  • DeepEval
  • Promptfoo
  • Custom harness
  • LlamaIndex evals
  • None
  • Other
Q07
選択式

What is the primary programming language you use for RAG development?

  • Python
  • JavaScript/TypeScript
  • Java
  • Go
  • C
  • Rust
  • Other
Q08
オピニオンスケール

How critical is retrieval quality to the overall success of your RAG system?

スケール: 1 – 7
最小:Not at all critical最大:Extremely critical
Q09
選択式

How many years of professional experience do you have in software, data, or ML?

  • 0–1
  • 2–4
  • 5–9
  • 10–14
  • 15+
  • Prefer not to say
Q10
メッセージ

Thank you for completing this survey! Your input is valuable and will help improve RAG systems and developer tooling. All results will be reported in aggregate only.

Q11
オピニオンスケール

In the last 30 days, how well did retrieved context meet your task requirements?

スケール: 1 – 7
最小:Far below needs最大:Far above needs
Q12
オピニオンスケール

Over the last 30 days, how much do you trust the correctness of cited evidence in your RAG system's outputs?

スケール: 1 – 7
最小:No trust at all最大:Complete trust
Q13
選択式

Which metrics best reflect your RAG quality today? Select all that apply.

  • Precision@k
  • Recall@k
  • MRR
  • nDCG
  • Answer faithfulness
  • Context precision/recall
  • Groundedness score
  • Human ratings
  • Production usage signals
  • Custom internal metrics
Q14
選択式

What is your primary vector store or retriever backend?

  • Pinecone
  • Weaviate
  • Milvus
  • FAISS
  • Elasticsearch/OpenSearch
  • pgvector
  • Chroma
  • Vespa
  • Not applicable
  • Other
Q15
オピニオンスケール

How satisfied are you with your RAG system overall today?

スケール: 1 – 7
最小:Not at all satisfied最大:Extremely satisfied
Q16
選択式

In which region do you primarily work?

  • Africa
  • Asia
  • Europe
  • North America
  • Oceania
  • South America
  • Prefer not to say
Q17
選択式

How do you set or tune top-k and related retrieval parameters?

  • Manual experimentation
  • Grid/Random search
  • Bayesian optimization
  • Vendor auto-tuning
  • Learned retrieval policy
  • Not tuned
  • Other
Q18
ランク付け

Rank your top 3 preferred citation/grounding display styles.

  1. Inline per sentence
  2. Numbered endnotes
  3. Collapsible evidence panel
  4. Top-k sources with scores
  5. Link to full passages
  6. Show only on demand
ドラッグして順位を付ける
Q19
自由回答(長文)

If you use a custom metric, please briefly describe it and how you compute it.

Q20
選択式

Which embedding model is your primary choice?

  • OpenAI text-embedding-3
  • OpenAI small embedding
  • Cohere Embed
  • VoyageAI
  • Jina embeddings
  • E5 family
  • Instructor
  • BGE family
  • Local model
  • Other
Q21
ランク付け

Rank your top 3 priorities for improving your RAG system in the next 3 months.

  1. Improve retrieval precision/recall
  2. Better grounding/citations
  3. Reduce latency
  4. Lower cost per query
  5. Scale to more data sources
  6. Harden evaluation pipeline
  7. Security/compliance
  8. Developer ergonomics
ドラッグして順位を付ける
Q22
選択式

What is the approximate size of your organization (number of employees)?

  • 1–10
  • 11–50
  • 51–200
  • 201–1,000
  • 1,001–5,000
  • 5,001+
  • Prefer not to say
Q23
オピニオンスケール

In the last 30 days, how often did you encounter irrelevant or off-topic passages in retrieval results?

スケール: 1 – 5
最小:Never最大:Very often
Q24
オピニオンスケール

In the last 30 days, how frequently did you observe hallucinations in your RAG outputs despite grounding?

スケール: 1 – 5
最小:Never最大:Very frequently
Q25
プルダウン

How automated is your evaluation workflow?

  • None (manual only)
  • Some scripts
  • CI-integrated checks
  • Continuous eval in production
Q26
選択式

Do you use a reranker after initial retrieval?

  • Yes
  • No
  • Experimenting
Q27
自由回答(長文)

Based on your responses in this survey, please share any additional thoughts or experiences about your RAG retrieval or grounding challenges.

Q28
選択式

What is the primary industry or domain for your RAG work?

  • Technology
  • Finance
  • Healthcare/Life sciences
  • Retail/CPG
  • Education
  • Government/Public sector
  • Manufacturing
  • Media/Entertainment
  • Other
  • Prefer not to say
Q29
オピニオンスケール

In the last 30 days, how often did you encounter missing context (key information not retrieved)?

スケール: 1 – 5
最小:Never最大:Very often
Q30
自由回答(長文)

Please describe a recent grounding failure you encountered and its impact on your work.

Q31
プルダウン

How often do you run RAG benchmarks?

  • Before each release
  • Weekly
  • Biweekly
  • Monthly
  • Quarterly
  • Ad hoc only
Q32
選択式

Which reranker do you use most often?

  • Cohere Rerank
  • Voyage Rerank
  • Jina Reranker
  • Cross-encoder (e.g., MS MARCO)
  • Self-hosted reranker
  • Other
Q33
オピニオンスケール

In the last 30 days, how often did you encounter stale or outdated content in retrieval results?

スケール: 1 – 5
最小:Never最大:Very often
Q34
AIインタビュー

We'd like to understand more about your experience with grounding and citation quality. An AI moderator will ask you a couple of follow-up questions.

Q35
プルダウン

What is your end-to-end RAG latency target per query?

  • < 200 ms
  • 200–500 ms
  • 500 ms–1 s
  • 1–2 s
  • 2–5 s
  • > 5 s
  • No specific target
Q36
オピニオンスケール

In the last 30 days, how often did you encounter duplicate or near-duplicate chunks in retrieval results?

スケール: 1 – 5
最小:Never最大:Very often

含まれる機能

  • AIによる深掘り

    自由回答に合わせてAIが追加で質問し、固定のフォームでは拾えない具体的な内容を引き出します。

  • 注意確認設問

    急いだ回答や質の低い回答者を除外する仕組みを標準で備えています。

  • AIが作成する設問文

    文言、設問の順序、条件分岐をAIが調査の目的に合わせて作成します。

  • 自動レポート

    回答が集まると、テーマ、引用、わかりやすい要約が自動で作成されます。

よくあるご質問

「RAG System Quality & Grounding Assessment (Developer Survey)」テンプレートにはどのような設問が含まれていますか?

すぐに使える設問が36問含まれており、最初の設問は次のとおりです:「Welcome! Thank you for participating in this survey about your experience with Retrieval-Augmented Generation (RAG) syst…」・「Have you built or maintained a RAG system in the last 6 months?」・「Which role best describes your day-to-day work?」。すべての設問は上でプレビューでき、自由に編集できます。

このアンケートの回答にはどのくらい時間がかかりますか?

回答者は通常、36問を約15分で回答し終えます。

テンプレートは編集できますか?

はい。公開前であれば、すべての設問、選択肢、順序を編集できます。設問の追加や削除のほか、調査の目的に合わせた作り直しをAIエディターに依頼することもできます。

このテンプレートは無料で使えますか?

はい。エディターで開けば、すぐに編集を始められます。お試しにアカウントは不要で、無料プランでアンケートを公開できます。

公開の準備はできましたか?

このテンプレートをエディターで開いてみてください。最初の回答者が目にする前に、すべてを自由に変更できます。

関連テンプレート

似たテーマのほかの調査もご覧ください。

すべて見る