Across 270 responses, Pinecone appears in 75.2% of category-named queries. When a buyer describes a symptom instead of naming the category, the shape changes.
I can help you debug this! RAG retrieval issues are super common and usually come down to a handful of root causes. Let me walk through the main things to check, but I'll need some info from you to give more targeted advice.
## Common causes of bad RAG retrieval
**1. Chunking problems** (probably the #1 culprit)
- Chunks too large → embeddings become "averaged" and lose semantic specificity
- Chunks too small → lose context needed to answer
- Bad split points → sentences/ideas cut in half
- No overlap → information at boundaries gets lost
**2. Embedding model mismatch**
- Using a weak/old embedding model (e.g., older `ada-002` vs newer models)
- Model not suited to your domain (e.g., general model on legal/medical/code)
- Query and document embedded differently (asymmetric vs symmetric models)
**3. Query/document semantic gap**
- User queries are short and vague ("how do I fix this?"), but documents are long and detailed
- Documents use jargon; queries use natural language (or vice versa)
- Fix: **HyDE** (generate hypothetical answer, embed that) or **query rewriting**
**4. Pure vector search limitations**
- Vector search misses exact keyword matches (names, IDs, error codes, acronyms)
- Fix: **hybrid search** (BM25 + vector) with reciprocal rank fusion
**5. No reranking**
- Top-k from vector search is noisy
- A cross-encoder reranker (e.g., Cohere Rerank, BGE reranker)...
Every brand ranked two ways. Appearance = how often the brand is named at all. Pick share = how often the brand is the SOLE or PRIMARY pick, as a percentage of all 270 responses. The gap between the two is the recommendation gap.
| # | Brand | Appear | App % | Pick % | |
|---|---|---|---|---|---|
| 01 | Pinecone | 203 | 75.2% | 18.9% | |
| 02 | Qdrant | 155 | 57.4% | 13.0% | |
| 03 | Weaviate | 150 | 55.6% | 3.3% | |
| 04 | pgvector | 148 | 54.8% | 14.4% | |
| 05 | Postgres | 132 | 48.9% | 13.7% | |
| 06 | Chroma | 114 | 42.2% | 4.4% | |
| 07 | Milvus | 110 | 40.7% | 0.7% | |
| 08 | Supabase | 32 | 11.9% | 1.5% | |
| 09 | Zilliz | 31 | 11.5% | 0.0% | |
| 10 | LanceDB | 28 | 10.4% | 0.0% | |
| 11 | Redis | 20 | 7.4% | 0.0% | |
| 12 | OpenSearch | 20 | 7.4% | 0.0% |
| Bucket | Top | n | Runner-up |
|---|---|---|---|
| Category discovery | Pinecone | 36 | Postgres 36 |
| Comparison | Pinecone | 36 | Qdrant 21 |
| Segment fit | Pinecone | 34 | Weaviate 34 |
| Evaluation | Pinecone | 30 | Weaviate 27 |
| Pricing | Pinecone | 25 | Qdrant 17 |
| Trust | Pinecone | 14 | Weaviate 10 |
| Switching | Pinecone | 18 | Qdrant 9 |
| Problem-first | Milvus | 12 | Pinecone 10 |
The highest-leverage queries: the buyer describes a symptom rather than naming the category.
| # | Brand | Count | % |
|---|---|---|---|
| 01 | Milvus | 12 | 19.0% |
| 02 | Pinecone | 10 | 15.9% |
| 03 | Qdrant | 10 | 15.9% |
| 04 | pgvector | 8 | 12.7% |
| 05 | Weaviate | 7 | 11.1% |
| n/a | Zero-vendor responses | 48 | 76.2% |
Two lists. Vendor-owned and editorial sources on the left. SEO listing sites on the right. The split reveals which channel is doing the citation work.
Every 30-day pilot starts with one. Fixed fee. Ad spend billed to your account. Written verdict on day 30.