Two related decisions: which embedding model converts text to vectors, and which vector database stores and queries them. They interact.
Embedding models in 2026
Closed API models
- text-embedding-3-large (OpenAI) — 3072 dimensions, very strong general-purpose.
- text-embedding-3-small (OpenAI) — 1536 dimensions, cheaper, almost as good for most use cases.
- voyage-3 / voyage-3-large (Voyage AI) — specialized for RAG, often outperforms OpenAI on retrieval benchmarks.
- embed-english-v3 / embed-multilingual-v3 (Cohere) — strong multilingual.
Open-weight models
- BGE-M3 (BAAI) — open-source, strong multilingual.
- jina-embeddings-v3 (Jina AI) — open-source, long context (up to 8K).
- nomic-embed-text-v1.5 — open, supports variable dimensions.
- mxbai-embed-large-v1 (Mixedbread) — strong open option.
Selection criteria
1. Performance on your data
The most important factor. Benchmarks like MTEB give general signals, but YOUR data is different. Test on a small set.
2. Dimensions
- Smaller (384-768): faster, less storage, slightly less accurate.
- Larger (1024-3072): slower, more storage, slightly more accurate.
For most apps under 10M vectors, dimension impact is marginal. Don't over-optimize.
3. Multilingual support
If your data spans multiple languages, use a multilingual model (Cohere multilingual, BGE-M3, jina). Single-language models lose much of their performance on other languages.
4. Long context support
Most embedding models work on chunks up to 512 tokens. Some (jina-embeddings-v3) handle 8K. If your chunks are large, pick a long-context model.
5. Cost
API models: $0.00002-$0.0001 per 1K tokens. For 10M tokens of documents, that's $0.20 to $1. Self-hosted: just compute cost.
For one-time corpus embedding, even premium models are cheap. For real-time queries at high volume, costs add up.
Vector database options
Hosted services
Pinecone (most popular)
- Pros: zero ops, very mature, fast.
- Cons: closed-source, lock-in, can get expensive.
Qdrant Cloud
- Pros: open-source core (can self-host later), fast, mature.
- Cons: smaller ecosystem than Pinecone.
Weaviate Cloud
- Pros: open-source, supports hybrid search natively.
- Cons: more complex API.
MongoDB Atlas Vector Search / Postgres Aurora pgvector / Snowflake Cortex
- Pros: you already use these databases.
- Cons: vector capabilities can lag specialized engines.
Self-hosted
pgvector (PostgreSQL extension)
- Pros: you probably already run Postgres; no new infra.
- Cons: not as fast as specialized engines past ~5M vectors.
Qdrant
- Pros: open-source, fast, easy to self-host.
- Cons: another service to operate.
ChromaDB
- Pros: local-first, dead simple to start with.
- Cons: limited scale; mostly for prototyping.
Milvus
- Pros: handles billions of vectors.
- Cons: operational complexity.
Selection grid
| Scale | Recommendation |
|---|---|
| Prototype / < 100K vectors | ChromaDB or pgvector |
| Production / 100K-10M vectors | pgvector, Qdrant, or Pinecone |
| Large production / > 10M vectors | Pinecone, Qdrant, Weaviate, or Milvus |
| Already on Postgres | Try pgvector first |
| Zero-ops budget | Pinecone or Voyage AI |
For most teams in 2026: start with pgvector if you have Postgres, otherwise Pinecone. Migrate when scale forces it.
Hybrid search support
Production RAG should use hybrid (dense + sparse) search. Native support varies:
| DB | Hybrid native? |
|---|---|
| Weaviate | Yes (BM25 + vector + RRF) |
| Qdrant | Yes (since v1.10) |
| Pinecone | Yes (since 2024) |
| pgvector | Manual (combine pgvector with tsvector + RRF in SQL) |
If you want hybrid out of the box, Weaviate / Qdrant / Pinecone make it easier. With pgvector, you can do it but you're writing the RRF combination yourself.
Metadata filtering
You often want to filter by metadata: only retrieve chunks from documents owned by this user, in this date range, of this type.
results = vector_db.search(
query_embedding,
filter={
"user_id": current_user.id,
"doc_type": "support_article",
"created_at": {"$gte": "2025-01-01"}
},
top_k=5
)
All modern vector DBs support metadata filtering. Use it — RAG without filtering returns irrelevant cross-tenant data.
Performance benchmarks
For a 1M-vector index, typical query latency:
| DB | p50 latency |
|---|---|
| Pinecone | 30-80ms |
| Qdrant (self-hosted) | 20-50ms |
| pgvector (HNSW index) | 50-200ms |
| Weaviate | 30-80ms |
All good enough for interactive RAG. Network round-trip often dominates.
Cost benchmarks (rough)
For 5M vectors at 1536 dimensions:
| Solution | Approx monthly cost |
|---|---|
| Pinecone (Serverless) | $50-150 |
| Pinecone (Pod-based) | $200-1000 |
| Qdrant Cloud | $50-200 |
| pgvector (own Postgres) | $0 marginal (already running PG) |
| Self-hosted Qdrant | $30-100 (cloud VM) |
For most teams, the cost difference is modest. Optimize for fit-with-stack first; cost second.
When to migrate vector DBs
You start with one; you might need to move:
- Outgrowing pgvector at 5-10M vectors → Qdrant or Pinecone.
- Outgrowing managed budget → self-hosted Qdrant.
- Need better hybrid support → Weaviate or Qdrant.
Plan migration as a re-index operation. Treat the embedding model AND the vector DB as separable choices — you might keep the same embeddings while changing DBs.
Common mistakes
- Defaulting to OpenAI embeddings without testing. Voyage, Cohere, BGE often perform better for specific tasks.
- Using different embedding models for chunks vs queries. Must be the same.
- No metadata filtering. Cross-tenant data leakage at best, irrelevant retrieval at worst.
- Storing too much in vectors. Embeddings + metadata; keep raw text in object storage.
- Picking pgvector at 50M+ scale. It can do it, but specialized engines are smoother.
- Re-embedding everything to try a new model. Cost / time real; test on a sample first.
Takeaway
Embedding model: start with text-embedding-3-large or voyage-3; test against open alternatives. Vector DB: start with pgvector or Pinecone depending on your stack; migrate when scale demands. Always use metadata filtering. Default to hybrid search where supported.