Best Embedding Models in 2026: Pricing, Dimensions, and Real Workload Costs Compared
Turn this article into takeaways for your work.
Each assistant summarizes the article only for you and suggests best practices for your work.
There's no single best embedding model in 2026, only a best one for your corpus size, context length needs, and the cloud bill you're already paying. OpenAI's text-embedding-3-small is still the default nobody gets questioned for picking. Voyage AI's voyage-4 family usually wins on quality per dollar once you price a real workload, partly because its free tier is large enough to cover most pilots outright. Qwen3-Embedding and BGE-M3 are the open-weight options worth running yourself when a third-party API call on your data is a non-starter.
This guide prices eleven commercial and open-weight embedding families at a concrete workload (10 million tokens indexed once, plus 2 million query tokens a month) instead of the advertised headline rate, because that's where the real ranking lives. Figures marked (reported) are approximate. This is a technical buying decision, not just a benchmark one: who runs it, what switching costs later, and whether the vendor still exists in its current form a year from now all matter more than a leaderboard rank.
Key Facts
- Google's own docs state that
gemini-embedding-001and the newergemini-embedding-2produce incompatible vector spaces, so upgrading forces a full re-embed of existing data (source: Google AI for Developers). - Cohere's Embed 5 Pro and Fast models share one embedding space, so you can index with Pro and query with the cheaper Fast model without re-indexing (source: Cohere, "Embed 5", 30 September 2026).
- MongoDB acquired Voyage AI for roughly $220 million in February 2025 and has folded its models into Atlas Vector Search, while Voyage's standalone API stays live (source: SiliconANGLE).
- BAAI's BGE-M3 is MIT-licensed and the only model here that natively produces dense, sparse, and ColBERT-style multi-vector output from a single pass (source: BAAI/bge-m3).
- Qwen's own model card reports Qwen3-Embedding-8B scoring 70.58 on MTEB multilingual as of its June 2025 release, a vendor claim rather than an independent ranking (source: Qwen/Qwen3-Embedding-8B).
Quick Comparison Table
| Model | Best For | Price (per 1M tokens) | Max Dimensions | Max Context | Key Limitation |
|---|---|---|---|---|---|
| OpenAI text-embedding-3-small/large | One vendor bill alongside GPT API usage | $0.02 / $0.13 | 1,536 / 3,072 | 8,191 tokens | No native multimodal input |
| Cohere Embed 5 (Pro/Fast) | Splitting cheap query traffic from quality indexing | $0.12 / $0.08 | 2,048 | 128,000 tokens | No published always-free tier |
| Voyage AI voyage-4 family | Best retrieval quality per dollar, large free tier | $0.02 to $0.12 (200M free) | 2,048 | 32,000 tokens | Now MongoDB-owned, roadmap tied to Atlas |
| Jina Embeddings v5 | Small multilingual footprint, low entry cost | ~$0.05 (reported) | 1,024 | 8,192 to 32,768 tokens | Exact per-SKU rate needs dashboard login |
| Google Gemini Embedding 2 | One call embeds text, image, video, audio, PDF | $0.20 standard / $0.10 batch | 3,072 (MRL) | 8,192 tokens | Incompatible with prior gemini-embedding-001 |
| Mistral Embed | EU-based or Mistral-stack teams | $0.10 | 1,024 (fixed) | ~8,000 tokens | No Matryoshka truncation offered |
| Amazon Titan Text Embeddings V2 | AWS-native orgs consolidating the Bedrock bill | $0.02 | 1,024 | 8,192 tokens | Fewer independent benchmark citations |
| BAAI BGE-M3 (open weight) | Self-hosted hybrid dense+sparse+ColBERT | $0 API fee | 1,024 | 8,192 tokens | You own uptime, scaling, GPU cost |
| intfloat E5 (multilingual-e5-large-instruct) | Well-documented open baseline | $0 API fee | 1,024 | 512 tokens | Short context forces aggressive chunking |
| Nomic Embed v2-MoE (open weight) | Small, efficient self-hosted multilingual model | $0 API fee | 768 (truncatable to 256) | 512 tokens | Same short context ceiling as E5 |
| Qwen3-Embedding (open weight) | Top open-weight benchmark performance | $0 API fee | 4,096 (configurable 32 to 4,096) | 32,000 tokens | 8B variant needs real GPU memory |
What Embeddings Actually Cost: Why the Headline Rate Lies
Every vendor publishes a per-million-token rate, and most comparisons stop there. That number is close to meaningless alone, because the real drivers of your bill (how much text you index once, how many queries you run monthly, and whether a free tier absorbs either) vary by an order of magnitude between a pilot and production. New to the concept? Our plain-English explainer on embeddings covers it in two minutes, and vector databases covers where those numbers end up living.
The workload below is deliberately modest: 10 million tokens embedded once (roughly 15,000 to 25,000 pages, depending on chunk size) plus 2 million query tokens a month, a realistic volume for an internal knowledge base or a mid-size SaaS product's RAG layer feeding a team's AI agents. Annualized, that's 34 million total tokens in year one.
Pricing at a Real Workload (10M Indexed + 2M/Month Queried)
| Model | Indexing (10M tokens) | Querying (24M tokens/yr) | Year-One Total |
|---|---|---|---|
| OpenAI text-embedding-3-small | $0.20 | $0.48 | $0.68 |
| OpenAI text-embedding-3-large | $1.30 | $3.12 | $4.42 |
| Cohere Embed 5 (Pro index / Fast query) | $1.20 | $1.92 | $3.12 |
| Voyage AI voyage-4 | $0.00 (within 200M free) | $0.00 (within 200M free) | $0.00 |
| Voyage AI voyage-4-large | $0.00 (within 200M free) | $0.00 (within 200M free) | $0.00 |
| Jina Embeddings v5-text-small | $0.00 (10M free tier) | $1.20 (reported) | ~$1.20 (reported) |
| Google gemini-embedding-2 (standard) | $2.00 | $4.80 | $6.80 |
| Google gemini-embedding-001 (reported) | $1.50 | $3.60 | $5.10 (reported) |
| Mistral Embed | $1.00 | $2.40 | $3.40 |
| Amazon Titan Text Embeddings V2 | $0.20 | $0.48 | $0.68 |
Open-weight models (BGE-M3, E5, Nomic, Qwen3-Embedding) carry no per-token fee here; see the self-hosting section for their real cost.
This is the flip the headline rate hides. Gemini's new multimodal flagship costs ten times what Titan V2 or OpenAI's small model charges for the identical text-only workload, and Voyage's free tier means a huge share of pilots, and even some small production deployments, pay nothing for models that independently rank near the top of the field. Past 200 million lifetime tokens on Voyage, the free allowance stops and you're back to $0.02 to $0.12 per million depending on tier.
The Re-Embedding Cost Nobody Prices (The Real Lock-In)
Picking an embedding model isn't a one-time decision. Every vector you store is tied to the exact model and version that produced it, and swapping vendors, or even upgrading within the same vendor, means regenerating every vector from scratch. Google's own documentation says it outright: gemini-embedding-001 and gemini-embedding-2 produce incompatible embedding spaces, so the newer model requires re-embedding everything indexed under the old one. That's not a hypothetical, it's the vendor's own migration note.
| Model | Cost to Re-Embed a 10M-Token Corpus |
|---|---|
| OpenAI text-embedding-3-small | $0.20 |
| OpenAI text-embedding-3-large | $1.30 |
| Cohere Embed 5 Pro | $1.20 |
| Voyage AI voyage-4 | $0.00 to $0.60 (depends on remaining free tier) |
| Jina Embeddings v5-text-small | $0.50 (reported) |
| Google gemini-embedding-2 | $2.00 |
| Mistral Embed | $1.00 |
| Amazon Titan Text Embeddings V2 | $0.20 |
| Open-weight (BGE-M3, E5, Nomic, Qwen3) | $0 API fee, full GPU re-inference pass |
The dollar figure is almost always trivial, which is exactly why it gets ignored during vendor selection. The real cost is operational: re-chunking source documents if your strategy changed too, standing up a parallel index, running a dual-write cutover so search doesn't go dark mid-migration, and re-validating retrieval quality before you delete the old vectors. That's a multi-day project triggered by an upgrade you didn't initiate, and it recurs every time a vendor ships a new generation. Favor a family with Matryoshka truncation or clear version-compatibility notes (OpenAI's v3, Voyage's v4, Cohere's Embed 5, Qwen3-Embedding) for fewer forced migrations.
1. OpenAI text-embedding-3-small and text-embedding-3-large
OpenAI's third-generation embeddings remain the default for teams who already pay OpenAI for chat completions and don't want a second vendor relationship. text-embedding-3-small produces 1,536-dimension vectors at $0.02 per million tokens; text-embedding-3-large goes up to 3,072 dimensions at $0.13 per million, both truncatable via Matryoshka learning with minimal quality loss. The Batch API halves both rates for workloads that tolerate a 24-hour turnaround, which is most bulk-indexing jobs.
| text-embedding-3-small | text-embedding-3-large | |
|---|---|---|
| Price (per 1M tokens) | $0.02 | $0.13 |
| Dimensions | 1,536 (truncatable) | 3,072 (truncatable) |
| Context length | 8,191 tokens | 8,191 tokens |
Best for: teams standardized on OpenAI for generation wanting one API key and invoice. Not ideal for: anyone needing multimodal embeddings or EU-only processing without the Enterprise tier. Source: OpenAI API pricing.
2. Cohere Embed 5
Cohere shipped Embed 5 on 30 September 2026, replacing the Embed v3/v4 line with Pro and Fast variants sharing one embedding space. That shared space is the headline feature: index once with Pro for quality, then query with the cheaper, 2.4x-faster Fast model without re-indexing. Both support a 128,000-token context, dramatically longer than any other model here, and dimensions down to 256 via Matryoshka compression.
Best for: teams with real query volume who want to decouple indexing cost from serving cost. Not ideal for: teams wanting a transparent self-serve free tier to prototype against before committing, since Cohere doesn't publish one comparable to Voyage's. Source: Cohere, "Embed 5".
3. Voyage AI (voyage-4 family)
Voyage AI, acquired by MongoDB for roughly $220 million in February 2025, still runs a standalone API and ships the voyage-4 family: voyage-4 ($0.06/M), voyage-4-large ($0.12/M), voyage-4-lite ($0.02/M), plus voyage-code-4 for code and voyage-context-4 for contextualized chunk embeddings, all with the first 200 million tokens free per account. That free tier covers the entire 10M-plus-24M-token workload modeled above at zero cost, the single biggest factor in Voyage's price-to-quality reputation. Every v4 model supports configurable dimensions (256, 512, 1,024, 2,048) and an 8x longer context window than OpenAI or Titan at 32,000 tokens.
Best for: teams prototyping or running mid-size production RAG wanting generous free usage and flexible dimension sizing. Not ideal for: teams wary of a roadmap now controlled by a database vendor rather than an independent AI lab. Source: Voyage AI pricing docs.
4. Jina Embeddings v5
Jina's v5 series ships four variants built on the Qwen3 and EuroBERT tokenizers: v5-text-small (1,024 dimensions, 32,768-token context, 119 languages) and v5-text-nano (768 dimensions, 8,192-token context, 15-plus languages), plus multimodal omni variants. New accounts get 10 million free tokens. Jina's public page doesn't post an exact per-model rate (you need a dashboard login); the text-small SKU reportedly runs about $0.05 per million tokens.
Best for: budget-conscious teams needing broad multilingual coverage in a small footprint. Not ideal for: teams needing a published, citable price before budget sign-off, since Jina's rate requires a logged-in dashboard. Source: Jina Embeddings v5-text announcement for specs.
5. Google Gemini Embedding (gemini-embedding-2)
Google's current flagship, gemini-embedding-2, is a genuine category shift: Google's first multimodal embedding model, mapping text, images, video, audio, and PDFs into one unified vector space in a single call. Standard pricing is $0.20 per million tokens for text, dropping to $0.10 on the Batch tier, with image, audio, and video priced separately and considerably higher. Dimensions run from 128 to 3,072 via Matryoshka learning, but context is a relatively tight 8,192 tokens. The older text-only gemini-embedding-001 (2,048-token context) is still documented and reportedly priced around $0.15 per million standard, but Google states plainly the two models produce incompatible embedding spaces, so there's no free upgrade path.
Best for: teams on Google Cloud or Vertex AI wanting one call to cover text, images, and PDFs together. Not ideal for: anyone who embedded a corpus under gemini-embedding-001 and assumed an in-place upgrade; budget a full re-embed. Source: Gemini API pricing and Gemini Embedding docs.
6. Mistral Embed
Mistral Embed produces fixed 1,024-dimension vectors from a roughly 8,000-token context window at $0.10 per million tokens, with a code-focused sibling, Codestral Embed, at $0.15 per million. Unlike every other commercial model here, Mistral doesn't offer a selectable or Matryoshka-truncatable dimension: 1,024 or nothing, which simplifies storage planning but removes a cost lever every competitor gives you.
Best for: EU-headquartered teams or shops already standardized on Mistral's LLMs who want one European vendor relationship. Not ideal for: teams that want to shrink vector storage later without re-embedding, since there's no truncation option. Source: Mistral AI API pricing.
7. Amazon Titan Text Embeddings V2
Titan Text Embeddings V2 is AWS's Bedrock-native option: $0.02 per million input tokens (matching OpenAI's small model), flexible output at 256, 512, or 1,024 dimensions, and an 8,192-token context window. Its appeal isn't benchmark leadership, it's that the charge lands on the same AWS bill as the rest of your infrastructure, under the IAM, VPC, and procurement controls you already have for Bedrock.
Best for: AWS-committed teams avoiding a second vendor contract and data-processing agreement. Not ideal for: teams chasing the top of independent retrieval benchmarks, where Titan is rarely cited as a leader against OpenAI, Voyage, or the open-weight frontier. Source: AWS, "Get started with Amazon Titan Text Embeddings V2".
8. BAAI BGE-M3 (open weight)
BGE-M3 is worth knowing even if you never self-host anything else: it's the only model here producing dense, sparse (lexical), and ColBERT-style multi-vector output from a single pass, so one model powers both semantic and keyword-style hybrid search. It's MIT-licensed, outputs 1,024-dimension dense vectors, handles up to 8,192 tokens, and covers more than 100 languages.
Best for: technical teams with GPU infrastructure wanting hybrid dense-plus-sparse retrieval and strict data sovereignty. Not ideal for: teams without the appetite to own uptime, autoscaling, and GPU cost themselves. Source: BAAI/bge-m3 model card.
9. intfloat E5 (multilingual-e5-large-instruct)
E5 is the open-weight baseline most RAG tooling defaults to when a tutorial needs "a free embedding model." The multilingual-e5-large-instruct variant is MIT-licensed, outputs 1,024-dimension vectors across 94 languages, and integrates cleanly with sentence-transformers and most vector database SDKs. Its real limitation is the 512-token context window, short enough to force aggressive chunking, well below the page-level chunks Voyage or Cohere allow.
Best for: teams wanting a thoroughly documented, widely supported open baseline. Not ideal for: long-document retrieval, where 512 tokens forces many small, context-poor chunks. Source: intfloat/multilingual-e5-large-instruct model card.
10. Nomic Embed v2-MoE (open weight)
Nomic's v2-MoE is a mixture-of-experts model with 475 million total parameters but only 305 million active at inference, cheaper to run than its quality class suggests. It outputs 768-dimension vectors, truncatable to 256 via Matryoshka learning for roughly a 3x storage reduction, trained across roughly 100 languages on 1.6 billion pairs. Like E5, it caps out at a 512-token context window.
Best for: teams wanting a small, GPU-efficient self-hosted model for resource-constrained deployments. Not ideal for: the same long-document use case E5 struggles with; the short context is shared across most of the open "base" generation of embedding models. Source: nomic-ai/nomic-embed-text-v2-moe model card.
11. Qwen3-Embedding (open weight)
Qwen3-Embedding is the strongest open-weight performer here on paper: it ships in 0.6B, 4B, and 8B sizes, supports a 32,000-token context matching Voyage, and lets you configure output dimensions from 32 up to 4,096 depending on the size you run. It's Apache 2.0 licensed and covers over 100 languages. Qwen's own model card reports the 8B variant scoring 70.58 on MTEB multilingual as of its June 2025 release, a vendor-published figure rather than an independent rank.
Best for: technical teams with GPU budget who want the best open-weight retrieval quality available and full control over dimension sizing. Not ideal for: teams without the infrastructure to serve an 8B-parameter model at low latency; the smaller 0.6B and 4B variants trade some of that quality back for easier hosting. Source: Qwen/Qwen3-Embedding-8B model card.
Dimensions, Context Length, and What They Cost You in Storage
Dimension count decides your vector database bill independent of what you pay the embedding vendor. A higher-dimension vector captures more nuance but costs proportionally more disk, memory, and latency to store and query, which is why Matryoshka support matters as much as the raw benchmark score. The context window a model accepts matters just as much on the input side: a short window forces smaller chunks, which means more chunks and more vectors, for the same source corpus.
| Model | Dimensions | Max Context | Multilingual | Matryoshka/Truncation |
|---|---|---|---|---|
| OpenAI text-embedding-3-large | up to 3,072 | 8,191 tokens | Yes | Yes |
| Cohere Embed 5 | up to 2,048 | 128,000 tokens | 100+ languages | Yes |
| Voyage voyage-4 family | up to 2,048 | 32,000 tokens | Yes (multilingual-2 legacy) | Yes |
| Jina v5-text-small | 1,024 | 32,768 tokens | 119 languages | Partial (model-dependent) |
| Gemini gemini-embedding-2 | up to 3,072 | 8,192 tokens | 100+ languages | Yes |
| Mistral Embed | 1,024 (fixed) | ~8,000 tokens | Not emphasized | No |
| AWS Titan Text Embeddings V2 | 256 / 512 / 1,024 | 8,192 tokens | Yes | Selectable, not Matryoshka-trained |
| BAAI BGE-M3 | 1,024 | 8,192 tokens | 100+ languages | No (fixed dense output) |
| intfloat E5 (multilingual-large) | 1,024 | 512 tokens | 94 languages | No |
| Nomic Embed v2-MoE | 768 | 512 tokens | ~100 languages | Yes (to 256) |
| Qwen3-Embedding-8B | up to 4,096 | 32,000 tokens | 100+ languages | Yes (32 to 4,096) |
To make the storage math concrete: at 32-bit float precision, a million vectors at 3,072 dimensions runs about 12.3 GB, the same million at 1,024 dimensions runs about 4.1 GB, and truncated to 256 via Matryoshka, about 1 GB. That's arithmetic (4 bytes per dimension, times dimension count, times vector count), not a vendor claim, but it's what determines whether your vector database bill scales linearly or explosively as your corpus grows. Model the storage cost at your chosen dimension before you commit, not after. The vector database roundup prices 12 engines at a 1-million-vector workload, so you can run that math against real storage rates.
Open-Weight Models Head to Head
| BGE-M3 | E5 (multilingual-large-instruct) | Nomic Embed v2-MoE | Qwen3-Embedding-8B | |
|---|---|---|---|---|
| License | MIT | MIT | Apache 2.0 | Apache 2.0 |
| Parameters | Not disclosed as MoE | ~560M (24-layer) | 475M total, 305M active (MoE) | 8B |
| Dimensions | 1,024 | 1,024 | 768 (to 256) | up to 4,096 |
| Context | 8,192 tokens | 512 tokens | 512 tokens | 32,000 tokens |
| Distinct capability | Dense + sparse + ColBERT in one pass | Broadest tooling support | Smallest active-parameter footprint | Longest context, largest dimension ceiling |
API vs Self-Hosted: Who Should Actually Run the Open-Weight Models
Self-hosting BGE-M3, E5, Nomic, or Qwen3-Embedding removes the per-token fee, but it doesn't remove cost, it moves it onto your infrastructure budget and your team's time. A GPU serving one of these models around the clock costs real money regardless of volume processed, so the trade pays off once token volume exceeds what a metered API would charge for a dedicated instance, or once data residency rules make an external API call a non-starter regardless of price.
| Factor | Hosted API (OpenAI, Voyage, Cohere, etc.) | Self-hosted open weight |
|---|---|---|
| Cost structure | Per-token, scales with usage | Fixed GPU/CPU cost, scales with uptime not usage |
| Time to first embedding | Minutes (API key) | Hours to days (serving infra, batching, autoscaling) |
| Who owns uptime | Vendor | Your team |
| Data residency | Leaves your network unless using a VPC/private offering | Stays entirely in your environment |
| Best at | Low to moderate, spiky volume | Very high, steady volume, or strict data rules |
If your data can't leave your network for compliance reasons, that single constraint usually overrides the cost math entirely, go open-weight regardless of volume. Our guide on choosing AI knowledge base software covers the retrieval-grounding and permission-aware search questions that sit one layer above this decision.
Benchmarks: What the Vendors Document
The Massive Text Embedding Benchmark (MTEB) is the standard reference for retrieval quality, but its live leaderboard shifts as new models land, so the table below sticks to what each vendor documents about its own model, labeled as a vendor claim rather than an independent rank.
| Model | What the Vendor Documents | Source |
|---|---|---|
| Qwen3-Embedding-8B | Reports 70.58 on MTEB multilingual (self-reported, June 2025) | Qwen3-Embedding-8B card |
| BGE-M3 | Documents hybrid dense/sparse/multi-vector retrieval as its differentiator, not a single leaderboard score | BGE-M3 card |
| Voyage voyage-4 family | Markets retrieval quality improvements over prior voyage-3 generation; no third-party score cited on the pricing page itself | Voyage pricing docs |
| OpenAI text-embedding-3 | Documents MRL dimension truncation with "minimal" quality loss; no competitive benchmark table on the pricing page | OpenAI pricing |
If a benchmark score matters to your decision, pull it from the live MTEB leaderboard yourself at evaluation time rather than trusting any single article, including this one, to have the current number.
How to Choose: Decision Framework
| If you need... | Pick | Why |
|---|---|---|
| The cheapest possible rate at huge scale | Titan Text Embeddings V2 or OpenAI text-embedding-3-small | Both sit at $0.02/M with no free-tier ceiling to track |
| Best quality-per-dollar with room to prototype free | Voyage voyage-4 or voyage-4-large | 200M free tokens covers most pilots entirely |
| One call for text, images, video, and PDFs | Gemini gemini-embedding-2 | Only model here with native multimodal unification |
| Strict data residency, nothing leaves your network | BGE-M3 or Qwen3-Embedding (self-hosted) | Open weight, run inside your own VPC |
| Code-specific retrieval | Voyage voyage-code-4 or Mistral Codestral Embed | Purpose-trained on code and technical queries |
| To decouple indexing cost from query cost | Cohere Embed 5 (Pro index, Fast query) | Shared embedding space, no re-indexing needed |
| Long documents without aggressive pre-chunking | Voyage voyage-4 (32K) or Jina v5-text-small (32,768) | Far longer context than OpenAI, Titan, or Mistral (8K) |
| To minimize future forced re-embeds | Any model with Matryoshka support (OpenAI v3, Voyage v4, Cohere 5, Qwen3) | Dimension changes don't require a full re-embed |
| An AWS-native bill with existing IAM controls | Amazon Titan Text Embeddings V2 | Lands on the Bedrock invoice you already reconcile |
Stage and Team Fit
| Model | Ideal Team | Signal |
|---|---|---|
| OpenAI text-embedding-3 | Startups to mid-market on OpenAI for generation | Already paying for GPT API access |
| Cohere Embed 5 | Mid-market to enterprise with real query volume | Query cost matters as much as indexing cost |
| Voyage AI voyage-4 | Startups piloting RAG, MongoDB Atlas users | Wants to prototype at zero marginal cost |
| Jina Embeddings v5 | Budget-conscious teams needing multilingual reach | Small footprint, broad language coverage |
| Gemini Embedding 2 | Google Cloud or Vertex AI shops | Needs multimodal in one embedding call |
| Mistral Embed | EU-based or Mistral-stack teams | Data residency or vendor preference in Europe |
| AWS Titan V2 | AWS-committed orgs | Wants one cloud bill, one IAM model |
| BGE-M3 | Technical teams with GPU infra | Needs hybrid dense+sparse+ColBERT, data sovereignty |
| E5 / Nomic Embed | Resource-constrained self-hosters | Wants the most battle-tested or smallest open option |
| Qwen3-Embedding | Technical teams chasing open-weight benchmark leadership | Has GPU budget for an 8B model |
For the agent-building side, when an agent should call retrieval versus a tool, see RAG for AI agents and what agentic RAG changes. Deciding which model powers the generation step on top of these embeddings? Claude vs ChatGPT vs Gemini covers that half of the stack. Once your pipeline is live, AI agent observability tools cover catching retrieval regressions before users do, and developers assembling the surrounding stack can check AI agent frameworks for developers, several of which ship default integrations for the models above.
Frequently Asked Questions about Embedding Models
What's the actual difference between a 1,024-dimension and a 3,072-dimension embedding?
More dimensions let a model encode finer distinctions in meaning, which can improve retrieval accuracy, but every added dimension costs more to store and search. Models with Matryoshka Representation Learning (OpenAI's v3, Voyage's v4, Cohere's Embed 5, Gemini's gemini-embedding-2, Qwen3-Embedding) let you truncate a high-dimension vector to a cheaper size with minimal quality loss, so you can test both and pick the smallest one that still hits your accuracy bar.
Should I use a hosted API or self-host an open-weight model like BGE-M3 or Qwen3-Embedding?
Self-hosting removes the per-token fee but adds a fixed GPU bill and the operational work of serving, scaling, and monitoring the model yourself. It pays off at very high steady-state volume or when data residency rules mean your text genuinely cannot leave your network. At low to moderate volume, a metered API is almost always cheaper once you count engineering time.
What happens to my vector database when I switch embedding models or vendors?
Every vector is tied to the exact model and version that generated it, so switching requires regenerating every vector from the original source text, not converting one model's output into another's. Google's own documentation confirms this applies even within the same vendor: gemini-embedding-001 and gemini-embedding-2 produce incompatible spaces. Budget for both the usually-small re-embedding fee and the larger engineering cost of a dual-write migration.
Is OpenAI's text-embedding-3-small good enough, or do I need text-embedding-3-large?
text-embedding-3-small at $0.02 per million tokens is the right default for most general-purpose retrieval, and its output truncates further if you need to save storage. Reach for text-embedding-3-large's 3,072 dimensions in a nuanced domain (legal, scientific, technical) where retrieval misses are expensive, and test the quality difference against your own corpus before paying the 6.5x premium.
Can I index with one embedding model and query with a cheaper one?
Only if the vendor explicitly guarantees the two models share an embedding space. Cohere's Embed 5 Pro and Fast are built for exactly this: index with Pro for quality, query with Fast for speed and lower cost. Don't assume this works across unrelated models or vendors; mismatched embedding spaces return meaningless similarity scores.
Are open-weight embedding models actually competitive with the commercial APIs?
On paper, yes. Qwen3-Embedding's own published claims put its 8B variant near the top of the MTEB multilingual leaderboard, and BGE-M3 offers hybrid dense, sparse, and multi-vector retrieval most commercial APIs don't expose. The gap isn't model quality, it's operational: you take on serving, scaling, and uptime a hosted API handles for you.
How much does it cost to embed a typical internal knowledge base?
For a 10-million-token corpus (roughly 15,000 to 25,000 pages) plus 2 million query tokens a month, year-one costs range from $0 on Voyage AI's free tier to around $6.80 on Gemini's multimodal flagship, with most mid-tier options landing between $0.68 and $4.42. Run your own token count through the pricing table above before budgeting.
Why does context length matter if I'm chunking my documents anyway?
A longer context window lets you embed bigger, more coherent chunks (a full section instead of three sentences), which usually improves retrieval because the model has more surrounding context. Models capped at 512 tokens, like E5 and Nomic Embed v2-MoE, force smaller chunks and more vectors per document, raising both vector count and storage bill for the same source corpus.
What to Do Next
Don't pick an embedding model off a benchmark score. Pull a representative sample (500 to 1,000 real chunks from your actual corpus), embed it with your top two or three candidates, run your real queries against each index, and compare retrieval precision on what matters for your use case. Then price the full workload, not the headline rate, using the real-workload table above as your template. Only after both checks pass should you commit, because the cost of being wrong isn't the embedding bill, it's the re-embedding project six months later when you find a better option after your corpus has already grown. The retrieval layer these embeddings feed has its own pricing quirks, which the RAG tools roundup works through on one workload.

On this page
- Key Facts
- Quick Comparison Table
- What Embeddings Actually Cost: Why the Headline Rate Lies
- Pricing at a Real Workload (10M Indexed + 2M/Month Queried)
- The Re-Embedding Cost Nobody Prices (The Real Lock-In)
- 1. OpenAI text-embedding-3-small and text-embedding-3-large
- 2. Cohere Embed 5
- 3. Voyage AI (voyage-4 family)
- 4. Jina Embeddings v5
- 5. Google Gemini Embedding (gemini-embedding-2)
- 6. Mistral Embed
- 7. Amazon Titan Text Embeddings V2
- 8. BAAI BGE-M3 (open weight)
- 9. intfloat E5 (multilingual-e5-large-instruct)
- 10. Nomic Embed v2-MoE (open weight)
- 11. Qwen3-Embedding (open weight)
- Dimensions, Context Length, and What They Cost You in Storage
- Open-Weight Models Head to Head
- API vs Self-Hosted: Who Should Actually Run the Open-Weight Models
- Benchmarks: What the Vendors Document
- How to Choose: Decision Framework
- Stage and Team Fit
- What to Do Next