Best Embedding Models in 2026: Pricing, Dimensions, and Real Workload Costs Compared

Turn this article into takeaways for your work.

Each assistant summarizes the article only for you and suggests best practices for your work.

There's no single best embedding model in 2026, only a best one for your corpus size, context length needs, and the cloud bill you're already paying. OpenAI's text-embedding-3-small is still the default nobody gets questioned for picking. Voyage AI's voyage-4 family usually wins on quality per dollar once you price a real workload, partly because its free tier is large enough to cover most pilots outright. Qwen3-Embedding and BGE-M3 are the open-weight options worth running yourself when a third-party API call on your data is a non-starter.

This guide prices eleven commercial and open-weight embedding families at a concrete workload (10 million tokens indexed once, plus 2 million query tokens a month) instead of the advertised headline rate, because that's where the real ranking lives. Figures marked (reported) are approximate. This is a technical buying decision, not just a benchmark one: who runs it, what switching costs later, and whether the vendor still exists in its current form a year from now all matter more than a leaderboard rank.

Key Facts

  • Google's own docs state that gemini-embedding-001 and the newer gemini-embedding-2 produce incompatible vector spaces, so upgrading forces a full re-embed of existing data (source: Google AI for Developers).
  • Cohere's Embed 5 Pro and Fast models share one embedding space, so you can index with Pro and query with the cheaper Fast model without re-indexing (source: Cohere, "Embed 5", 30 September 2026).
  • MongoDB acquired Voyage AI for roughly $220 million in February 2025 and has folded its models into Atlas Vector Search, while Voyage's standalone API stays live (source: SiliconANGLE).
  • BAAI's BGE-M3 is MIT-licensed and the only model here that natively produces dense, sparse, and ColBERT-style multi-vector output from a single pass (source: BAAI/bge-m3).
  • Qwen's own model card reports Qwen3-Embedding-8B scoring 70.58 on MTEB multilingual as of its June 2025 release, a vendor claim rather than an independent ranking (source: Qwen/Qwen3-Embedding-8B).

Quick Comparison Table

Model Best For Price (per 1M tokens) Max Dimensions Max Context Key Limitation
OpenAI text-embedding-3-small/large One vendor bill alongside GPT API usage $0.02 / $0.13 1,536 / 3,072 8,191 tokens No native multimodal input
Cohere Embed 5 (Pro/Fast) Splitting cheap query traffic from quality indexing $0.12 / $0.08 2,048 128,000 tokens No published always-free tier
Voyage AI voyage-4 family Best retrieval quality per dollar, large free tier $0.02 to $0.12 (200M free) 2,048 32,000 tokens Now MongoDB-owned, roadmap tied to Atlas
Jina Embeddings v5 Small multilingual footprint, low entry cost ~$0.05 (reported) 1,024 8,192 to 32,768 tokens Exact per-SKU rate needs dashboard login
Google Gemini Embedding 2 One call embeds text, image, video, audio, PDF $0.20 standard / $0.10 batch 3,072 (MRL) 8,192 tokens Incompatible with prior gemini-embedding-001
Mistral Embed EU-based or Mistral-stack teams $0.10 1,024 (fixed) ~8,000 tokens No Matryoshka truncation offered
Amazon Titan Text Embeddings V2 AWS-native orgs consolidating the Bedrock bill $0.02 1,024 8,192 tokens Fewer independent benchmark citations
BAAI BGE-M3 (open weight) Self-hosted hybrid dense+sparse+ColBERT $0 API fee 1,024 8,192 tokens You own uptime, scaling, GPU cost
intfloat E5 (multilingual-e5-large-instruct) Well-documented open baseline $0 API fee 1,024 512 tokens Short context forces aggressive chunking
Nomic Embed v2-MoE (open weight) Small, efficient self-hosted multilingual model $0 API fee 768 (truncatable to 256) 512 tokens Same short context ceiling as E5
Qwen3-Embedding (open weight) Top open-weight benchmark performance $0 API fee 4,096 (configurable 32 to 4,096) 32,000 tokens 8B variant needs real GPU memory

What Embeddings Actually Cost: Why the Headline Rate Lies

Every vendor publishes a per-million-token rate, and most comparisons stop there. That number is close to meaningless alone, because the real drivers of your bill (how much text you index once, how many queries you run monthly, and whether a free tier absorbs either) vary by an order of magnitude between a pilot and production. New to the concept? Our plain-English explainer on embeddings covers it in two minutes, and vector databases covers where those numbers end up living.

The workload below is deliberately modest: 10 million tokens embedded once (roughly 15,000 to 25,000 pages, depending on chunk size) plus 2 million query tokens a month, a realistic volume for an internal knowledge base or a mid-size SaaS product's RAG layer feeding a team's AI agents. Annualized, that's 34 million total tokens in year one.

Pricing at a Real Workload (10M Indexed + 2M/Month Queried)

Model Indexing (10M tokens) Querying (24M tokens/yr) Year-One Total
OpenAI text-embedding-3-small $0.20 $0.48 $0.68
OpenAI text-embedding-3-large $1.30 $3.12 $4.42
Cohere Embed 5 (Pro index / Fast query) $1.20 $1.92 $3.12
Voyage AI voyage-4 $0.00 (within 200M free) $0.00 (within 200M free) $0.00
Voyage AI voyage-4-large $0.00 (within 200M free) $0.00 (within 200M free) $0.00
Jina Embeddings v5-text-small $0.00 (10M free tier) $1.20 (reported) ~$1.20 (reported)
Google gemini-embedding-2 (standard) $2.00 $4.80 $6.80
Google gemini-embedding-001 (reported) $1.50 $3.60 $5.10 (reported)
Mistral Embed $1.00 $2.40 $3.40
Amazon Titan Text Embeddings V2 $0.20 $0.48 $0.68

Open-weight models (BGE-M3, E5, Nomic, Qwen3-Embedding) carry no per-token fee here; see the self-hosting section for their real cost.

This is the flip the headline rate hides. Gemini's new multimodal flagship costs ten times what Titan V2 or OpenAI's small model charges for the identical text-only workload, and Voyage's free tier means a huge share of pilots, and even some small production deployments, pay nothing for models that independently rank near the top of the field. Past 200 million lifetime tokens on Voyage, the free allowance stops and you're back to $0.02 to $0.12 per million depending on tier.

The Re-Embedding Cost Nobody Prices (The Real Lock-In)

Picking an embedding model isn't a one-time decision. Every vector you store is tied to the exact model and version that produced it, and swapping vendors, or even upgrading within the same vendor, means regenerating every vector from scratch. Google's own documentation says it outright: gemini-embedding-001 and gemini-embedding-2 produce incompatible embedding spaces, so the newer model requires re-embedding everything indexed under the old one. That's not a hypothetical, it's the vendor's own migration note.

Model Cost to Re-Embed a 10M-Token Corpus
OpenAI text-embedding-3-small $0.20
OpenAI text-embedding-3-large $1.30
Cohere Embed 5 Pro $1.20
Voyage AI voyage-4 $0.00 to $0.60 (depends on remaining free tier)
Jina Embeddings v5-text-small $0.50 (reported)
Google gemini-embedding-2 $2.00
Mistral Embed $1.00
Amazon Titan Text Embeddings V2 $0.20
Open-weight (BGE-M3, E5, Nomic, Qwen3) $0 API fee, full GPU re-inference pass

The dollar figure is almost always trivial, which is exactly why it gets ignored during vendor selection. The real cost is operational: re-chunking source documents if your strategy changed too, standing up a parallel index, running a dual-write cutover so search doesn't go dark mid-migration, and re-validating retrieval quality before you delete the old vectors. That's a multi-day project triggered by an upgrade you didn't initiate, and it recurs every time a vendor ships a new generation. Favor a family with Matryoshka truncation or clear version-compatibility notes (OpenAI's v3, Voyage's v4, Cohere's Embed 5, Qwen3-Embedding) for fewer forced migrations.

1. OpenAI text-embedding-3-small and text-embedding-3-large

OpenAI's third-generation embeddings remain the default for teams who already pay OpenAI for chat completions and don't want a second vendor relationship. text-embedding-3-small produces 1,536-dimension vectors at $0.02 per million tokens; text-embedding-3-large goes up to 3,072 dimensions at $0.13 per million, both truncatable via Matryoshka learning with minimal quality loss. The Batch API halves both rates for workloads that tolerate a 24-hour turnaround, which is most bulk-indexing jobs.

text-embedding-3-small text-embedding-3-large
Price (per 1M tokens) $0.02 $0.13
Dimensions 1,536 (truncatable) 3,072 (truncatable)
Context length 8,191 tokens 8,191 tokens

Best for: teams standardized on OpenAI for generation wanting one API key and invoice. Not ideal for: anyone needing multimodal embeddings or EU-only processing without the Enterprise tier. Source: OpenAI API pricing.

2. Cohere Embed 5

Cohere shipped Embed 5 on 30 September 2026, replacing the Embed v3/v4 line with Pro and Fast variants sharing one embedding space. That shared space is the headline feature: index once with Pro for quality, then query with the cheaper, 2.4x-faster Fast model without re-indexing. Both support a 128,000-token context, dramatically longer than any other model here, and dimensions down to 256 via Matryoshka compression.

Best for: teams with real query volume who want to decouple indexing cost from serving cost. Not ideal for: teams wanting a transparent self-serve free tier to prototype against before committing, since Cohere doesn't publish one comparable to Voyage's. Source: Cohere, "Embed 5".

3. Voyage AI (voyage-4 family)

Voyage AI, acquired by MongoDB for roughly $220 million in February 2025, still runs a standalone API and ships the voyage-4 family: voyage-4 ($0.06/M), voyage-4-large ($0.12/M), voyage-4-lite ($0.02/M), plus voyage-code-4 for code and voyage-context-4 for contextualized chunk embeddings, all with the first 200 million tokens free per account. That free tier covers the entire 10M-plus-24M-token workload modeled above at zero cost, the single biggest factor in Voyage's price-to-quality reputation. Every v4 model supports configurable dimensions (256, 512, 1,024, 2,048) and an 8x longer context window than OpenAI or Titan at 32,000 tokens.

Best for: teams prototyping or running mid-size production RAG wanting generous free usage and flexible dimension sizing. Not ideal for: teams wary of a roadmap now controlled by a database vendor rather than an independent AI lab. Source: Voyage AI pricing docs.

4. Jina Embeddings v5

Jina's v5 series ships four variants built on the Qwen3 and EuroBERT tokenizers: v5-text-small (1,024 dimensions, 32,768-token context, 119 languages) and v5-text-nano (768 dimensions, 8,192-token context, 15-plus languages), plus multimodal omni variants. New accounts get 10 million free tokens. Jina's public page doesn't post an exact per-model rate (you need a dashboard login); the text-small SKU reportedly runs about $0.05 per million tokens.

Best for: budget-conscious teams needing broad multilingual coverage in a small footprint. Not ideal for: teams needing a published, citable price before budget sign-off, since Jina's rate requires a logged-in dashboard. Source: Jina Embeddings v5-text announcement for specs.

5. Google Gemini Embedding (gemini-embedding-2)

Google's current flagship, gemini-embedding-2, is a genuine category shift: Google's first multimodal embedding model, mapping text, images, video, audio, and PDFs into one unified vector space in a single call. Standard pricing is $0.20 per million tokens for text, dropping to $0.10 on the Batch tier, with image, audio, and video priced separately and considerably higher. Dimensions run from 128 to 3,072 via Matryoshka learning, but context is a relatively tight 8,192 tokens. The older text-only gemini-embedding-001 (2,048-token context) is still documented and reportedly priced around $0.15 per million standard, but Google states plainly the two models produce incompatible embedding spaces, so there's no free upgrade path.

Best for: teams on Google Cloud or Vertex AI wanting one call to cover text, images, and PDFs together. Not ideal for: anyone who embedded a corpus under gemini-embedding-001 and assumed an in-place upgrade; budget a full re-embed. Source: Gemini API pricing and Gemini Embedding docs.

6. Mistral Embed

Mistral Embed produces fixed 1,024-dimension vectors from a roughly 8,000-token context window at $0.10 per million tokens, with a code-focused sibling, Codestral Embed, at $0.15 per million. Unlike every other commercial model here, Mistral doesn't offer a selectable or Matryoshka-truncatable dimension: 1,024 or nothing, which simplifies storage planning but removes a cost lever every competitor gives you.

Best for: EU-headquartered teams or shops already standardized on Mistral's LLMs who want one European vendor relationship. Not ideal for: teams that want to shrink vector storage later without re-embedding, since there's no truncation option. Source: Mistral AI API pricing.

7. Amazon Titan Text Embeddings V2

Titan Text Embeddings V2 is AWS's Bedrock-native option: $0.02 per million input tokens (matching OpenAI's small model), flexible output at 256, 512, or 1,024 dimensions, and an 8,192-token context window. Its appeal isn't benchmark leadership, it's that the charge lands on the same AWS bill as the rest of your infrastructure, under the IAM, VPC, and procurement controls you already have for Bedrock.

Best for: AWS-committed teams avoiding a second vendor contract and data-processing agreement. Not ideal for: teams chasing the top of independent retrieval benchmarks, where Titan is rarely cited as a leader against OpenAI, Voyage, or the open-weight frontier. Source: AWS, "Get started with Amazon Titan Text Embeddings V2".

8. BAAI BGE-M3 (open weight)

BGE-M3 is worth knowing even if you never self-host anything else: it's the only model here producing dense, sparse (lexical), and ColBERT-style multi-vector output from a single pass, so one model powers both semantic and keyword-style hybrid search. It's MIT-licensed, outputs 1,024-dimension dense vectors, handles up to 8,192 tokens, and covers more than 100 languages.

Best for: technical teams with GPU infrastructure wanting hybrid dense-plus-sparse retrieval and strict data sovereignty. Not ideal for: teams without the appetite to own uptime, autoscaling, and GPU cost themselves. Source: BAAI/bge-m3 model card.

9. intfloat E5 (multilingual-e5-large-instruct)

E5 is the open-weight baseline most RAG tooling defaults to when a tutorial needs "a free embedding model." The multilingual-e5-large-instruct variant is MIT-licensed, outputs 1,024-dimension vectors across 94 languages, and integrates cleanly with sentence-transformers and most vector database SDKs. Its real limitation is the 512-token context window, short enough to force aggressive chunking, well below the page-level chunks Voyage or Cohere allow.

Best for: teams wanting a thoroughly documented, widely supported open baseline. Not ideal for: long-document retrieval, where 512 tokens forces many small, context-poor chunks. Source: intfloat/multilingual-e5-large-instruct model card.

10. Nomic Embed v2-MoE (open weight)

Nomic's v2-MoE is a mixture-of-experts model with 475 million total parameters but only 305 million active at inference, cheaper to run than its quality class suggests. It outputs 768-dimension vectors, truncatable to 256 via Matryoshka learning for roughly a 3x storage reduction, trained across roughly 100 languages on 1.6 billion pairs. Like E5, it caps out at a 512-token context window.

Best for: teams wanting a small, GPU-efficient self-hosted model for resource-constrained deployments. Not ideal for: the same long-document use case E5 struggles with; the short context is shared across most of the open "base" generation of embedding models. Source: nomic-ai/nomic-embed-text-v2-moe model card.

11. Qwen3-Embedding (open weight)

Qwen3-Embedding is the strongest open-weight performer here on paper: it ships in 0.6B, 4B, and 8B sizes, supports a 32,000-token context matching Voyage, and lets you configure output dimensions from 32 up to 4,096 depending on the size you run. It's Apache 2.0 licensed and covers over 100 languages. Qwen's own model card reports the 8B variant scoring 70.58 on MTEB multilingual as of its June 2025 release, a vendor-published figure rather than an independent rank.

Best for: technical teams with GPU budget who want the best open-weight retrieval quality available and full control over dimension sizing. Not ideal for: teams without the infrastructure to serve an 8B-parameter model at low latency; the smaller 0.6B and 4B variants trade some of that quality back for easier hosting. Source: Qwen/Qwen3-Embedding-8B model card.

Dimensions, Context Length, and What They Cost You in Storage

Dimension count decides your vector database bill independent of what you pay the embedding vendor. A higher-dimension vector captures more nuance but costs proportionally more disk, memory, and latency to store and query, which is why Matryoshka support matters as much as the raw benchmark score. The context window a model accepts matters just as much on the input side: a short window forces smaller chunks, which means more chunks and more vectors, for the same source corpus.

Model Dimensions Max Context Multilingual Matryoshka/Truncation
OpenAI text-embedding-3-large up to 3,072 8,191 tokens Yes Yes
Cohere Embed 5 up to 2,048 128,000 tokens 100+ languages Yes
Voyage voyage-4 family up to 2,048 32,000 tokens Yes (multilingual-2 legacy) Yes
Jina v5-text-small 1,024 32,768 tokens 119 languages Partial (model-dependent)
Gemini gemini-embedding-2 up to 3,072 8,192 tokens 100+ languages Yes
Mistral Embed 1,024 (fixed) ~8,000 tokens Not emphasized No
AWS Titan Text Embeddings V2 256 / 512 / 1,024 8,192 tokens Yes Selectable, not Matryoshka-trained
BAAI BGE-M3 1,024 8,192 tokens 100+ languages No (fixed dense output)
intfloat E5 (multilingual-large) 1,024 512 tokens 94 languages No
Nomic Embed v2-MoE 768 512 tokens ~100 languages Yes (to 256)
Qwen3-Embedding-8B up to 4,096 32,000 tokens 100+ languages Yes (32 to 4,096)

To make the storage math concrete: at 32-bit float precision, a million vectors at 3,072 dimensions runs about 12.3 GB, the same million at 1,024 dimensions runs about 4.1 GB, and truncated to 256 via Matryoshka, about 1 GB. That's arithmetic (4 bytes per dimension, times dimension count, times vector count), not a vendor claim, but it's what determines whether your vector database bill scales linearly or explosively as your corpus grows. Model the storage cost at your chosen dimension before you commit, not after. The vector database roundup prices 12 engines at a 1-million-vector workload, so you can run that math against real storage rates.

Open-Weight Models Head to Head

BGE-M3 E5 (multilingual-large-instruct) Nomic Embed v2-MoE Qwen3-Embedding-8B
License MIT MIT Apache 2.0 Apache 2.0
Parameters Not disclosed as MoE ~560M (24-layer) 475M total, 305M active (MoE) 8B
Dimensions 1,024 1,024 768 (to 256) up to 4,096
Context 8,192 tokens 512 tokens 512 tokens 32,000 tokens
Distinct capability Dense + sparse + ColBERT in one pass Broadest tooling support Smallest active-parameter footprint Longest context, largest dimension ceiling

API vs Self-Hosted: Who Should Actually Run the Open-Weight Models

Self-hosting BGE-M3, E5, Nomic, or Qwen3-Embedding removes the per-token fee, but it doesn't remove cost, it moves it onto your infrastructure budget and your team's time. A GPU serving one of these models around the clock costs real money regardless of volume processed, so the trade pays off once token volume exceeds what a metered API would charge for a dedicated instance, or once data residency rules make an external API call a non-starter regardless of price.

Factor Hosted API (OpenAI, Voyage, Cohere, etc.) Self-hosted open weight
Cost structure Per-token, scales with usage Fixed GPU/CPU cost, scales with uptime not usage
Time to first embedding Minutes (API key) Hours to days (serving infra, batching, autoscaling)
Who owns uptime Vendor Your team
Data residency Leaves your network unless using a VPC/private offering Stays entirely in your environment
Best at Low to moderate, spiky volume Very high, steady volume, or strict data rules

If your data can't leave your network for compliance reasons, that single constraint usually overrides the cost math entirely, go open-weight regardless of volume. Our guide on choosing AI knowledge base software covers the retrieval-grounding and permission-aware search questions that sit one layer above this decision.

Benchmarks: What the Vendors Document

The Massive Text Embedding Benchmark (MTEB) is the standard reference for retrieval quality, but its live leaderboard shifts as new models land, so the table below sticks to what each vendor documents about its own model, labeled as a vendor claim rather than an independent rank.

Model What the Vendor Documents Source
Qwen3-Embedding-8B Reports 70.58 on MTEB multilingual (self-reported, June 2025) Qwen3-Embedding-8B card
BGE-M3 Documents hybrid dense/sparse/multi-vector retrieval as its differentiator, not a single leaderboard score BGE-M3 card
Voyage voyage-4 family Markets retrieval quality improvements over prior voyage-3 generation; no third-party score cited on the pricing page itself Voyage pricing docs
OpenAI text-embedding-3 Documents MRL dimension truncation with "minimal" quality loss; no competitive benchmark table on the pricing page OpenAI pricing

If a benchmark score matters to your decision, pull it from the live MTEB leaderboard yourself at evaluation time rather than trusting any single article, including this one, to have the current number.

How to Choose: Decision Framework

If you need... Pick Why
The cheapest possible rate at huge scale Titan Text Embeddings V2 or OpenAI text-embedding-3-small Both sit at $0.02/M with no free-tier ceiling to track
Best quality-per-dollar with room to prototype free Voyage voyage-4 or voyage-4-large 200M free tokens covers most pilots entirely
One call for text, images, video, and PDFs Gemini gemini-embedding-2 Only model here with native multimodal unification
Strict data residency, nothing leaves your network BGE-M3 or Qwen3-Embedding (self-hosted) Open weight, run inside your own VPC
Code-specific retrieval Voyage voyage-code-4 or Mistral Codestral Embed Purpose-trained on code and technical queries
To decouple indexing cost from query cost Cohere Embed 5 (Pro index, Fast query) Shared embedding space, no re-indexing needed
Long documents without aggressive pre-chunking Voyage voyage-4 (32K) or Jina v5-text-small (32,768) Far longer context than OpenAI, Titan, or Mistral (8K)
To minimize future forced re-embeds Any model with Matryoshka support (OpenAI v3, Voyage v4, Cohere 5, Qwen3) Dimension changes don't require a full re-embed
An AWS-native bill with existing IAM controls Amazon Titan Text Embeddings V2 Lands on the Bedrock invoice you already reconcile

Stage and Team Fit

Model Ideal Team Signal
OpenAI text-embedding-3 Startups to mid-market on OpenAI for generation Already paying for GPT API access
Cohere Embed 5 Mid-market to enterprise with real query volume Query cost matters as much as indexing cost
Voyage AI voyage-4 Startups piloting RAG, MongoDB Atlas users Wants to prototype at zero marginal cost
Jina Embeddings v5 Budget-conscious teams needing multilingual reach Small footprint, broad language coverage
Gemini Embedding 2 Google Cloud or Vertex AI shops Needs multimodal in one embedding call
Mistral Embed EU-based or Mistral-stack teams Data residency or vendor preference in Europe
AWS Titan V2 AWS-committed orgs Wants one cloud bill, one IAM model
BGE-M3 Technical teams with GPU infra Needs hybrid dense+sparse+ColBERT, data sovereignty
E5 / Nomic Embed Resource-constrained self-hosters Wants the most battle-tested or smallest open option
Qwen3-Embedding Technical teams chasing open-weight benchmark leadership Has GPU budget for an 8B model

For the agent-building side, when an agent should call retrieval versus a tool, see RAG for AI agents and what agentic RAG changes. Deciding which model powers the generation step on top of these embeddings? Claude vs ChatGPT vs Gemini covers that half of the stack. Once your pipeline is live, AI agent observability tools cover catching retrieval regressions before users do, and developers assembling the surrounding stack can check AI agent frameworks for developers, several of which ship default integrations for the models above.

Frequently Asked Questions about Embedding Models

What's the actual difference between a 1,024-dimension and a 3,072-dimension embedding?

More dimensions let a model encode finer distinctions in meaning, which can improve retrieval accuracy, but every added dimension costs more to store and search. Models with Matryoshka Representation Learning (OpenAI's v3, Voyage's v4, Cohere's Embed 5, Gemini's gemini-embedding-2, Qwen3-Embedding) let you truncate a high-dimension vector to a cheaper size with minimal quality loss, so you can test both and pick the smallest one that still hits your accuracy bar.

Should I use a hosted API or self-host an open-weight model like BGE-M3 or Qwen3-Embedding?

Self-hosting removes the per-token fee but adds a fixed GPU bill and the operational work of serving, scaling, and monitoring the model yourself. It pays off at very high steady-state volume or when data residency rules mean your text genuinely cannot leave your network. At low to moderate volume, a metered API is almost always cheaper once you count engineering time.

What happens to my vector database when I switch embedding models or vendors?

Every vector is tied to the exact model and version that generated it, so switching requires regenerating every vector from the original source text, not converting one model's output into another's. Google's own documentation confirms this applies even within the same vendor: gemini-embedding-001 and gemini-embedding-2 produce incompatible spaces. Budget for both the usually-small re-embedding fee and the larger engineering cost of a dual-write migration.

Is OpenAI's text-embedding-3-small good enough, or do I need text-embedding-3-large?

text-embedding-3-small at $0.02 per million tokens is the right default for most general-purpose retrieval, and its output truncates further if you need to save storage. Reach for text-embedding-3-large's 3,072 dimensions in a nuanced domain (legal, scientific, technical) where retrieval misses are expensive, and test the quality difference against your own corpus before paying the 6.5x premium.

Can I index with one embedding model and query with a cheaper one?

Only if the vendor explicitly guarantees the two models share an embedding space. Cohere's Embed 5 Pro and Fast are built for exactly this: index with Pro for quality, query with Fast for speed and lower cost. Don't assume this works across unrelated models or vendors; mismatched embedding spaces return meaningless similarity scores.

Are open-weight embedding models actually competitive with the commercial APIs?

On paper, yes. Qwen3-Embedding's own published claims put its 8B variant near the top of the MTEB multilingual leaderboard, and BGE-M3 offers hybrid dense, sparse, and multi-vector retrieval most commercial APIs don't expose. The gap isn't model quality, it's operational: you take on serving, scaling, and uptime a hosted API handles for you.

How much does it cost to embed a typical internal knowledge base?

For a 10-million-token corpus (roughly 15,000 to 25,000 pages) plus 2 million query tokens a month, year-one costs range from $0 on Voyage AI's free tier to around $6.80 on Gemini's multimodal flagship, with most mid-tier options landing between $0.68 and $4.42. Run your own token count through the pricing table above before budgeting.

Why does context length matter if I'm chunking my documents anyway?

A longer context window lets you embed bigger, more coherent chunks (a full section instead of three sentences), which usually improves retrieval because the model has more surrounding context. Models capped at 512 tokens, like E5 and Nomic Embed v2-MoE, force smaller chunks and more vectors per document, raising both vector count and storage bill for the same source corpus.

What to Do Next

Don't pick an embedding model off a benchmark score. Pull a representative sample (500 to 1,000 real chunks from your actual corpus), embed it with your top two or three candidates, run your real queries against each index, and compare retrieval precision on what matters for your use case. Then price the full workload, not the headline rate, using the real-workload table above as your template. Only after both checks pass should you commit, because the cost of being wrong isn't the embedding bill, it's the re-embedding project six months later when you find a better option after your corpus has already grown. The retrieval layer these embeddings feed has its own pricing quirks, which the RAG tools roundup works through on one workload.

About the author

Camellia

Camellia

Principal Product Marketing Strategist

Camellia is Principal Product Marketing Strategist at Rework, helping B2B buyers pick the right software with confidence. With 6+ years in product marketing and 150+ SaaS tools evaluated across CRM, project management, and sales engagement, Camellia turns competitive intelligence into clear, honest comparisons. Readers get vendor evaluations they can trust to cut through marketing noise and decide faster.