Pinecone vs Weaviate vs Qdrant: The Vector Database Shortlist for 2026

Turn this article into takeaways for your work.

Each assistant summarizes the article only for you and suggests best practices for your work.

If you're building a retrieval-augmented generation pipeline or giving an agent long-term memory, these three names show up on almost every shortlist. Pinecone sells the fully managed, no-ops version of a vector database. Weaviate sells an open-core database with a managed cloud wrapped around it. Qdrant sells a fully open-source engine first, with a managed cloud as one deployment option among several. All three do the core job: store millions of embeddings and return the nearest ones fast. What actually differs is how each one bills you, what you're allowed to run yourself, and what happens to the bill when corpus and traffic both grow at once.

This is written for the person who has to make the call and live with it: a technical founder, a head of engineering, or a CTO who needs this in production, not a benchmark leaderboard. We fetched each vendor's pricing page on 2 October 2026, built one workload, priced it across all three, then repeated it at roughly ten times the scale. The ranking changes. That's the part worth your time. If you only need Pinecone against Weaviate, Pinecone vs Weaviate drops Qdrant and goes deeper on compliance, vendor risk, and built-in vectorizers.

TL;DR

Pinecone Weaviate Qdrant
Deployment model Fully managed SaaS only Open-core, self-managed or Weaviate Cloud Fully open source, self-managed, Cloud, or Hybrid Cloud
License Proprietary (no self-hosting) BSD-3-Clause core, newer enterprise features under a separate commercial license Apache 2.0, entire engine
Pricing model Pay-per-read-unit and write-unit, scales with index size times query volume Pay for stored vector dimensions plus object storage, flat tiers with a monthly floor Pay for provisioned vCPU, RAM, and disk; self-host for infrastructure cost only
Cheapest at small scale (1M vectors, 50K queries/month) Builder plan, $20/month flat, covers it Flex plan floor, $45/month Free tier RAM cap is too small; needs a small paid cluster or self-hosted ~5 to 6 GB RAM box
Cost behavior at 10x scale Jumps hardest, because read cost multiplies corpus size by query count Grows linearly with corpus size, stays near the same floor in our model Grows linearly with RAM needed; cloud rate isn't published, self-host stays license-free
Best known for Zero-ops serverless simplicity Native hybrid search and GraphQL-style filtering Self-hostable performance and the deepest quantization options
Self-hosting Not available at any tier Yes, BSD-3-Clause core, enterprise features gated Yes, full engine, Apache 2.0

What Each One Is Actually Built For

Pinecone is the database for a team that does not want to think about infrastructure at all. There's no cluster to size, no quantization to configure, no self-hosted fallback. You create a serverless index, pick a cloud and region, and Pinecone handles sharding, compaction, and scaling behind an API. That simplicity is the entire pitch, and the ceiling too: you're fully dependent on Pinecone's roadmap, regions, and pricing formula, with no exit ramp to your own infrastructure if the math stops working.

Weaviate sits in the middle. The core engine is open source (BSD-3-Clause) and you can run it yourself on a laptop, a Kubernetes cluster, or whatever you already operate. Weaviate Cloud is the managed version of that engine, and also the company's main channel for newer capabilities like its AI-native Query Agent. As of 2026, Weaviate gates some newer code (shipped in the repository's wl directory) under a separate commercial license rather than BSD-3-Clause, worth knowing before you assume every new feature ships fully open.

Qdrant is built open source first. Written in Rust, distributed entirely under Apache 2.0 with no carved-out enterprise directory, its managed Cloud, Hybrid Cloud (Qdrant-managed control plane on your own infrastructure), and Private Cloud tiers are deployment options layered on the same codebase everyone can run for free. If self-hosting without a licensing asterisk is the requirement, Qdrant is the only one of the three that clears it today.

For teams building RAG for an AI agent or implementing the RAG assistant pattern, the practical question isn't which engine returns the nearest neighbor fastest. All three are fast enough for the vast majority of production workloads. It's which billing model and deployment model survive your actual growth curve.

Licensing and Self-Hosting, Verified

Licensing is the fact that goes stale fastest here, as infrastructure vendors re-license.

Pinecone Weaviate Qdrant
License Proprietary, closed source BSD-3-Clause for the core (most of the repository); code in the wl directory is "Copyright Weaviate BV 2026-present, all rights reserved," available only under a separate enterprise license Apache License 2.0, full codebase
Self-hosting allowed No, SaaS only on Pinecone's infrastructure Yes, for everything under BSD-3-Clause; enterprise wl features require a commercial license key Yes, entire engine, no license key required
Source visible No Yes, including the wl enterprise code, but using it without a license isn't permitted Yes, fully
Managed options Serverless only Weaviate Cloud (Flex, Premium Shared, Premium Dedicated), Bring Your Own Cloud on Premium Qdrant Cloud, Hybrid Cloud (your infra, Qdrant-managed control plane), Private Cloud
What this means for you You accept full vendor lock-in from day one, no exit path to your own infra You can self-host the database itself for free; some newer platform features are commercial-only You keep the real option to run the exact production engine yourself at zero license cost

Pricing at a Real Workload (and the Flip)

Headline rates are close to useless here because all three meter completely differently: Pinecone on read and write units, Weaviate on stored vector dimensions plus object storage, Qdrant on provisioned compute. So we built one workload, priced it on each vendor's own published formula, then grew it 10x to see what breaks.

Scenario A, the common starting point: 1,000,000 vectors at 768 dimensions (a typical size for embeddings from OpenAI's text-embedding-3-small or similar models), 50,000 read queries a month, and about 15,000 upserts a month as the corpus gets refreshed. This is a realistic single-product RAG deployment, not a toy.

Scenario B, the same company a year later: 10,000,000 vectors, same 768 dimensions, 500,000 queries a month, about 150,000 upserts a month. Corpus and traffic both grew 10x, which is the normal shape of a RAG product that's actually working.

The figures below are our calculations from each vendor's own published rates and formulas, not a vendor-issued quote. We've rounded for readability; use each vendor's own calculator or console for a figure you'd budget against.

Pinecone Weaviate Qdrant
Storage needed About 3.1 GB of raw vector data (1M x 768 x 4 bytes), call it 3.4 GB with metadata Same underlying data, billed as 768,000,000 "vector dimensions" plus ~3 GiB object storage RAM needed per Qdrant's own sizing formula (vectors x dimensions x 4 bytes x 1.5 for HNSW overhead): about 4.3 GiB, plus 20% headroom, roughly 5.2 GiB
Which tier fits Builder plan. Storage, read units (170,000 RU/month at roughly 3.4 RU per query), and write units (under 50,000 WU) all land comfortably inside Builder's caps Free tier's 100,000-object cap is 10x too small; must use Flex Free tier's 1 GiB RAM cap is about 5x too small; needs a small paid cluster or a self-hosted box sized to ~6 GB RAM
Real metered cost $20/month flat (Builder's fixed rate covers the whole workload with room to spare) Dimensions ($0.00465/1M) plus storage: about $3.93/month in real usage, but Flex's monthly floor is $45, so you pay $45/month Qdrant Cloud's per-GB-RAM rate isn't published on its pricing page; it's generated through Qdrant's own calculator (cloud.qdrant.io/calculator). Self-hosted: $0 license cost, your own ~6 GB RAM instance
Pinecone Weaviate Qdrant
Storage needed About 30.7 GB raw, roughly 34 GB with metadata overhead 7,680,000,000 "vector dimensions," plus ~30 GiB object storage About 43 GiB RAM by the sizing formula, plus 20% headroom, roughly 52 GiB
Which tier fits Exceeds Builder's 10 GB storage cap and its 2M read-unit cap by a wide margin; must move to Standard Still fits under Flex (no published object-count ceiling on Flex itself, unlike Free) Needs a materially larger paid cluster, or a self-hosted box sized to ~52 GB RAM
Real metered cost Storage: 34 GB x $0.33/GB = $11.22. Reads: Standard's read-unit rate is 1 RU per GB of index size per query, so 500,000 queries x ~34 RU each = 17,000,000 RU, at $16 to $18 per million = $272 to $306. Writes: roughly 457,500 WU at $4 to $4.50 per million = $1.83 to $2.06. Total: roughly $285 to $319/month Dimensions: 7,680 x $0.00465 = $35.71. Storage: ~$3.60. Total: roughly $39/month, still under Flex's $45 floor, so you'd likely pay $45/month Same structural answer: no published per-unit cloud rate, self-hosted stays license-free regardless of scale

The flip, in plain terms: at the smaller scale, all three land in roughly the same $20 to $45 a month range, because minimums and flat tiers dominate actual usage. At 10x scale, Pinecone's bill grows roughly 14 to 16x (from $20 to $285 to $319), while our calculation shows Weaviate's growing only about linearly with vector count and staying near the same $45 floor. The reason isn't that Pinecone got more expensive per vector; it's that Pinecone's read-unit formula charges per query based on the size of the entire namespace being searched, so cost scales with corpus size multiplied by query volume, not added to it. Grow both 10x and the read bill alone can grow close to 100x before storage and writes are even counted. Weaviate's dimension-based rate doesn't show an explicit per-query multiplier on its pricing page, which makes it more predictable as query volume climbs, though confirm your real number in Weaviate Cloud's Running Costs console since the published page itself points you there.

Qdrant's limitation here is transparency, not cost: its Cloud pricing page names what you're billed for (vCPU, RAM, disk, backups, inference tokens) but not the rate, so you can't build a number without the calculator. What you can always fall back on, and what neither Pinecone nor Weaviate's core offers, is running the identical engine yourself under Apache 2.0 for the cost of your own compute. If you're already treating agent cost as a lever you actively manage, that self-hosting option is the real reason technical teams keep Qdrant on the shortlist even when its cloud sticker price is unclear upfront.

Pure vector search misses exact matches: product codes, part numbers, names, anything where the right answer uses specific tokens that semantic search alone can blur past. All three support combining dense vector search with keyword or sparse retrieval, but the mechanics differ.

Pinecone Weaviate Qdrant
Native hybrid search Yes, dense plus sparse vectors in a single index, or two linked indexes queried separately and merged client-side Yes, dense vector search fused with BM25F keyword search natively Yes, via the Query API's prefetch mechanism, combining named dense and sparse vectors in one request
Fusion method Alpha-weighted scaling of query vectors (you control the balance) Relative Score Fusion (default since v1.24) or Ranked Fusion, controlled by an alpha parameter from 0 (pure keyword) to 1 (pure vector) Reciprocal Rank Fusion (RRF, with configurable constant and weighting since v1.16/1.17) or Distribution-Based Score Fusion (DBSF)
Multi-stage retrieval Not a first-class concept; typically handled at the application layer Not built as a distinct multi-stage primitive Yes, native: use a cheap representation to get candidates, then re-score with a larger model or technique like ColBERT or Matryoshka embeddings
Known limitation The single-index approach produces unbounded sparse scores that need normalization, adding complexity versus pure-dense queries Multi-vector collections don't support diversity selection (MMR) alongside hybrid search The most configuration knobs of the three, which is powerful but means more decisions before you get predictable results

Metadata and Payload Filtering

A vector-only search with no filtering is rarely what production systems need. You almost always want to combine "find similar" with "and only from this tenant, this date range, this document type."

Pinecone Weaviate Qdrant
Filter integration Metadata filters apply within the query Pre-filtering: the eligible set is built before the vector search runs, using an inverted index plus HNSW; ACORN is the default strategy since v1.34 for low-correlation filters Payload index built ahead of time; filtering happens during the HNSW traversal itself rather than as a separate pass
Supported operators $eq, $ne, $gt, $gte, $lt, $lte, $in, $nin, $exists, plus $and/$or/$not for combining them Standard comparison and range operators, plus indexRangeFilters for efficient numeric/date range queries Match, Match Any, Match Except, Prefix Match, full-text match (all-terms or any-term), phrase match, numeric and datetime ranges, geospatial (bounding box, radius, polygon with interior exclusion rings), array length, is-null, is-empty, has-ID
Data types String, 64-bit integer, float, boolean, list of strings (no nested objects) Standard scalar types plus geo-coordinates Arbitrary JSON via dot-notation paths, including nested object and array element filtering
Notable limit 40 KB metadata per record; $in/$nin capped at 10,000 values per query Falls back to brute-force search when a filter is highly selective (roughly under 15% of the dataset), by design, to avoid exhaustive HNSW traversal None published as a hard cap; depth of nested filtering is effectively Qdrant's standout feature here

Quantization and Memory Tradeoffs

Quantization is how you keep RAM costs sane once your vector count climbs into the tens of millions. This is also one of the more technical differentiators, which matters most to the head-of-engineering reader deciding what a production cluster actually costs to run.

Pinecone Weaviate Qdrant
User-configurable quantization No, serverless indexing and compression are handled internally and aren't exposed as options Yes: Product Quantization, Binary Quantization, Scalar Quantization, and Rotational Quantization (8-bit, 4-bit, 1-bit variants) Yes: Scalar Quantization (4x compression), Binary Quantization (up to 32x), Product Quantization (up to 64x), and TurboQuant (up to 32x, with 4-bit, 2-bit, 1.5-bit, and 1-bit encoding)
Typical compression Not applicable, not user-facing PQ roughly 24x, BQ up to 32x (lossy, works best with certain embedding models), 8-bit RQ about 4x while holding 98 to 99% recall with no training phase required Scalar 4x with minimal accuracy loss; binary and product methods reach 32 to 64x for memory-constrained deployments
Configuration storage tiers Not applicable Not a named tier system; set per collection Three explicit strategies: RAM-optimized (fastest, highest memory), balanced hybrid (quantized vectors in RAM, originals on disk), and storage-optimized (both on disk, slowest)
Who this matters most to Teams who want zero tuning decisions Teams who want quantization but prefer it is bundled with query-time fusion in the same engine Teams running tens of millions of vectors who need to choose the exact memory-versus-recall point on the curve, including self-hosted

Multi-Tenancy and Namespace Models

If you're building a SaaS product on top of any of these, whether you're storing document chunks for retrieval or agent memory per customer, you need per-customer isolation without paying for a dedicated index per tenant. The three approaches aren't interchangeable.

Pinecone Weaviate Qdrant
Isolation unit Namespace within an index Dedicated shard per tenant Payload-based partition (default, for many small tenants), dedicated shard (for a few large tenants), or a tiered mix of both
Scale (published) 100 namespaces per serverless index on the Starter plan, up to 1,000,000 on Enterprise (Pinecone says it can go higher on request) Tested to roughly 170,000 active tenants across a 9-node cluster (about 18,000 to 19,000 per node); Weaviate states support for 1,000,000+ concurrent tenants with around 20 nodes Cloud clusters default to a 1,000 collection limit; payload partitioning avoids needing a collection per tenant in the first place
Query behavior Queries target one namespace at a time; metadata filters apply only within that namespace Each tenant effectively gets its own index, so queries never scan another tenant's data Payload-partitioned tenants share a collection but are filtered out at query time via an indexed is_tenant field, co-locating each tenant's vectors for fast sequential reads
Best fit Clean, simple per-customer split when you don't need tens of thousands of tenants Workloads needing hard per-tenant index isolation (compliance-sensitive multi-tenant SaaS) Very large numbers of small tenants where per-tenant dedicated infrastructure would be wasteful

Regional Availability

Pinecone Weaviate Qdrant
AWS us-east-1, us-west-2, eu-west-1, eu-central-1, ap-southeast-1 US East (N. Virginia) and Europe (Frankfurt) confirmed for Shared/Flex; broader AWS, GCP, and Azure coverage on Premium Dedicated, not publicly enumerated "Broad choice" of AWS regions on Standard clusters per Qdrant's own docs; exact list isn't published, visible when you create a cluster
GCP us-central1, europe-west4 Included in Premium's stated multi-cloud coverage, not itemized publicly Included in Standard's stated coverage, not itemized publicly
Azure eastus2 Included in Premium's stated multi-cloud coverage, not itemized publicly Included in Standard's stated coverage, not itemized publicly
Free/entry tier restriction Starter plan is locked to AWS us-east-1 only; other regions require a paid plan Free tier has limited region choice versus Flex and Premium Free tier has a limited choice of cloud provider and region versus Standard

Qdrant does not publish an exhaustive region list; what's shown here is from its own cloud documentation.

When Pinecone Is the Right Call

Pick Pinecone when the team's highest priority is shipping without anyone owning infrastructure. There's no cluster to size, no quantization strategy to pick, no self-hosted fallback to maintain, which is exactly the point for a small team that wants the database to be a line item, not a project. If cost or the missing self-hosted option is what's pushing you away, the Pinecone alternatives guide covers 12 replacements.

Pinecone
Best for Small teams with no dedicated infra engineer, fast RAG prototypes, products where query volume stays roughly proportional to a modest corpus
Real limitation No self-hosting at any price, and the read-unit model means cost can grow faster than either corpus size or query volume alone would suggest once both climb together
Not ideal for Large, high-traffic corpora where the read-unit multiplier turns into real budget risk, or any team that wants an exit path to its own infrastructure

When Weaviate Is the Right Call

Pick Weaviate when hybrid search is a core requirement, not a nice-to-have, and when you want the option to self-host the core engine without re-architecting if you later leave the managed cloud. The Weaviate alternatives guide covers what you'd give up by leaving it.

Weaviate
Best for Teams that need strong native BM25-plus-vector hybrid search out of the box, and teams who want a genuine self-hosted exit path for the core database
Real limitation The newer enterprise-grade platform features live behind a separate commercial license (the wl directory), so "open source" no longer means every current capability is free to self-host
Not ideal for Teams that need the deepest, most configurable quantization and multi-tenancy controls; Qdrant goes further on both

When Qdrant Is the Right Call

Pick Qdrant when you need to know, with certainty, that you can run the exact production database yourself if your vendor relationship or your budget changes, and when fine-grained control over the memory-versus-recall tradeoff matters. The Qdrant alternatives guide covers 11 other options for when the opaque cloud pricing is the sticking point.

Qdrant
Best for Technical teams with the capacity to self-host, workloads needing the deepest quantization and filtering options, very large multi-tenant SaaS products
Real limitation Cloud pricing isn't transparent without opening the calculator, so budgeting ahead of a sales conversation takes more work than with Pinecone or Weaviate
Not ideal for Small teams that explicitly do not want to think about infrastructure sizing at all, even with a managed cloud in front of it

Decision Framework

Pick this If this describes your situation
Pinecone You have no infrastructure engineer, you want zero-config scaling, and your corpus and traffic growth will likely stay roughly proportional
Pinecone You need the fastest possible path from zero to a production RAG endpoint and you're comfortable with full vendor lock-in
Weaviate Hybrid (keyword plus vector) search quality is a top-three requirement, not an afterthought
Weaviate You want to self-host the core database for free today, while still having a managed cloud on-ramp if the team grows
Qdrant You need a hard guarantee that you can self-host the exact production engine, under a fully permissive license, with no commercial carve-out
Qdrant You're running tens of millions of vectors and need granular quantization and multi-tenancy controls to keep memory costs under control
Run a two-week bake-off Your corpus size and query volume are both still unknown. Load your real embeddings into a free tier or self-hosted instance of each and measure actual RAM, actual latency, and actual filter performance before committing a budget line

Frequently Asked Questions about Pinecone, Weaviate, and Qdrant

Which of Pinecone, Weaviate, or Qdrant is cheapest?

It depends entirely on scale. At roughly 1 million vectors and 50,000 queries a month, Pinecone's Builder plan ($20/month flat) is the cheapest in our calculation. At 10 million vectors and 500,000 queries a month, our model shows Pinecone's metered cost jumping to roughly $285 to $319/month while Weaviate's stays near its $45/month floor, because Pinecone's read-unit pricing multiplies index size by query volume. Qdrant doesn't publish a flat cloud rate, but self-hosting under Apache 2.0 avoids licensing cost entirely at any scale.

Can I self-host Pinecone?

No. Pinecone is a proprietary, fully managed service with no self-hosted or on-premise option at any plan tier.

Is Weaviate still open source in 2026?

Most of it, yes. The core database remains BSD-3-Clause and is free to self-host. Since 2026, code in the repository's wl directory is described in Weaviate's own LICENSE file as copyright Weaviate BV, available only under a separate enterprise license, so some newer platform capabilities are no longer free to run yourself.

What license is Qdrant under, and does that cover self-hosting?

The entire Qdrant engine is Apache License 2.0, with no separate enterprise carve-out in the open-source repository. You can self-host the full production database for free; Qdrant Cloud, Hybrid Cloud, and Private Cloud are optional managed layers on top of that same code.

Which one has the best hybrid search?

Weaviate's BM25F-plus-vector fusion is the most mature native implementation of the three and was built into the engine from early on. Qdrant's Query API hybrid search, with Reciprocal Rank Fusion and Distribution-Based Score Fusion, is newer but adds multi-stage re-ranking that neither Pinecone nor Weaviate offers as a first-class primitive. Pinecone supports hybrid search too, but the single-index approach requires you to normalize unbounded sparse scores yourself.

Does metadata filtering slow down vector search on these databases?

Less than it used to. Weaviate pre-builds an eligible set before the vector search runs and falls back to brute-force search only when a filter is extremely selective. Qdrant builds payload indexes ahead of time and filters during the HNSW traversal itself rather than after. Pinecone applies metadata filters within the targeted namespace at query time. All three are built to avoid the older pattern of running a full vector search and then filtering the results afterward.

Do any of these support true multi-region, multi-cloud deployment?

Weaviate's Premium Dedicated tier and Qdrant's Standard tier both advertise coverage across AWS, GCP, and Azure, though neither publishes a complete, itemized region list; you confirm the exact list when you provision. Pinecone publishes its full region list directly: five AWS regions, two GCP regions, and one Azure region as of 2 October 2026, and locks its Starter plan to a single AWS region until you upgrade.

Has either Weaviate or Qdrant been acquired or pivoted away from being a vector database?

No. Qdrant closed a $50 million Series B in March 2026 to keep building its vector search infrastructure. Weaviate took a strategic investment from Ricoh's innovation fund in March 2026 aimed at combining Ricoh's data-capture technology with Weaviate's database. Pinecone remains an independently operated company with no acquisition on record. All three are still shipping as the same category of product their reputation describes.

What to Do Next

Don't pick a vector database off a feature table, including this one. Take your actual embedding dimension, vector count, and a realistic 12-month growth estimate for corpus size and query volume, then run the same math we did against each vendor's current calculator. If you're still deciding whether you need a dedicated vector database at all versus a simpler option, our guide to choosing AI knowledge base software is a useful gut check first. For the simplest alternative, staying on Postgres, pgvector vs Pinecone vs Weaviate tests whether that's enough.

Then spend a week with the free tier of whichever two are still on your shortlist. Load a representative sample of your real data, not a synthetic benchmark set, and measure what a spec sheet can't tell you: actual query latency under your real filter patterns, actual RAM consumption at your real dimension count, and how much engineering time cluster or index configuration eats in that week. That week of hands-on testing will tell you more than any comparison article, including this one. The vector database roundup prices the wider field, including pgvector, Milvus, and the engines built into Redis and Elasticsearch.

About the author

Camellia

Camellia

Principal Product Marketing Strategist

Camellia is Principal Product Marketing Strategist at Rework, helping B2B buyers pick the right software with confidence. With 6+ years in product marketing and 150+ SaaS tools evaluated across CRM, project management, and sales engagement, Camellia turns competitive intelligence into clear, honest comparisons. Readers get vendor evaluations they can trust to cut through marketing noise and decide faster.