Best RAG Tools in 2026: 13 Frameworks, Managed Platforms, and Search Engines Compared

Turn this article into takeaways for your work.

Each assistant summarizes the article only for you and suggests best practices for your work.

If your team needs to ground an LLM in your own documents, you're choosing between two purchases marketed as if they're one: a framework you assemble yourself (LangChain, LlamaIndex, Haystack, Dify), or a managed endpoint that runs retrieval for you (Vectara, Ragie, Azure AI Search, Amazon Bedrock Knowledge Bases, Google Vertex AI Search, Pinecone Assistant, Contextual AI). A third group, search infrastructure like Vespa and data-parsing tools like Unstructured, sits underneath both. This guide splits all 13 by what they actually are, prices each against its own vendor pricing page, and prices one real workload (50,000 pages, 10,000 queries a month) across the managed options so you can see where the ranking flips.

The pattern behind the retrieval-augmented generation category this year is a build-versus-buy swing, not a model-quality story. What changed is how much of the pipeline, chunking, embedding, reranking, vector storage, you're still expected to build yourself versus pay someone else to run.

Key Facts

  • Enterprise generative AI spending hit $37 billion in 2025, up 3.2 times from $11.5 billion in 2024, per Menlo Ventures' 2025 State of Generative AI in the Enterprise report.
  • 76% of enterprise AI use cases are now purchased rather than built in-house, up from 53% purchased about 18 months earlier, per the same Menlo Ventures report, the exact reversal that's reshaping whether a team reaches for a framework or a managed endpoint.
  • The vector database category grew 377% year over year, the fastest growth of any LLM-related technology Databricks tracked, per its State of AI: Enterprise Adoption & Growth Trends report (November 2025).
  • 57% of organizations building AI agents rely on base models plus prompt engineering and RAG rather than fine-tuning, per LangChain's State of Agent Engineering 2026 survey of 1,300-plus practitioners.
  • The global vector database market is projected to grow from $1.5 billion in 2023 to $4.3 billion by 2028, a 23.3% CAGR, per MarketsAndMarkets.

Quick Comparison Table

Tool Group Best For Starting Price Key Strength Key Limitation
LangChain Framework Widest integration ecosystem, agent-plus-RAG builds Free (MIT); LangSmith from $39/seat/mo Largest ecosystem and community You own the retrieval pipeline
LlamaIndex Framework Messy, document-heavy ingestion Free (MIT); LlamaCloud from $50/mo Best-in-class document parsing Parsing costs scale with volume
Haystack (deepset) Framework Explicit, debuggable pipelines, EU data residency Free (Apache 2.0); deepset Platform Enterprise is custom Inspectable pipeline-as-DAG design No self-serve mid-tier
Dify Framework / low-code Fast, non-engineer-built RAG chatbots Free self-hosted; Cloud from $590/yr Visual builder ships fast SaaS resale restricted by license
Vectara Managed (RaaS) Regulated enterprises buying a platform From $100,000/yr Built-in hallucination scoring No self-serve tier published
Ragie Managed (RaaS) Startups wanting a RAG endpoint today Free; Starter $100/mo Unlimited retrievals on every paid tier $250/mo per extra connector
Azure AI Search Managed (RaaS) Azure teams wanting flat, predictable cost Free; Basic ~$75/mo Capacity pricing, queries free in-tier You size and pay for capacity
Amazon Bedrock Knowledge Bases Managed (RaaS) AWS teams, lowest marginal cost per query Pay-as-you-go, no minimum Cheapest at real usage AWS-ecosystem lock-in
Google Vertex AI Search (Agent Search) Managed (RaaS) Google Cloud teams, structured + unstructured in one $1.50 / 1,000 queries (Standard) Grounded answers over structured and unstructured data Confusing post-rename docs
Pinecone Assistant Managed (RaaS) Teams already on Pinecone wanting a packaged layer $20/mo flat (Builder) Granular pay-for-use metering Five metered dimensions to forecast
Contextual AI Managed (component) Swapping in one best-in-class RAG component Pay-as-you-go, $25 free credit Transparent, published component pricing Multimodal parsing costs 13x text
Vespa Search infrastructure Hybrid lexical + vector search at scale Free self-hosted (Apache 2.0); Cloud from $0.05/vCPU-hr Native hybrid retrieval No managed RAG layer, you build generation
Unstructured Data infrastructure Parsing messy source documents before ingestion 10,000 free pages; then $0.015/page Every connector on every tier Not a retrieval tool by itself

Three Kinds of "RAG Tool," and Why Comparing Them on Price Is a Mistake

A reader shopping for "RAG tools" usually lands on a page ranking LangChain against Pinecone Assistant against Azure AI Search as if they compete for the same budget line. They don't.

Group What You're Actually Buying Who Runs the Infrastructure Who Maintains the Glue Code
Frameworks (LangChain, LlamaIndex, Haystack, Dify) Open-source code you assemble into a pipeline You (self-hosted) or a mix of self-hosted compute plus the vendor's cloud add-on You
Managed RAG-as-a-Service (Vectara, Ragie, Azure AI Search, Bedrock Knowledge Bases, Vertex AI Search, Pinecone Assistant, Contextual AI) A hosted endpoint: documents in, answers or ranked chunks out The vendor The vendor, mostly
Infrastructure components (Vespa, Unstructured) A piece of the pipeline (search engine, or document parsing), not a full RAG solution You (Vespa self-hosted) or the vendor (Vespa Cloud, Unstructured API) You, for everything above that piece

A framework license costing nothing and a managed endpoint costing $500 a month aren't comparable line items. The framework's real cost shows up in engineering hours: picking an embedding model, tuning chunking, operating a vector store, building an eval harness, absorbing version churn. The RAG assistant pattern looks identical from the outside in both cases. What differs is who's on the hook when retrieval quality degrades at 2am.

Framework licensing, verified

Before you assume any of these are simply "free," check the license. One of the four has a commercial restriction that catches teams building a product to resell.

Tool License Commercial SaaS Restriction Self-Host Cost
LangChain MIT None Free
LlamaIndex MIT None Free
Haystack (deepset) Apache 2.0 None Free
Dify Modified Apache 2.0 ("Dify Open Source License") Yes: cannot operate a multi-tenant SaaS on the Dify source without written permission from LangGenius Free, single-tenant use
Vespa Apache 2.0 None Free
Unstructured (open-source library) Apache 2.0 None Free

Group 1: Frameworks You Assemble Yourself

These four are free to license. None is free to run well. Each gives you a code layer for chunking, embedding, retrieving, and generating, plus an optional paid cloud product for observability, parsing, or hosting.

1. LangChain

LangChain is the default orchestration layer for LLM apps: chains, agents, retrievers, and memory, with the largest integration catalog on this list (hundreds of vector store, loader, and tool connectors). LangGraph, its stateful-agent layer, has become the common way teams build agentic RAG, where the model decides when and what to retrieve instead of following a fixed retrieve-then-generate sequence. LangChain vs LlamaIndex compares it directly with LlamaIndex, and the LangChain alternatives guide covers 12 frameworks you could use instead.

License MIT, no restrictions
Paid layer LangSmith (observability, tracing, evals)
Pricing Developer: free, 5,000 traces/mo, 1 seat. Plus: $39/seat/mo, 10,000 traces/mo, unlimited seats, plus usage billed in LangChain Standard Units at $1.00/LSU. Enterprise: custom, self-hosted/hybrid, SSO, SLA
Best for Teams wanting the widest integration surface who are comfortable assembling their own retrieval pipeline
Not ideal for Teams that want a single, opinionated path and no decisions to make

2. LlamaIndex

LlamaIndex is more retrieval-first than LangChain's broader agent focus, built specifically to connect LLMs to your own data. Its standout is LlamaParse, a document-parsing service that handles tables, scanned PDFs, and complex layouts noticeably better than a naive text extractor, which matters more than any retrieval algorithm when your source documents are the actual problem. The LlamaIndex alternatives guide separates replacing the framework from replacing LlamaCloud and LlamaParse.

License MIT, no restrictions
Paid layer LlamaCloud / LlamaParse
Pricing Free: 10,000 credits/mo, up to 100 users. Starter: $50/mo, 40,000 credits included, pay-as-you-go to 400,000. Pro: $500/mo, 400,000 credits plus a one-time 800,000-credit bonus, pay-as-you-go to $5,000/mo. Enterprise: custom, volume discounts, 5x rate limits, SSO. 1,000 credits = $1.25; basic page parsing costs as little as 1 credit, agentic/VLM parsing costs more
Best for Teams whose bottleneck is messy source documents, not agent orchestration
Not ideal for Teams needing the broadest non-RAG tool integration catalog

3. Haystack (deepset)

Haystack predates the current LLM wave; deepset originally built it for extractive question-answering and enterprise search, which shows in its architecture. Pipelines are explicit, inspectable directed graphs rather than implicit chains, easier to trace when a bad answer shows up. deepset is based in Berlin and carries a stronger EU public-sector footprint than the two US frameworks above, relevant if data residency matters to you.

License Apache 2.0, no restrictions
Paid layer deepset AI Platform (formerly deepset Cloud)
Pricing Studio: free, 1 workspace, 1 user, 100 pipeline hours/mo, 50 files (10MB max each), 2 development pipelines. Enterprise: Not published, contact sales; unlimited workspaces, users, and pipelines
Best for Teams that want an explicit, debuggable pipeline and may need EU data residency
Not ideal for Teams needing a self-serve tier between "free with 100 pipeline hours" and "enterprise sales call"

4. Dify

Dify is the odd one out here: a visual, low-code app builder with a RAG knowledge-base module built in, rather than a code-first framework. Non-engineers can assemble a working document-grounded chatbot on its canvas in an afternoon. Read the license before you plan to resell it, though: the self-hosted version bars running it as a multi-tenant SaaS without written permission from LangGenius, the company behind it.

License Modified Apache 2.0 ("Dify Open Source License")
Commercial restriction Cannot operate a multi-tenant SaaS on the open-source code without written permission from LangGenius
Pricing Sandbox: free, 200 message credits, 1 workspace, 1 member, 5 apps. Professional: $590/yr billed annually, 5,000 message credits/mo, 3 members, 50 apps, 5GB knowledge storage. Team: $1,590/yr billed annually, 10,000 message credits/mo, 50 members, 200 apps, 20GB storage. Enterprise: custom, commercial license, SSO
Best for Teams that want a working RAG assistant shipped by non-specialist builders
Not ideal for Teams needing to bolt on custom retrieval logic like hybrid search weighting

The page notes annual billing saves 17% over paying monthly; it does not publish a separate monthly rate, but that discount implies a monthly-billed Professional plan would run close to $59/month if it were offered.

Group 2: Managed RAG as a Service

Here you're buying an endpoint, not a codebase. Documents in, a query in, ranked chunks or a generated answer out. The pipeline decisions (chunking, embedding model, reranking) are the vendor's problem, exactly what the Menlo Ventures "76% purchased, not built" shift above describes.

5. Vectara

Vectara was one of the original RAG-as-a-service vendors, pairing its own "Boomerang" retrieval model with a "Mockingbird" generative model and hallucination-correction scoring built into the response. Its pricing page shows only annual enterprise contracts after a 30-day free trial. Older third-party reviews describe a free self-serve "Growth" tier; that no longer appears on Vectara's own page, so treat it as stale.

Pricing SaaS from $100,000/yr, VPC from $250,000/yr, on-premises from $500,000/yr, after a 30-day free trial. No self-serve metered tier published
Best for Regulated enterprises buying a company-wide platform
Not ideal for Teams evaluating a single RAG use case on a limited budget

6. Ragie

Ragie is the narrower, developer-first alternative: upload documents or connect a source, call a retrieval endpoint, get ranked chunks back. Every paid tier includes unlimited retrievals, hybrid search, reranking, entity extraction, and recency bias, none of it gated behind an Enterprise tier.

Pricing Developer: free, 1,000 max retrievals, 1,000 pages. Starter: $100/mo, unlimited retrievals, 10,000 pages. Pro: $500/mo, unlimited retrievals, 60,000 pages. Enterprise: custom, unlimited pages
Overage $0.02/page fast processing, $0.05/page hi-res, $0.002/page/mo storage, $250/mo per additional connector beyond the first (free) one
Best for Startups wanting a RAG endpoint today without operating a vector database
Not ideal for Teams pulling from several data sources simultaneously, where connector fees add up

Microsoft's managed search service is the retrieval layer underneath Azure AI Foundry RAG apps: full-text, vector, and hybrid search in one product, plus a semantic ranker add-on. It prices on provisioned capacity (search units and storage), not per query, the opposite metering model from most of this group.

Pricing Free: 50MB, 3 indexes. Basic: approximately $75/mo, 15GB storage, 15 indexes. Standard S1: approximately $245 to $250/mo, 160GB storage, 50 indexes. Standard S2/S3 and Storage Optimized L1/L2 scale further
Metering Capacity-based: query volume is free once a tier is provisioned. Agentic retrieval, semantic ranking, and document image extraction are separate add-ons, each with its own free monthly allowance
Best for Azure-committed teams wanting flat, predictable monthly cost regardless of query volume
Not ideal for Light, spiky usage, where you'd be paying for provisioned capacity you don't consume

8. Amazon Bedrock Knowledge Bases

Point it at an S3 bucket, pick an embedding model, and AWS handles chunking, embedding, indexing (into OpenSearch Serverless, Aurora, Pinecone, or another supported vector store), and retrieval. It's billed purely on storage and API calls, no provisioned tier to size, which makes it the cheapest hyperscaler option at real-world volume (see the workload pricing below).

Pricing Index storage: $5.00/GB of raw data per month. Standard Retrieve API: $1.00 per 1,000 calls. Agentic Retrieve (LLM-driven query planning): $4.00 per 1,000 calls plus the underlying Retrieve charge
Included at no extra charge Managed document parsing, embedding generation, and reranking, unless you bring a custom embedding or reranking model
Best for AWS teams wanting the lowest marginal cost per query
Not ideal for Teams wanting a GUI-first setup experience, or heavy agentic-retrieval usage where the 4x per-call premium adds up

Google's managed search and RAG product is now called Agent Search, part of the rebranded Gemini Enterprise Agent Platform (formerly Vertex AI Search). Standard Edition handles search without generation; Enterprise Edition adds grounded generative answers over both structured and unstructured data, the cleanest story of the three hyperscalers if your data spans BigQuery tables and documents alike.

Pricing Standard Edition: $1.50 per 1,000 queries. Enterprise Edition (includes core generative answers): $4.00 per 1,000 queries. Advanced Generative Answers add-on: an additional $4.00 per 1,000 queries on either edition. Indexed data storage: $5.00/GB/mo, first 10GB free
Best for Google Cloud teams wanting generative answers grounded in both structured and unstructured data
Not ideal for Anyone trying to quickly find current pricing; the Vertex AI Search to Agent Search rename has made the documentation harder to navigate than it was a year ago

(reported, vendor-sourced)

10. Pinecone Assistant

Pinecone Assistant is a higher-level product built on the standalone Pinecone vector database: upload documents, it handles chunking and embedding, and you call a chat-style endpoint instead of writing retrieval logic. Raw Pinecone remains available separately for teams building their own pipeline with a Group 1 framework.

Pricing Starter: free. Builder: $20/mo flat, 3GB storage, 2M input tokens, 1M output tokens, 2M context-processing tokens, 10,000 ingestion units included. Standard: $50/mo minimum usage, $300 trial credit included. Enterprise: $500/mo minimum usage, 99.95% uptime SLA
Overage beyond Builder Storage $3/GB/mo, input tokens $8/million, output tokens $15/million, context processing $5/million, ingestion $0.0005/unit
Best for Teams already using Pinecone who want a packaged assistant layer instead of building retrieval-to-generation glue
Not ideal for Teams wanting one simple metered dimension instead of five to forecast

11. Contextual AI

Founded by former Google DeepMind and Meta AI researchers, Contextual AI sells RAG as purpose-built components (parsing, reranking, grounded generation, an LMUnit evaluation model) rather than one black-box endpoint, alongside a packaged enterprise "RAG Agent" product. Unlike Vectara, it publishes self-serve, component-level pricing you can model before a sales call.

Pricing On-demand, pay-as-you-go, $25 in free credits to start, no stated minimum. Parse (text-only): $3 per 1,000 pages. Parse (multimodal): $40 per 1,000 pages. Rerank-v2: $0.05 per million tokens. Rerank-v2-mini: $0.02 per million tokens. Generate: $3 per million input tokens, $15 per million output tokens. LMUnit: $3 per million input tokens. Enterprise: custom
Best for Technical teams wanting to swap in one best-in-class RAG component, not replace their whole stack
Not ideal for Document sets with a meaningful share of scanned pages or embedded images; multimodal parsing costs more than 13 times text-only parsing

Group 3: Infrastructure You Still Need

Neither of these two is a RAG product you'd compare feature-for-feature against Group 2. They're pieces every tool above either includes internally or expects you to bring yourself.

12. Vespa

Vespa began inside Yahoo as a web-scale search engine, now maintained as its own company. It's not a RAG product out of the box; it's the vector database and search engine you'd build a RAG pipeline on once you've outgrown a managed vector store and need hybrid lexical plus vector retrieval, where exact matches on SKUs or error codes matter as much as semantic similarity.

License Apache 2.0, fully self-hostable for free
Vespa Cloud pricing Billed hourly per resource, four plans. Startup: $0.05/vCPU-hr, $0.005/GB memory-hr, $0.0002/GB disk-hr. Basic: $0.10/vCPU-hr, $0.01/GB memory-hr, $0.0004/GB disk-hr. Commercial: $0.145/vCPU-hr, $0.0145/GB memory-hr, $0.0005/GB disk-hr. Enterprise: $0.18/vCPU-hr, $0.018/GB memory-hr, $0.0007/GB disk-hr, $20,000/mo committed minimum
Discounts Annual commitments: 15% off (8% on enclave deployments inside your own cloud account)
Best for Engineering teams with search expertise who need hybrid retrieval at genuine scale
Not ideal for Teams wanting a managed RAG layer; Vespa stops at retrieval, you build the generation step

13. Unstructured

Unstructured isn't a retrieval tool. It solves the step before retrieval: parsing messy source files, PDFs, scanned documents, PowerPoint decks, HTML, images, into clean, chunked text any tool above can embed and index. Every framework and managed service here either runs something like it internally or expects an equivalent step before ingestion; data readiness is the unglamorous prerequisite most RAG evaluations skip. See our guide to data readiness for AI for the broader checklist. The AI data pipeline roundup prices it against LlamaParse, Reducto, Azure, and Textract on a 10,000-page job.

License The open-source unstructured Python library: Apache 2.0, no restrictions. The hosted API and commercial platform are separately licensed
Pricing 10,000 free pages to start. Pay-as-you-go: $0.015/page, all features included. Business: custom pricing, multi-user accounts, dedicated instance or VPC deployment
Best for Any team on this list whose source documents, not retrieval logic, are the actual bottleneck
Not ideal for A genuinely large one-time archive (hundreds of thousands of pages); the per-page rate adds up fast at that scale

Pricing One Real Workload: 50,000 Pages, 10,000 Queries a Month

Headline rates don't tell you what you'll actually pay. Here's the same workload, roughly 50,000 pages of source documents (product docs, support articles, contracts, call it 2GB once parsed to text), reindexed monthly, serving about 10,000 queries a month, priced across every managed option in Group 2. These are calculations built from each vendor's own 2026-10-02 rate card against the stated assumptions, not a vendor-quoted scenario price, so treat the estimated rows as directional.

Vendor Estimated Monthly Cost How It's Metered
Amazon Bedrock Knowledge Bases ~$20 $10 storage (2GB at $5/GB) + $10 for 10,000 standard Retrieve calls at $1/1,000
Google Vertex AI Search, Standard ~$15 $15 for 10,000 queries at $1.50/1,000; storage free under the 10GB allowance
Azure AI Search, Basic ~$75 Flat capacity tier; 15GB storage covers the corpus, queries free within capacity
Google Vertex AI Search, Enterprise (with core generative answers) ~$40 $40 for 10,000 queries at $4/1,000; add another $40/mo for the Advanced Generative Answers add-on if needed
Ragie, Pro $500 flat Flat tier; 60,000 pages included covers the 50,000-page corpus, retrievals are unlimited so query volume is free
Pinecone Assistant, Builder ~$20 to $40 (estimate) $20/mo flat base, likely ingestion and context-token overage on a full monthly 50,000-page reindex beyond the 10,000 included ingestion units
Contextual AI ~$150 one-time parse, then ~$100 to $150/mo $150 one-time for 50,000 pages at $3/1,000 text-only parse; ongoing Generate cost estimated assuming roughly 2,000 input and 300 output tokens per query across 10,000 queries/mo
Vectara $8,333/mo equivalent No metered tier; the $100,000/yr SaaS floor applies regardless of workload size

The ranking flips hard once you plug in a real workload. Bedrock and Vertex AI Search Standard stay cheap because they meter on storage and API calls, not a provisioned tier or a flat floor. Vectara's enterprise-only pricing turns a 50,000-page corpus into the same $100,000-a-year commitment as a 5-million-page one, which only makes sense once you're buying a company-wide platform.

Who Maintains the Glue Code? The Hidden Cost of the Framework Path

A free license is not a free decision. Pick LangChain, LlamaIndex, or Haystack, and you're also picking, and maintaining:

  • An embedding model. Which one, and what re-embedding your whole corpus costs when a better one ships. Our embedding models roundup prices eleven at a real workload.
  • A chunking strategy. Fixed-size, semantic, or document-structure-aware, re-tuned as your document mix changes.
  • A vector store. None of the three frameworks stores vectors itself; you separately pay for and operate Pinecone, a self-hosted Vespa, or another vector database on top of whichever one you pick. Our vector database roundup prices 12 at one workload.
  • Reranking and hybrid search. Built into most of Group 2 by default; a manual integration if you're assembling your own pipeline.
  • An eval harness. Something that tells you retrieval quality degraded before your users do. Our LLM observability roundup compares 12 platforms for tracing and evaluation.
  • Version churn. LangChain in particular has shipped breaking changes between major versions; that's the tradeoff for being the most actively developed framework here.

That list is what the 76%-purchased, 24%-built-internally shift from Menlo Ventures above describes at the macro level. None of it shows up on a framework's pricing page, which only covers what the vendor sells you (LangSmith traces, LlamaParse credits, deepset pipeline hours). The rest is headcount. Evaluating agent orchestration on top of your RAG pipeline too? Our multi-agent framework roundup and open-source AI agent framework roundup cover the same tradeoff one layer up the stack, and our guide to choosing AI knowledge base software gives a structured way to weigh a build against a purchase.

None of this argues against frameworks. A team with strong platform engineers and an unusual requirement (custom hybrid search weighting, a fine-tuned embeddings model, multi-hop retrieval per the agentic RAG pattern) will outgrow any managed endpoint eventually. "Free" and "fast to ship" just describe different tools, and conflating them is how a 10-person team ends up three months into a LangChain build that a $500-a-month Ragie subscription would have shipped in a week.

How to Choose: Decision Framework

If you need... Pick
The widest integration ecosystem, building agents that also retrieve LangChain
Your bottleneck is messy source documents (PDFs, scanned tables) LlamaIndex with LlamaParse, or Unstructured for parsing alone
An explicit, debuggable pipeline and EU data residency Haystack (deepset)
A working RAG chatbot built by non-engineers, fast Dify (check the SaaS-resale license clause first)
The lowest marginal cost per query, already on AWS Amazon Bedrock Knowledge Bases
Flat, predictable monthly cost regardless of query volume, already on Azure Azure AI Search
Generative answers grounded in structured and unstructured Google Cloud data Google Vertex AI Search (Agent Search)
A RAG endpoint today, no vector database to operate Ragie
You already use Pinecone and want a packaged assistant layer Pinecone Assistant
One best-in-class component (parsing or reranking), not a new stack Contextual AI
Hybrid lexical plus vector search at genuine scale, search engineers in-house Vespa
A regulated enterprise buying a full platform, budget isn't the constraint Vectara

Frequently Asked Questions

Frequently Asked Questions about RAG Tools

What's the real difference between a RAG framework and RAG-as-a-service?

A framework like LangChain, LlamaIndex, or Haystack is a code library you assemble yourself: you choose the vector store, embedding model, and chunking strategy, and maintain the glue code as everything changes. RAG-as-a-service options like Ragie, Vectara, and Bedrock Knowledge Bases are hosted endpoints: documents in, answers out, with the pipeline decisions made for you. Frameworks cost engineering time. Managed services cost a metered bill.

Is LangChain actually free?

The LangChain framework itself is open-source under the MIT license, free to use and self-host with no restrictions. What costs money is LangSmith, its observability and tracing product: free for small usage (5,000 traces/month, 1 seat), then $39/seat/month on the Plus tier plus usage-based charges.

Which RAG tool is cheapest for a small team?

For a modest corpus and query volume, Ragie's Starter plan ($100/month) or Amazon Bedrock Knowledge Bases' pay-as-you-go pricing (roughly $20/month for the 50,000-page workload above) are the least expensive managed options. With engineering time instead of budget, self-hosting LangChain, LlamaIndex, or Haystack costs nothing in licensing, though you'll still pay separately for a vector database and an LLM API.

Can I use Dify's open-source version to build a product I resell to multiple clients?

Only with explicit written permission from LangGenius, the company behind Dify. Its license, a modified Apache 2.0, specifically prohibits running the Dify source code as a multi-tenant SaaS, defined as one workspace per tenant with separated data, without that permission. Self-hosting for your own single-tenant use is unrestricted.

What happened to Vectara's free tier?

As of this article's pricing check on 2 October 2026, Vectara's own pricing page shows only annual enterprise contracts: SaaS from $100,000/year, VPC from $250,000/year, on-premises from $500,000/year, after a 30-day free trial. Older reviews describe a self-serve "Growth" plan with a free monthly query allowance that no longer appears on Vectara's current pricing page. Budget for an enterprise contract if you're evaluating Vectara today.

Do I still need a vector database if I use a framework like LangChain or LlamaIndex?

Yes. LangChain, LlamaIndex, and Haystack are orchestration layers, not storage. You still need to pick and pay for somewhere to store embeddings, whether that's Pinecone, a self-hosted Vespa or other vector database, or a managed option like Azure AI Search, and that cost sits on top of the framework itself.

What to Do Next

Don't start by comparing feature lists. Start with one question: does your team have the engineering time to own chunking, embedding, reranking, and an eval harness, or would you rather pay someone else to own it?

If you have the time and an unusual retrieval requirement, pick a framework from Group 1, budget real engineering weeks against it (not just the license cost), and plan for a vector database as a separate line item. If you want this running this month, price two or three Group 2 options against your actual document count and query volume using the worked example above as a template, not their headline rate. Either way, evaluate the parsing step (Unstructured or LlamaParse) on its own if your source documents are scanned, inconsistent, or table-heavy: a bad retrieval result is more often a bad chunk than a bad model.

About the author

Camellia

Camellia

Principal Product Marketing Strategist

Camellia is Principal Product Marketing Strategist at Rework, helping B2B buyers pick the right software with confidence. With 6+ years in product marketing and 150+ SaaS tools evaluated across CRM, project management, and sales engagement, Camellia turns competitive intelligence into clear, honest comparisons. Readers get vendor evaluations they can trust to cut through marketing noise and decide faster.