Best Agent Memory Tools in 2026: 8 Platforms for Production AI Agents
Turn this article into takeaways for your work.
Each assistant summarizes the article only for you and suggests best practices for your work.
If an agent needs to remember a user across sessions, the fastest path for most teams is Mem0 or Zep for a managed API, Letta if memory and reasoning need to run as one system, or Cognee if the whole thing has to be self-hosted and backend-agnostic. This guide evaluates eight purpose-built agent memory platforms plus the memory features that already ship free inside OpenAI's and Anthropic's own APIs.
Agent memory is a young category and the vendor list from six months ago is already wrong in places. Zep quietly stopped maintaining its open source core in 2026. Graphlit, one of the more ambitious platforms here, is winding down entirely, new signups closed and customer data scheduled for deletion. This guide leads with the question most vendors would rather you skip: do you need a dedicated memory product at all, or does a Postgres table and a native model feature already cover your case?
Key Facts
- The global vector database market, the infrastructure layer under most agent memory products, is projected to reach $3.2 billion in 2026, up from $2.58 billion in 2025, growing at roughly 24% a year through 2034 (Fortune Business Insights).
- A bigger context window is not memory: on the RULER long-context benchmark, GPT-4's accuracy fell from 96.6% at 4,000 tokens to 81.2% at 128,000 tokens, a 15.4-point drop even on a model advertising support for that full range (RULER, arXiv:2404.06654).
- Anthropic reports that pairing its memory tool with context editing lifted performance 39% over baseline on an internal agentic-search evaluation, versus 29% from context editing alone, a vendor-reported figure worth testing against your own workload (Anthropic).
- OpenAI's Responses API keeps response objects for 30 days by default and Conversation objects indefinitely, at standard token pricing with no separate storage fee, but it stores transcripts, not extracted memories (OpenAI).
- Redis has been tri-licensed since version 8.0 in May 2025: AGPLv3, the only OSI-approved option of the three, or one of two source-available licenses, a reversal of its 2024 move away from open source that also produced the Linux Foundation's Valkey fork (Redis).
Do You Actually Need a Memory Product?
Before you shortlist anything, ask whether you're solving a real problem or buying insurance against one you don't have yet.
A dedicated agent memory layer earns its cost when your agent needs genuine cross-session recall (not just a long transcript), when you need entity and relationship tracking that a plain similarity search won't give you, or when nobody on your team has the bandwidth to own embedding refresh, deduplication, and summarization logic by hand. That's a real, specific job, and it's different from retrieval-augmented generation over your own documents, which a plain vector database already handles well.
It's a different story pre-product-market-fit, or when your agent only needs to recall the current conversation. Every frontier model vendor now ships some persistence for free: OpenAI's Responses and Conversations APIs store transcripts, and Anthropic's context editing and memory tool manage a long session at standard token pricing. Neither does true memory consolidation, but neither does most of what a one-person team needs in month one. A Postgres table with a pgvector column and a weekly summarization job covers a surprising amount of ground, free, on infrastructure you already run. pgvector vs Pinecone vs Weaviate works out when that stops being enough.
The honest rule: start with what you already pay for. Move to a dedicated memory platform once you can name the specific gap it closes, not because the category is trending. For where memory fits inside the rest of your stack, see how AI agents use retrieval, choosing an AI agent platform, and how to choose AI knowledge base software.
Quick Comparison Table
| Tool | Best For | Starting Price | Open Source? | Key Limitation |
|---|---|---|---|---|
| Mem0 | A drop-in memory API with an unusually generous free tier | Free (10K adds/mo); $19/mo Starter | Yes, Apache 2.0 core | Pro tier jumps from $19 to $249/mo with no step in between |
| Zep | A temporal knowledge graph built in, not bolted on | Free (10K credits/mo); $125/mo Flex | Partial: Graphiti only (Apache 2.0); hosted core is closed | Community Edition discontinued in 2026, no free self-hosted path anymore |
| Letta (formerly MemGPT) | Memory and agent reasoning running as one self-hostable system | Free; $20/mo Pro plus per-agent and per-second metering | Yes, Apache 2.0 | Per-active-agent and per-tool-second pricing is hard to forecast at scale |
| Cognee | A self-hosted, backend-agnostic knowledge graph memory layer | Free (self-hosted); hosted Cognee Cloud from $1 per 1M tokens | Yes, Apache 2.0 | Hosted cloud tier is new and thin; production use leans on self-hosting |
| LangMem / LangGraph persistence | Teams already orchestrating agents in LangGraph | Free (MIT); cost is whatever you pay to host LangGraph | Yes, MIT | Not a standalone product; unusable outside LangGraph |
| Supermemory | Memory plus built-in connectors to Gmail, Notion, Drive, and OneDrive | Free ($5/mo usage included); $19/mo Pro | Core repo MIT, some SDKs Apache 2.0 | Four separate usage meters make bills hard to predict without a pilot |
| Redis | Teams that already run Redis and want to extend it into a memory store | Free (30MB); pay-as-you-go from $0.007/hr | Tri-licensed: AGPLv3 or source-available | Not purpose-built for memory. You write the retrieval and consolidation logic |
| Papr | A flat, operation-metered API, simpler than Mem0's tier curve | Free (1K ops/mo); $100/mo Starter | No, closed source | $100 Starter jumps straight to $500 Growth with nothing in between |
Licensing and Self-Hosting
This is the table to check before legal or security signs off on a self-hosted deployment. Licensing in this category has moved twice in 2026 alone.
| Tool | License | Self-Host Option | What Changed Recently |
|---|---|---|---|
| Mem0 | Apache 2.0 (core) | Yes, full open source server | No major change in 2026 |
| Zep | Apache 2.0, Graphiti component only | Graphiti yes; the core Zep memory service no | Zep discontinued Community Edition in 2026; the hosted core is now closed |
| Letta | Apache 2.0 | Yes, full framework and server | Rebranded from MemGPT in 2024; license unchanged since |
| Cognee | Apache 2.0 | Yes, multiple backend choices (Neo4j, Amazon Neptune, relational or vector stores) | None noted; stable Apache 2.0 throughout |
| LangMem / LangGraph | MIT | Yes, runs inside your own LangGraph deployment | None noted |
| Supermemory | MIT (core repo), Apache 2.0 (some SDKs) | Yes, documented as locally runnable | None noted |
| Redis | AGPLv3 or source-available (RSALv2 / SSPLv1), your choice | Yes, under whichever license you select | Reversed its 2024 SSPL move in May 2025 with Redis 8.0's tri-license |
| Papr | Closed source | No | None noted |
Pricing at Real Usage, Not the Headline Rate
Headline prices rarely survive contact with a real workload, because every vendor here meters a different unit: add events, "episodes," tokens processed, flat operations, or raw compute. To make them comparable, we modeled one scenario: a support agent handling roughly 10,000 conversations a month, writing about 300,000 memory events and running about 30,000 retrievals, assuming an average memory entry of 50 to 80 tokens (up to 350 bytes).
| Tool | Modeled Monthly Cost | How It's Metered |
|---|---|---|
| Mem0 | $249/mo (Pro) | Per add and retrieval request; 300K adds exceeds Starter's 50K cap, forcing the jump to Pro |
| Zep | About $562/mo on Flex Plus, cheaper than roughly $750/mo on the entry Flex tier | Per "episode" (up to 350 bytes) written; retrieval, storage, and users are unmetered |
| Supermemory | Roughly $19 to $39/mo (Pro) | Per token ingested ($5 to $10 per 1M) plus per-query search ($5 per 1M) |
| Papr | $500/mo (Growth), since Starter's 50K operation cap is too small | Flat rate per memory operation |
| Cognee Cloud | Roughly $15 to $30/mo plus $5 per extra workspace | Per 1M tokens processed |
| Letta | Roughly $20 to $50/mo platform fee, with underlying LLM token costs billed separately and uncapped | $20 base, $0.10 per active agent/mo, $0.00015 per tool-execution second |
| Redis | Roughly $20 to $60/mo | Hourly RAM and compute, not memory events; you build the retrieval logic |
| LangGraph | Not directly comparable: metered on infrastructure, not memory events | LangChain Standard Units at $1.00/unit against vCPU-hours and GiB-hours of compute and storage |
These are illustrative estimates from each vendor's own published rates and a stated workload assumption, not a quote: your actual mix of short facts versus long documents will move the numbers, so treat this as a method to re-run on your own traffic rather than a guaranteed bill. The flip worth noticing: Zep's $125/mo sticker looks cheapest at a glance, but its per-episode credit metering means that at 300,000 memory events a month, the pricier $375/mo Flex Plus tier wins, because its larger credit bucket and cheaper overage rate beat 25 top-up blocks on Flex. Papr shows the opposite problem: nothing sits between $100 and $500, so a workload barely above Starter's cap still pays for Growth's full 750,000-operation ceiling.
1. Mem0: The Default Starting Point
Mem0 is the closest thing this category has to a default choice, and the free tier is the reason why: 10,000 memory writes and 1,000 retrievals a month, no card required. The Starter tier at $19/mo multiplies both ceilings by five. The jump to Pro is steep ($249/mo for 500,000 writes, 50,000 retrievals, and graph memory for entity tracking), and there's no middle option, so budget for it rather than being surprised.
The product itself is a straightforward memory API: SDKs across Python and Node, an Apache 2.0 open source server you can self-host for free, and the same interface works against OpenAI, Anthropic, or a local model. That portability, decoupling memory from your choice of LLM, is Mem0's real strength for a team that hasn't committed to one model vendor long-term.
| Dimension | Detail |
|---|---|
| Best for | Fast setup, model-agnostic memory, generous free tier for prototyping |
| Not ideal for | Teams that will outgrow Starter and need predictable costs mid-tier |
| Self-host | Yes, Apache 2.0 |
| Entity/relationship tracking | Yes, gated behind the Pro tier ($249/mo) |
2. Zep: Built on a Temporal Knowledge Graph
Zep's differentiator is Graphiti, its open source temporal knowledge graph framework: it tracks not just what a user said, but when a fact became true and when it stopped being true, which matters for agents that need to reason about state changes ("they used to be on the free plan, now they're on Enterprise") rather than just retrieve similar text.
The catch is Zep's open source story changed in 2026: the company discontinued Zep Community Edition, the self-hosted version of its core memory service, and now concentrates its open source effort entirely on Graphiti. If you want to self-host the full memory service rather than just the graph library underneath it, that option is gone. What's left is a credit-metered hosted product: free up to 10,000 credits a month, $125/mo for 50,000 (Flex), $375/mo for 200,000 (Flex Plus), with storage, retrieval, and users unmetered, only writing an "episode" costs a credit.
| Dimension | Detail |
|---|---|
| Best for | Temporal/relationship reasoning, teams that want a hosted graph without building one |
| Not ideal for | Teams that need to self-host the full memory service, not just Graphiti |
| Self-host | Graphiti only (Apache 2.0); core service is hosted-only now |
| Pricing model | Credits per write episode; retrieval and storage are free |
3. Letta (formerly MemGPT): Memory and Reasoning as One System
Letta is the commercial continuation of MemGPT, the 2023 Berkeley research project that modeled agent memory after an operating system: a main context window acts like RAM, archival memory acts like disk, and the agent decides what to page in and out through tool calls. The rebrand happened in 2024, the Apache 2.0 license hasn't changed since, and the GitHub repository confirms the full framework and server remain open source and self-hostable.
What sets Letta apart from Mem0 or Zep is architectural: memory isn't a separate API call, it's part of the same stateful agent, sharing context with reasoning instead of syncing between two systems. Pricing reflects that integration: a $20/mo base plan, then $0.10 per active agent/mo and $0.00015 per second of tool execution, with LLM token costs passed through separately. That's granular and genuinely hard to forecast until you've run a full billing cycle, the honest tradeoff for a platform-level product instead of a simple metered API.
| Dimension | Detail |
|---|---|
| Best for | Teams that want memory and reasoning as one self-hostable system |
| Not ideal for | Teams that want one predictable line-item cost instead of multi-axis metering |
| Self-host | Yes, Apache 2.0, full framework and server |
| Pricing model | Base fee plus per-active-agent and per-tool-second usage, LLM costs separate |
4. Cognee: The Self-Hosted, Backend-Agnostic Option
Cognee is built for teams who want the memory layer itself to be infrastructure they own, not a vendor they depend on. It's Apache 2.0, it ingests data in essentially any format, and it builds a structured knowledge graph your agents can query and update across sessions, running against whichever backend you already operate: Neo4j, Amazon Neptune, or a standard relational or vector store.
The hosted option, Cognee Cloud, is new and intentionally thin: a free tier (1 million tokens, one workspace), a Standard tier at $1 per 1 million tokens processed plus $5 per additional workspace, and an Enterprise bring-your-own-cloud tier with vendor support. For most teams, the honest expectation is that production use means self-hosting and owning the operational overhead, with the cloud tier more a way to try it than a mature managed product yet.
| Dimension | Detail |
|---|---|
| Best for | Fully self-hosted, backend-agnostic memory with no vendor lock-in |
| Not ideal for | Teams that want a mature managed service rather than infrastructure to operate |
| Self-host | Yes, Apache 2.0, multiple backend choices |
| Pricing model | Self-host free; hosted Cognee Cloud at $1 per 1M tokens processed |
5. LangMem and LangGraph Persistence: Memory Inside Your Orchestrator
If you're already building on LangGraph, LangMem isn't a separate purchase decision, it's a library (MIT, free) that stores memories inside LangGraph's own structured store. There's no standalone LangMem pricing because there's no standalone LangMem service: your cost is whatever you pay to run LangGraph, self-hosted (free, MIT) or through LangGraph Platform, where the Plus tier runs $39 per seat a month with 10,000 base traces included, plus usage billed in LangChain Standard Units at $1.00 each against compute and storage.
The tradeoff is specialization: this only makes sense if LangGraph is already your orchestration layer, not a product you'd adopt independently. And because LangGraph Platform meters infrastructure (vCPU-hours, GiB-hours) rather than memory events, you can't isolate "how much memory costs" the way Mem0 or Zep's per-write pricing lets you. That's fine if you reason about agent cost as a whole; it's a real limitation if you need to budget memory as its own line item. For frameworks more broadly, see our roundup of AI agent frameworks for developers.
| Dimension | Detail |
|---|---|
| Best for | Teams already standardized on LangGraph for agent orchestration |
| Not ideal for | Anyone not using LangGraph, or anyone who needs to budget memory cost separately |
| Self-host | Yes, MIT, free |
| Pricing model | No separate memory pricing; bundled into LangGraph Platform compute/storage billing |
6. Supermemory: Memory With Connectors Built In
Supermemory's pitch is that memory shouldn't require manual feeding: the Pro tier ($19/mo, roughly $20 of usage included) ships Google Drive, Notion, and OneDrive connectors, and Max ($100/mo) adds Gmail, so memory can ingest from where work already happens instead of waiting for your app to push data in. The core repository is MIT licensed, with some SDKs under Apache 2.0, and the vendor documents it as runnable fully locally if you want to self-host rather than pay.
Where it gets complicated is billing: four separate usage meters (memory ingestion by token, SuperRAG retrieval by token, search by query, and a catch-all "operations" meter at $100 per million), each with its own rate and its own plain-text versus rich-content split. That's flexible, but you can't eyeball a monthly bill from the plan price alone; you need a week or two of real traffic before the number stabilizes.
| Dimension | Detail |
|---|---|
| Best for | Teams that want memory ingestion from existing productivity tools, not just API calls |
| Not ideal for | Teams that want one simple, predictable price |
| Self-host | Documented as runnable locally; core repo MIT |
| Pricing model | Four separate usage meters: memory, SuperRAG, search, operations |
7. Redis: Your Existing Cache, Stretched Into a Memory Store
Redis isn't a purpose-built agent memory product, and it doesn't pretend to be one on its pricing page. It's positioned as infrastructure: caching, a vector database and semantic search, session management, and, explicitly, a memory store for AI agents and RAG applications, all on the same engine most engineering teams already run for something else. That's the appeal: if Redis is already in your stack, extending it into a fast, low-latency memory layer avoids adding a new vendor, a new bill, and a new set of credentials.
It's also the limitation. Redis gives you the storage primitive, not the memory logic: you write your own consolidation, deduplication, decay, and entity-extraction code, none of which ships the way Mem0's graph memory or Cognee's knowledge graph does. Before committing to self-hosting it commercially, get legal sign-off on licensing: since Redis 8.0 in May 2025 it ships tri-licensed under AGPLv3, RSALv2, or SSPLv1, after a 2024 detour into source-available-only licensing that pushed the Linux Foundation to fork the last BSD version as Valkey. AGPLv3 is the only one of the three that's OSI-approved open source; the other two carry commercial-use restrictions worth reading closely.
| Dimension | Detail |
|---|---|
| Best for | Teams already running Redis who want to avoid adding a new vendor for memory |
| Not ideal for | Teams that want memory logic (consolidation, entity tracking) included, not built |
| Self-host | Yes, under your choice of AGPLv3, RSALv2, or SSPLv1 |
| Pricing model | Hourly RAM/compute, free up to 30MB, pay-as-you-go from $0.007/hr |
8. Papr: The Simple Metered API
Papr's pitch is simplicity: one metered unit (memory operations), a free developer tier (1,000 operations, 1GB, 2,500 active memories), and two paid tiers, Starter at $100/mo (50,000 operations, 10GB, 100,000 active memories) and Growth at $500/mo (750,000 operations, 100GB, 1 million active memories). There's no graph-memory upsell and no multi-axis credit system to model, genuinely easier to reason about than Mem0's or Supermemory's pricing pages.
The tradeoff shows up at the edges. Papr is closed source, so there's no self-hosted fallback, and the gap between tiers is large: a team needing 80,000 operations a month has no $150 or $250 option and must buy the full $500 Growth tier. Papr also carries the smallest public footprint here, with fewer third-party integrations and less independent benchmarking than Mem0, Zep, or Letta, which matters if you want evidence beyond the vendor's own claims.
| Dimension | Detail |
|---|---|
| Best for | Teams that want the simplest possible pricing model to reason about |
| Not ideal for | Teams whose usage will land between tiers, or who need a self-hosted fallback |
| Self-host | No, closed source |
| Pricing model | Flat rate per memory operation across three tiers |
The One That Died: Graphlit
Graphlit is the cautionary example this category needed. It was pitched as a "context layer for AI agents," semantic infrastructure with ingestion, knowledge graphs, and RAG chatbots in one platform, and its own pricing page, when we fetched it, still listed four live tiers. But a second page on the same domain tells the real story: Graphlit's own sunset notice confirms the service is winding down, is no longer accepting new customers, and has scheduled customer accounts and project data for deletion once each account's offboarding window closes.
| Status | Detail |
|---|---|
| New signups | Closed; not accepting new customers |
| Free customer export deadline | August 1, 2026 (already passed) |
| Paid customer timeline | Individual, set by direct notice to each account |
| Data | Scheduled for deletion after each account's offboarding window closes |
| Export scope | Content metadata, extracted text and transcripts export; embeddings and vector indexes do not, because they're model- and index-specific |
The lesson generalizes beyond Graphlit: a comprehensive-sounding platform in this category can close its doors the same year it's being compared favorably to Zep and Letta elsewhere. Check a vendor's own legal or status page before you build on it, not just its marketing homepage, which in Graphlit's case still reads as if nothing has changed (Graphlit sunset notice).
What OpenAI and Anthropic Ship for Free
If you're comparing this against an older guide that mentions "OpenAI Assistants," that product is gone: OpenAI deprecated the Assistants API on August 26, 2025 and fully removed it a year later, on August 26, 2026, a date already past as of this writing (OpenAI). What replaced it, the Responses plus Conversations APIs, is what matters now, alongside Anthropic's separate context tooling for Claude.
| Option | What It Actually Does | Cost | What It's Missing |
|---|---|---|---|
| OpenAI Responses + Conversations API | Stores response and conversation history; Responses retained 30 days by default, Conversation objects indefinitely | Included in standard token pricing, no storage fee | No summarization or entity extraction. It remembers the transcript, not the meaning |
| Anthropic context editing + memory tool | Server-side clearing of stale tool results and thinking blocks, plus a file-based memory tool Claude can write to and read from across a session | Included in standard token pricing | You design what gets written to memory; there's no managed retrieval or consolidation layer |
| Plain Postgres (with pgvector) | A table you already run, extended with an embedding column and a scheduled job that summarizes old rows | Free, beyond what you already pay for the database | You build chunking, embedding refresh, decay, deduplication, and retrieval ranking yourself |
For a lot of teams under real production load, this trio genuinely is enough, especially pre-product-market-fit, when you don't yet know what "memory" should mean for your specific agent. For a side-by-side of the underlying models these features ship on, see Claude vs ChatGPT vs Gemini. Reach for one of the eight platforms above once you can name a gap none of these three options closes.
How to Choose: Decision Framework
| If you need... | Pick... | Why |
|---|---|---|
| A fast drop-in with the most generous free tier | Mem0 | 10,000 free writes a month and a $19 entry tier before any real commitment |
| Built-in temporal or relationship reasoning | Zep or Cognee | Graphiti and Cognee's knowledge graph both model entities and relationships, not just text similarity |
| Memory and agent logic in one self-hosted system | Letta | Memory is the agent's own context management, not a separate API call |
| Full self-hosting with backend flexibility | Cognee | Apache 2.0, runs against Neo4j, Neptune, or a relational/vector store you already operate |
| Memory already wired into your orchestrator | LangMem / LangGraph | No new vendor if you're already running LangGraph for agent orchestration |
| Memory plus ingestion from Gmail, Notion, Drive, or OneDrive | Supermemory | The only platform here with first-party productivity-app connectors |
| Lowest latency and you already run the infrastructure | Redis | Fastest path if your team operates Redis already; get licensing sign-off first |
| Simple pricing over Mem0's tier curve | Papr | Flat per-operation pricing, no graph-memory upsell |
| You're pre-PMF or your agent's job is genuinely simple | None of the above | Start with your model vendor's native memory plus a Postgres table; revisit once a specific gap appears |
Team Size and Who Runs It
| Tool | Fits Best At | Who Operates It |
|---|---|---|
| Mem0 | Solo developer through mid-size engineering team | One backend engineer; no dedicated infra team needed |
| Zep | Small to mid-size team that wants a managed graph | API consumer only; no infra ownership required |
| Letta | Teams building genuinely agentic, stateful products | Backend/ML engineer comfortable with a platform, not just an API |
| Cognee | Engineering teams with existing database operations capacity | Someone who already owns Neo4j, Neptune, or a vector store |
| LangMem / LangGraph | Any team already standardized on LangGraph | Whoever already runs your LangGraph deployment |
| Supermemory | Small teams that want ingestion handled, not built | One developer; connectors reduce the integration work |
| Redis | Teams with existing Redis/infrastructure operations | Platform or infrastructure engineer, plus legal review of licensing |
| Papr | Small teams with simple, bounded memory needs | One developer; minimal ops overhead |
What to Do Next
Don't pick a platform before you've modeled your own numbers. Take the metric that matters for your agent (writes per month, retrievals per month, and whether you need relationship tracking or just similarity search) and run it through two candidates: one managed option (Mem0 or Zep) and one self-hosted option (Letta or Cognee), using the method above rather than the headline rate. Spend a week on real traffic before committing to an annual plan, check a vendor's status or legal page the same day you check its pricing page, and confirm you need a dedicated memory layer before adding a ninth vendor to your stack. A Postgres table plus your model provider's native memory feature is the right answer too, when it's enough. Once memory is live, tracing the retrieval step behind a wrong answer is an observability job, and the LLM observability roundup compares 12 platforms for it.

On this page
- Key Facts
- Do You Actually Need a Memory Product?
- Quick Comparison Table
- Licensing and Self-Hosting
- Pricing at Real Usage, Not the Headline Rate
- 1. Mem0: The Default Starting Point
- 2. Zep: Built on a Temporal Knowledge Graph
- 3. Letta (formerly MemGPT): Memory and Reasoning as One System
- 4. Cognee: The Self-Hosted, Backend-Agnostic Option
- 5. LangMem and LangGraph Persistence: Memory Inside Your Orchestrator
- 6. Supermemory: Memory With Connectors Built In
- 7. Redis: Your Existing Cache, Stretched Into a Memory Store
- 8. Papr: The Simple Metered API
- The One That Died: Graphlit
- What OpenAI and Anthropic Ship for Free
- How to Choose: Decision Framework
- Team Size and Who Runs It
- What to Do Next