
Vectorize · Hindsight: Open-Source Long-Term Memory for AI Agents
Hindsight is an open-source (MIT) long-term memory system for AI agents from Vectorize. Agents store what they see and do through a retain API, then pull it back with recall or have Hindsight reason over it with reflect, across sessions, users, and projects. It suits developers who want persistent, self-hostable agent memory that works with their existing framework, coding agent, or MCP client.
Best for
Developers building agents that must remember users, projects, or their own past work across sessions, who want a framework-agnostic memory service they can self-host or run as managed Cloud
Not ideal for
Simple stateless workflows or single-session chatbots, teams that cannot absorb an LLM call on every memory write, and anyone who only needs document retrieval over a fixed corpus
Who it's for
AI engineers and developers building conversational or autonomous agents, and teams using CLI coding agents that want persistent project memory
Hindsight treats memory as its own service rather than a feature of an agent framework, and that is the main reason to pick it. The same banks are reachable from Python, TypeScript, Go, MCP, or a coding agent, and the observation and mental model layers do work that simpler memory stores leave to you, such as merging duplicate facts and keeping a current summary of a user or project. The cost is an LLM call on every write, plus another stateful service to run and secure if you self-host. The benchmark numbers are strong but come from Vectorize and its research partners. It pays off for long-lived agents and for coding agents that should remember project decisions. Short-lived workflows probably do not need it.
Who should use it
Developers building assistants, support or sales agents, and autonomous agents that must carry knowledge about users and tasks across sessions, and teams using Claude Code, Codex, or Cursor that want project decisions from git history and past sessions available to every new session.
Who should skip it
Teams with stateless or single-session agents, projects that only need retrieval over a fixed document set, and anyone unwilling to pay for an LLM call on each memory write or to operate PostgreSQL and a memory server.
Self-hosted
Free
Hindsight Cloud
Pay as you go
Enterprise
Custom
Free tier limits: Hindsight Cloud starts with free credits, but the amount is not published. Self-hosting is free with no usage limits.
Note: Cloud runs on prepaid credits (1 credit = $1) with optional auto-recharge, and API calls return an error when the balance runs out. There is no fixed monthly fee or per-seat pricing. Costs scale with how much you retain and how often you reflect or refresh mental models, and knowledge pages refresh after consolidation by default, which adds refresh charges as you write. Volume discounts are available on request. Self-hosted deployments pay their own LLM provider for extraction and reflection, or nothing extra with a local model.
API pricing
Hindsight Cloud is usage-based on prepaid credits: per million tokens for retain, recall, Iris file extraction, and mental model retrieval, per call for reflect and mental model refresh, plus storage for memories older than 30 days
Available models
Per-user memory for a customer-facing assistant
One bank per user (or one bank with user tags) keeps each customer's history isolated, and a mental model can hold a current summary of their preferences that the agent reads without a retrieval step.
Project memory for coding agents
The coding agents package builds a per-repo bank from commit history and past sessions and injects it when Claude Code, Codex, or Cursor starts, so decisions recorded outside the code are available in later sessions.
Adding memory to an existing app with minimal changes
Wrapping an OpenAI or Anthropic client with `hindsight-litellm` recalls relevant memories before each call and retains the conversation afterward, without restructuring the application.
Hindsight vs. LlamaIndex
LlamaIndex is a framework for ingesting, indexing, and retrieving documents for RAG, with agent orchestration on top. Hindsight is a separate memory service that extracts facts from an agent's interactions and consolidates them over time. Hindsight ships a LlamaIndex integration (memory tools or a BaseMemory implementation), so the two are usually combined rather than chosen between.
Hindsight vs. LangGraph
LangGraph keeps thread-level state in checkpoints and offers its own long-term memory store shared across threads, but you decide what to write and how to search it. Hindsight handles extraction, consolidation, and reflection for you, and its LangGraph/LangChain integration includes a drop-in BaseStore adapter, so it typically plugs into LangGraph rather than replacing it.
Hindsight vs. Mastra
Mastra builds memory into its TypeScript agent framework, with message history, working memory, semantic recall, and observational memory. Hindsight is framework-agnostic, runs as its own server with Python, TypeScript, Go, REST, and MCP access, and adds graph and temporal retrieval plus a reflect API, at the cost of another service to run.
What is Hindsight?
Hindsight is an open-source agent memory system from Vectorize. Agents send it content with retain, fetch relevant memories with recall, and ask it to reason over what it knows with reflect. It organizes memories into world facts, experiences, observations, and mental models inside isolated memory banks.
Is Hindsight free and open source?
Yes. The code is on GitHub at vectorize-io/hindsight under the MIT license, and Vectorize says self-hosting has no usage limits and no telemetry. You still pay your own LLM provider for extraction and reflection unless you run a local model.
How much does Hindsight Cloud cost?
Hindsight Cloud is pay-as-you-go on prepaid credits, with no monthly or seat fee. Retain costs $10.00 per million input tokens, recall $0.75 per million output tokens, reflect $0.05 per call, mental model refresh $0.05 per call, and storage $0.25 per million tokens per month for memories older than 30 days. New accounts get free credits (amount not published), and Enterprise pricing is custom.
Can Hindsight be self-hosted?
Yes. You can run it with a single Docker command, install it with pip, or deploy it on Kubernetes with the Helm chart. Development can use the embedded PostgreSQL (pg0). Production needs PostgreSQL 14+ with pgvector or another supported vector extension, or Oracle AI Database. An Enterprise plan adds bring-your-own-cloud and on-premises options with support SLAs.
Does Hindsight need an LLM?
Yes. Retain uses an LLM to extract facts, entities, and relationships, and reflect uses one to reason over memories. Hindsight supports 25+ providers, including OpenAI, Anthropic, Gemini, Groq, Bedrock, and Vertex AI, local options such as Ollama, LM Studio, and llama.cpp, and any OpenAI-compatible endpoint. Embeddings and reranking run locally by default.
Does Hindsight work with Claude Code or Cursor?
Yes. The `@vectorize-io/hindsight-coding-agents` package installs hooks, plugins, or MCP tools for Claude Code, Codex CLI, Cursor CLI, GitHub Copilot CLI, opencode, and other CLI agents, and builds a per-repository memory bank from git history and past sessions. It defaults to Hindsight Cloud, with self-hosted and local daemon modes. Any MCP client can also connect to a bank's MCP endpoint directly.
How is Hindsight different from RAG?
RAG retrieves chunks from a document corpus. Hindsight extracts structured facts from what an agent sees and does, links them by entity and time, consolidates them into observations and mental models, and updates those as new evidence arrives. It is aimed at memory that changes over time rather than search over static documents.
Are Hindsight's benchmark results independent?
Not fully. Vectorize's benchmarks site lists 94.6% on LongMemEval S and 92% on LoCoMo10, and the December 2025 paper reports 91.4% on LongMemEval. The README says the results were reproduced by collaborators at Virginia Tech and The Washington Post, but researchers from both are co-authors of the paper.
LlamaIndex
Developers building RAG systems and document-grounded agents who need intelligent data parsing, indexing, and retrieval — especially teams with large or complex document sets
FreemiumLangChain
Developers building production multi-agent systems that need fine-grained control over state, execution flow, and human-in-the-loop checkpoints — and who are willing to trade setup time for that control
FreeMastra
TypeScript-first AI agents and workflows for Node.js teams
FreemiumHindsight runs as a separate memory server that agents call with three operations. Retain sends content such as conversations, tool output, or documents to an LLM that extracts facts, entities, relationships, and timestamps into an isolated memory bank. Recall combines semantic, keyword, graph, and time-based search, reranks the results, and trims them to a token budget. Reflect is an agentic loop that searches the bank and returns a grounded answer. In the background, Hindsight consolidates related facts into observations that are refined as evidence changes, and keeps mental models, which are standing answers to questions you define (such as a user's preferences) that an agent can read without an LLM call. That consolidation layer is what separates it from a plain vector store or from the memory built into most agent frameworks. You can self-host it with Docker, pip, or Helm on PostgreSQL with pgvector (or Oracle AI Database), or use Hindsight Cloud, which is pay-as-you-go on prepaid credits, billed per token and per call with no seat fee. It is reachable from Python, TypeScript, and Go clients, a CLI, REST, and a built-in MCP endpoint, and a single package adds per-repository memory to Claude Code, Codex, Cursor, and other CLI coding agents. The tradeoffs are cost and operations: every retain is an LLM call, reflect needs a model with solid tool calling, and self-hosting adds a stateful service to run and secure. Vectorize's benchmark lead on LongMemEval and LoCoMo is self-published, and the README itself says Hindsight may be overkill for simple n8n-style workflows.
Vectorize shipped Hindsight v0.10.2 with bank aliases for zero-downtime bank ID migrations, per-scope consolidation strategies, opt-in CUDA acceleration for ONNX embeddings, lower reflect input usage, and a fix restoring OpenAI provider compatibility with GPT-6 models.
Trendshift records the vectorize-io/hindsight repository first reaching #1 on GitHub Trending on September 26, 2026. The repository had about 43,000 stars as of September 30, 2026.
Are you the founder? Claim this listing →