LaunchedEditorial Listing

Hindsight

Vectorize · Hindsight: Open-Source Long-Term Memory for AI Agents

Open Hindsight

Hindsight is an open-source (MIT) long-term memory system for AI agents from Vectorize. Agents store what they see and do through a retain API, then pull it back with recall or have Hindsight reason over it with reflect, across sessions, users, and projects. It suits developers who want persistent, self-hostable agent memory that works with their existing framework, coding agent, or MCP client.

PricingFreemium
Setupmedium
Runs onAPI · Self-hosted
APIYes
Open sourceYes
DocsYes
Agent MemoryLong-Term MemoryOpen SourceSelf-HostedMCPKnowledge GraphPython SDKTypeScript SDK

Best for

Developers building agents that must remember users, projects, or their own past work across sessions, who want a framework-agnostic memory service they can self-host or run as managed Cloud

Not ideal for

Simple stateless workflows or single-session chatbots, teams that cannot absorb an LLM call on every memory write, and anyone who only needs document retrieval over a fixed corpus

Who it's for

AI engineers and developers building conversational or autonomous agents, and teams using CLI coding agents that want persistent project memory

Capabilities

  • Retain, recall, and reflect APIs: store content, retrieve relevant memories within a token budget, or get an LLM-generated answer grounded in the bank
  • Four memory types: world facts, the agent's own experiences, consolidated observations, and mental models
  • Recall runs semantic, BM25 keyword, graph (entity, temporal, and causal links), and time-range retrieval in parallel, then merges them with reciprocal rank fusion and a cross-encoder reranker
  • Background consolidation turns related facts into observations that keep exact supporting quotes and a proof count, and are refined rather than overwritten when new evidence arrives
  • Mental models and knowledge pages: standing answers and wiki-style markdown pages that Hindsight rewrites as the bank learns, read back without a retrieval or LLM call
  • Isolated memory banks per user, agent, or project, with tags for per-user filtering inside a shared bank, disposition traits (skepticism, literalism, empathy) that shape reflect, and declarative bank templates
  • Built-in MCP server on every instance, with a per-bank endpoint exposing retain, recall, reflect, mental model, document, and bank tools
  • `hindsight-litellm` wrapper for OpenAI and Anthropic clients that recalls memories before each call and retains the conversation after it
  • Coding agents package that wires hooks, plugins, or MCP tools into Claude Code, Codex CLI, Cursor CLI, GitHub Copilot CLI, opencode, Cline CLI, and other CLI agents, building a per-repo bank from git history and past sessions
  • Clients for Python, TypeScript/Node.js, and Go, plus a CLI, REST API, and an embedded Python mode that starts the server in-process
  • 25+ LLM providers, including OpenAI, Anthropic, Gemini, Groq, Bedrock, Vertex AI, DeepSeek, and local Ollama, LM Studio, or llama.cpp, plus ChatGPT, Claude, Cursor, or GitHub Copilot subscriptions in place of an API key
  • Opt-in, per-bank Memory Defense that scans each retain against 45 secret and PII patterns and redacts or blocks matches
  • Production tooling: Prometheus metrics, an admin CLI for migrations and stuck operations, lifecycle webhooks, and tenant, auth, and storage extension points

Limitations

  • Every retain runs an LLM to extract facts and entities, so ingestion costs tokens and time. Models outside the tested list need at least 65,000 output tokens for reliable extraction, and there is a reduced-limit setting for models with less
  • Reflect is a tool-calling loop, and the docs call it the operation that leans hardest on tools, so models with weak or emulated tool calling are less reliable. The FAQ puts reflect at 1 to 10 seconds versus 50 to 500 ms for recall
  • Self-hosting adds a stateful service. Production needs PostgreSQL 14+ with a vector extension (the embedded pg0 database is for development). The installation guide lists 1.5 GB minimum (2 GB recommended) for the full API image plus 512 MB to 1 GB or more for PostgreSQL, the FAQ suggests 8 GB for production, and a GPU or external reranker is recommended to keep recall fast under load
  • The MCP endpoint is open with no authentication by default. You have to enable the API key extension before exposing a server
  • The LLM wrapper and the coding agents installer both default to Hindsight Cloud, so self-hosters have to point them at their own server or a local daemon
  • Benchmark results are published by Vectorize. The README says they were independently reproduced by collaborators at Virginia Tech and The Washington Post, but researchers from both are co-authors of the Hindsight paper
  • Still pre-1.0 (v0.10.x) with releases every one to three weeks, so expect frequent upgrades and occasional migrations
  • Cloud costs grow with write volume. The free starting credit amount is not published, and knowledge pages refresh after consolidation by default, which the billing docs warn can outgrow all other operations combined

Use cases

  • Giving a support or sales chatbot per-user memory by creating one bank per user, or one bank with user tags
  • Adding memory to an existing OpenAI or Anthropic SDK app by wrapping the client with `hindsight-litellm`
  • Giving Claude Code, Codex, or Cursor per-repository memory of past decisions drawn from git history and earlier sessions
  • Connecting an MCP-capable assistant to a bank's MCP endpoint so it can retain and recall across sessions
  • Defining a mental model such as "What are this user's preferences?" that an agent reads at startup instead of rediscovering context
  • Running reflect over accumulated experience, for example to ask which outreach messages got responses and why

Our take

Hindsight treats memory as its own service rather than a feature of an agent framework, and that is the main reason to pick it. The same banks are reachable from Python, TypeScript, Go, MCP, or a coding agent, and the observation and mental model layers do work that simpler memory stores leave to you, such as merging duplicate facts and keeping a current summary of a user or project. The cost is an LLM call on every write, plus another stateful service to run and secure if you self-host. The benchmark numbers are strong but come from Vectorize and its research partners. It pays off for long-lived agents and for coding agents that should remember project decisions. Short-lived workflows probably do not need it.

Who should use it

Developers building assistants, support or sales agents, and autonomous agents that must carry knowledge about users and tasks across sessions, and teams using Claude Code, Codex, or Cursor that want project decisions from git history and past sessions available to every new session.

Who should skip it

Teams with stateless or single-session agents, projects that only need retrieval over a fixed document set, and anyone unwilling to pay for an LLM call on each memory write or to operate PostgreSQL and a memory server.

Strengths

  • MIT-licensed and free to self-host, with a managed Cloud option and no seat fees
  • Combines vector, keyword, graph, and temporal retrieval instead of vector search alone
  • Consolidated observations and mental models give agents settled knowledge, not just raw snippets
  • Framework-agnostic: SDKs, REST, MCP, an LLM client wrapper, and 60+ listed integrations
  • Can run LLM, embeddings, and reranking on local models if data must stay on your hardware

Weaknesses

  • An LLM call on every retain adds cost and latency
  • Another stateful service to deploy, secure, and upgrade when self-hosted
  • MCP endpoint has no authentication unless you enable it
  • Benchmark leadership claims come from Vectorize and its paper co-authors

Hindsight pricing

Self-hosted

Free

  • MIT license, no usage limits
  • Docker, pip, or Helm deployment with embedded or external PostgreSQL
  • Includes MCP server and REST API
  • Community support via GitHub

Hindsight Cloud

Pay as you go

  • Retain $10.00 per 1M input tokens, recall $0.75 per 1M output tokens, reflect $0.05 per call
  • Iris file extraction $7.50 per 1M tokens, mental model retrieve $0.25 per 1M tokens, refresh $0.05 per call
  • Storage $0.25 per 1M tokens per month, free for each memory's first 30 days
  • 99.9% uptime SLA, 12x5 email support with a 12-hour response SLA

Enterprise

Custom

  • Bring-your-own-cloud or on-premises deployment
  • SSO and RBAC, dedicated infrastructure
  • Custom SLA up to 99.95%, up to 24x7 support with a 30-minute response SLA

Free tier limits: Hindsight Cloud starts with free credits, but the amount is not published. Self-hosting is free with no usage limits.

Note: Cloud runs on prepaid credits (1 credit = $1) with optional auto-recharge, and API calls return an error when the balance runs out. There is no fixed monthly fee or per-seat pricing. Costs scale with how much you retain and how often you reflect or refresh mental models, and knowledge pages refresh after consolidation by default, which adds refresh charges as you write. Volume discounts are available on request. Self-hosted deployments pay their own LLM provider for extraction and reflection, or nothing extra with a local model.

Technical specs

API pricing

Hindsight Cloud is usage-based on prepaid credits: per million tokens for retain, recall, Iris file extraction, and mental model retrieval, per call for reflect and mental model refresh, plus storage for memories older than 30 days

Available models

Any of 25+ LLM providers for retain and reflect (OpenAI, Anthropic, Gemini, Groq, Bedrock, Vertex AI, DeepSeek, Ollama, LM Studio, llama.cpp, OpenAI-compatible endpoints, LiteLLM)Default embeddings: BAAI/bge-small-en-v1.5 (runs locally)Default reranker: cross-encoder/ms-marco-MiniLM-L-6-v2 (runs locally)

Where Hindsight excels

Per-user memory for a customer-facing assistant

One bank per user (or one bank with user tags) keeps each customer's history isolated, and a mental model can hold a current summary of their preferences that the agent reads without a retrieval step.

Project memory for coding agents

The coding agents package builds a per-repo bank from commit history and past sessions and injects it when Claude Code, Codex, or Cursor starts, so decisions recorded outside the code are available in later sessions.

Adding memory to an existing app with minimal changes

Wrapping an OpenAI or Anthropic client with `hindsight-litellm` recalls relevant memories before each call and retains the conversation afterward, without restructuring the application.

Hindsight vs. competitors

Hindsight vs. LlamaIndex

LlamaIndex is a framework for ingesting, indexing, and retrieving documents for RAG, with agent orchestration on top. Hindsight is a separate memory service that extracts facts from an agent's interactions and consolidates them over time. Hindsight ships a LlamaIndex integration (memory tools or a BaseMemory implementation), so the two are usually combined rather than chosen between.

Hindsight vs. LangGraph

LangGraph keeps thread-level state in checkpoints and offers its own long-term memory store shared across threads, but you decide what to write and how to search it. Hindsight handles extraction, consolidation, and reflection for you, and its LangGraph/LangChain integration includes a drop-in BaseStore adapter, so it typically plugs into LangGraph rather than replacing it.

Hindsight vs. Mastra

Mastra builds memory into its TypeScript agent framework, with message history, working memory, semantic recall, and observational memory. Hindsight is framework-agnostic, runs as its own server with Python, TypeScript, Go, REST, and MCP access, and adds graph and temporal retrieval plus a reflect API, at the cost of another service to run.

Frequently asked questions

What is Hindsight?

Hindsight is an open-source agent memory system from Vectorize. Agents send it content with retain, fetch relevant memories with recall, and ask it to reason over what it knows with reflect. It organizes memories into world facts, experiences, observations, and mental models inside isolated memory banks.

Is Hindsight free and open source?

Yes. The code is on GitHub at vectorize-io/hindsight under the MIT license, and Vectorize says self-hosting has no usage limits and no telemetry. You still pay your own LLM provider for extraction and reflection unless you run a local model.

How much does Hindsight Cloud cost?

Hindsight Cloud is pay-as-you-go on prepaid credits, with no monthly or seat fee. Retain costs $10.00 per million input tokens, recall $0.75 per million output tokens, reflect $0.05 per call, mental model refresh $0.05 per call, and storage $0.25 per million tokens per month for memories older than 30 days. New accounts get free credits (amount not published), and Enterprise pricing is custom.

Can Hindsight be self-hosted?

Yes. You can run it with a single Docker command, install it with pip, or deploy it on Kubernetes with the Helm chart. Development can use the embedded PostgreSQL (pg0). Production needs PostgreSQL 14+ with pgvector or another supported vector extension, or Oracle AI Database. An Enterprise plan adds bring-your-own-cloud and on-premises options with support SLAs.

Does Hindsight need an LLM?

Yes. Retain uses an LLM to extract facts, entities, and relationships, and reflect uses one to reason over memories. Hindsight supports 25+ providers, including OpenAI, Anthropic, Gemini, Groq, Bedrock, and Vertex AI, local options such as Ollama, LM Studio, and llama.cpp, and any OpenAI-compatible endpoint. Embeddings and reranking run locally by default.

Does Hindsight work with Claude Code or Cursor?

Yes. The `@vectorize-io/hindsight-coding-agents` package installs hooks, plugins, or MCP tools for Claude Code, Codex CLI, Cursor CLI, GitHub Copilot CLI, opencode, and other CLI agents, and builds a per-repository memory bank from git history and past sessions. It defaults to Hindsight Cloud, with self-hosted and local daemon modes. Any MCP client can also connect to a bank's MCP endpoint directly.

How is Hindsight different from RAG?

RAG retrieves chunks from a document corpus. Hindsight extracts structured facts from what an agent sees and does, links them by entity and time, consolidates them into observations and mental models, and updates those as new evidence arrives. It is aimed at memory that changes over time rather than search over static documents.

Are Hindsight's benchmark results independent?

Not fully. Vectorize's benchmarks site lists 94.6% on LongMemEval S and 92% on LoCoMo10, and the December 2025 paper reports 91.4% on LongMemEval. The README says the results were reproduced by collaborators at Virginia Tech and The Washington Post, but researchers from both are co-authors of the paper.

Integrations & fit

Claude CodeCodexCursorGitHub CopilotopencodeClineAiderZedContinueRoo CodeOpenHandsLangGraphLangChainLlamaIndexCrewAIPydantic AIOpenAI Agents SDKGoogle ADKAgnoStrandsAutoGenMicrosoft Agent FrameworkVercel AI SDKHaystackn8nZapierDifyFlowiseChatGPTPerplexityObsidianPipecatVapiHermes AgentOpenClawLiteLLMMCP clientsPostgreSQL (pgvector)Oracle AI Database
Good fit forSolo / individual, Startup / small team, Enterprise
Pricing modelFreemium· Free tier available
See pricing on Hindsight →

Alternatives to consider

About Hindsight

Hindsight runs as a separate memory server that agents call with three operations. Retain sends content such as conversations, tool output, or documents to an LLM that extracts facts, entities, relationships, and timestamps into an isolated memory bank. Recall combines semantic, keyword, graph, and time-based search, reranks the results, and trims them to a token budget. Reflect is an agentic loop that searches the bank and returns a grounded answer. In the background, Hindsight consolidates related facts into observations that are refined as evidence changes, and keeps mental models, which are standing answers to questions you define (such as a user's preferences) that an agent can read without an LLM call. That consolidation layer is what separates it from a plain vector store or from the memory built into most agent frameworks. You can self-host it with Docker, pip, or Helm on PostgreSQL with pgvector (or Oracle AI Database), or use Hindsight Cloud, which is pay-as-you-go on prepaid credits, billed per token and per call with no seat fee. It is reachable from Python, TypeScript, and Go clients, a CLI, REST, and a built-in MCP endpoint, and a single package adds per-repository memory to Claude Code, Codex, Cursor, and other CLI coding agents. The tradeoffs are cost and operations: every retain is an LLM call, reflect needs a model with solid tool calling, and self-hosting adds a stateful service to run and secure. Vectorize's benchmark lead on LongMemEval and LoCoMo is self-published, and the README itself says Hindsight may be overkill for simple n8n-style workflows.

Updates from Hindsight

New FeatureHindsight v0.10.2 released

Vectorize shipped Hindsight v0.10.2 with bank aliases for zero-downtime bank ID migrations, per-scope consolidation strategies, opt-in CUDA acceleration for ONNX embeddings, lower reflect input usage, and a fix restoring OpenAI provider compatibility with GPT-6 models.

MilestoneHindsight reaches #1 on GitHub Trending

Trendshift records the vectorize-io/hindsight repository first reaching #1 on GitHub Trending on September 26, 2026. The repository had about 43,000 stars as of September 30, 2026.

Are you the founder? Claim this listing →