Ask a stateful LLM agent framework how it "remembers" a conversation, and the honest answer is usually: it doesn't, by itself — the caller resends the whole history every turn. That works until an agent runs a multi-step tool-calling loop that takes thirty seconds, the process restarts halfway through, and now you need to either resume exactly where it left off or explain to a user why their request silently vanished. The piece that makes resumption possible is a checkpointer, and the way one is actually wired into two different agent shapes — not just how it's implemented on top of Postgres — is what clarifies what "agent memory" really means underneath the abstraction.
What a checkpoint actually has to capture
A checkpoint is not just "the conversation so far." For a graph-based agent executing a sequence of steps, a checkpoint has to capture the full state object at a point in time — not just messages, but whatever typed state the graph is threading through — which node executes next, so resumption doesn't just restart from the top, a causal link to the previous checkpoint so history can be walked backward, and pending writes that haven't been merged into a checkpoint yet, because a step can fail after producing a partial result but before that result is committed as the new checkpoint. That last point is why a single state table isn't enough — you need three tables with genuinely distinct responsibilities:
-- One row per saved snapshot of the graph's state, per thread.
CREATE TABLE checkpoints (
thread_id TEXT NOT NULL,
checkpoint_id TEXT NOT NULL,
parent_checkpoint_id TEXT, -- forms a linked list per thread
checkpoint JSONB NOT NULL, -- serialized state
metadata JSONB NOT NULL, -- step number, source, etc.
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
PRIMARY KEY (thread_id, checkpoint_id)
);
-- Large channel values, out-of-line so the hot-path row stays small.
CREATE TABLE checkpoint_blobs (
thread_id TEXT NOT NULL,
checkpoint_id TEXT NOT NULL,
channel TEXT NOT NULL,
blob BYTEA,
PRIMARY KEY (thread_id, checkpoint_id, channel)
);
-- Writes produced by a step before they're folded into the next checkpoint —
-- the write-ahead log that makes crash recovery possible.
CREATE TABLE checkpoint_writes (
thread_id TEXT NOT NULL,
checkpoint_id TEXT NOT NULL,
task_id TEXT NOT NULL,
idx INT NOT NULL,
channel TEXT NOT NULL,
value BYTEA,
PRIMARY KEY (thread_id, checkpoint_id, task_id, idx)
);
Only the latest checkpoint per thread actually gets loaded on the next request — the parent chain exists for potential time-travel debugging and conversation branching, not because every request walks it. The split between checkpoints and checkpoint_blobs matters for a boring but real reason: agent state frequently contains large values — retrieved documents, tool outputs, images. Storing them inline in the row you query on every single step read means every read pays for data you often don't need just to know "what state are we in." Pulling large values into a separate table keyed the same way keeps the hot-path read cheap; large blobs get fetched only when a specific channel's value is actually needed.
Two usage patterns, deliberately different
The part that's easy to get wrong is treating "use a checkpointer" as one decision. In practice a system with more than one kind of agent ends up with two genuinely different wiring patterns, and conflating them produces exactly the bug the second pattern exists to avoid.
Pattern A — compiled-in, fully automatic. For a tool-calling, ReAct-style graph, the checkpointer is compiled directly into the graph. Loading and saving prior messages by thread ID happens automatically, with no manual plumbing. Before each model call, the agent node prepends an ephemeral preamble — system prompt, available tools, recalled memories, any active working context — onto the loaded message history, but the node's return value is only the new message, never the preamble. The preamble is never persisted. It gets rebuilt from scratch on the next turn. This is deliberate: system prompts change across deploys, recalled memories differ per query, and persisting a stale preamble would silently duplicate or contradict the fresh one computed on the next request.
Pattern B — fully stateless, manually orchestrated. For a retrieval-driven chat flow, the graph itself has no checkpointer compiled in at all, specifically because that graph carries a dozen or more per-request working fields — retrieved documents, search results, intent classification — that would leak stale values across unrelated requests if the whole graph state were checkpointed. Instead, a separate orchestration layer manually loads history before invoking the graph and manually appends the new turn afterward, treating the checkpointer purely as a message log rather than as full graph state.
# Pattern A: compiled in, fully automatic
graph = builder.compile(checkpointer=postgres_saver)
graph.invoke({"messages": [user_message]}, config={"configurable": {"thread_id": tid}})
# prior messages loaded, new ones saved — no manual step
# Pattern B: no checkpointer on the graph; orchestration owns persistence
history = checkpointer.load(thread_id) # explicit load
result = stateless_rag_graph.invoke({"messages": history + [user_message]})
checkpointer.save(thread_id, user_message, result) # explicit save
The architectural line underneath both patterns is the same: anything that changes per-deploy or per-query — system prompt, recalled memories, retrieved documents, intent flags — is rebuilt fresh every request and never persisted. Only the actual conversational turns are durable. Get this backwards even once — checkpoint a graph that carries per-request retrieval state — and a later request on the same thread can silently see a previous query's retrieved documents leak into context it has no business seeing.
Resumption at scale isn't "load everything"
For Pattern B, loading history naively — every message, every turn, forever — doesn't scale and doesn't help the model anyway; most of it is irrelevant to the current turn and it costs tokens on every single request. The actual load path layers several cheap mechanisms:
- Sliding window. Only the most recent N messages (roughly 60, i.e. ~30 user/assistant pairs) are pulled from the checkpointer directly.
- Presentation-markup stripping at read time. Whatever rich formatting the UI layer adds — citation tags, artifact markers — gets stripped before the text reaches the model, while the raw version stays in the checkpointer untouched for debugging. The reduction from this step alone is substantial: a message that's a thousand characters with markup intact can shrink by 70% or more once stripped down to what the model actually needs to read.
- Summarizing what falls outside the window, not dropping it. Messages older than the sliding window aren't discarded — they're summarized by a cheap, fast model tier, and that summary is cached (keyed by thread and message count, with a short TTL) so a second request against the same long thread doesn't re-pay the summarization cost.
- A token-budget trim pass as the last safety net. If the cleaned-and-summarized result still overflows a fixed token budget, the oldest remaining content is dropped first, not the most recent.
- An explicit "most recent exchange" marker. Inserted right before the final turn specifically to help the model resolve pronoun references like "it" or "that" back to the thing being discussed one turn ago, which otherwise gets buried under a wall of summarized history.
[ older turns, summarized + cached ] → [ sliding window of ~60 raw messages ]
→ [ MOST RECENT EXCHANGE marker ]
→ current turn
Idempotency without a distributed lock
A network retry that calls "save this turn" twice should never produce two copies of the same message. The mechanism here is almost embarrassingly simple once you see it: every message gets a deterministic ID derived by hashing the thread ID, the role, and the content together. The merge logic that appends messages deduplicates by that ID, so calling save twice with identical content is a safe no-op rather than a duplicate — with no distributed lock, no idempotency-key table, and no coordination between retrying callers. The ID is the idempotency key, derived from the data itself rather than generated and tracked separately.
A feature that got measured, then removed
Not every addition earns its cost, and this system has a documented example of reversing one. An earlier version included a cross-conversation "memory recall" layer — scoring a separate facts table by topical overlap and recency to surface things the user had mentioned in other conversations. It added somewhere between 100 and 200 milliseconds to every single request. After measurement, it was removed from the regular chat path specifically because the value was marginal: within-conversation context already comes from the checkpointer, and document knowledge already comes from retrieval, so the facts-table lookup was mostly restating what the other two sources already covered, at a real and constant latency cost. Notably, only the recall path was cut — the extraction step that writes facts to that table every several messages was kept, because a future feature might read it, and writing is nearly free compared to the per-request cost of reading and scoring it on every turn. It's a concrete example of a latency-vs-value tradeoff that got reversed after actually measuring it, rather than kept indefinitely on the assumption that more context can only help.
Two stores, deliberately not unified
A natural instinct is to ask why there's a separate application-level message table and a checkpointer-owned message log, instead of one source of truth. They serve genuinely different consumers: the application's own table is the system of record for what a user sees rendered in the UI, for usage/billing, and for the fact-extraction step above; the checkpointer is the system of record for exactly what the model sees as conversational context. Collapsing them would couple two things that change for different reasons — a UI rendering change has no business forcing a migration of what the model reads as history, and vice versa.
Shipping the two sides independently
Because the checkpointer-owning service and the application backend are deployed independently, the message-loading path supports three generations of request shape at once — a bare thread ID, a pre-built history string, or a fully structured conversation array — so one side can ship a new format before the other side has caught up, rather than requiring a synchronized, coordinated cutover between two services that don't share a deploy pipeline.
What this buys an agent in practice
Once state, blobs, and pending writes are modeled as three distinct concerns instead of one JSON blob, and "what gets checkpointed" is scoped deliberately per agent shape instead of applied uniformly: a multi-step tool-calling loop survives a process restart without the caller noticing anything beyond latency, because the write-ahead log replays exactly the pending work that hadn't yet been folded into a checkpoint — not the last clean snapshot, and not a half-applied mutation. Large tool outputs don't bloat every state read, because they're fetched on demand by channel. Conversation branching comes for free from the parent-pointer chain, useful for retrying a step with different parameters or building a debugging timeline. And a long-running thread stays affordable, because the sliding window plus cached summarization means token cost per request stays roughly flat instead of growing linearly with how long the conversation has been going on. The unglamorous part — three tables, a write-ahead log, a replay function, two different wiring patterns for two different agent shapes — is also the part that actually makes "the agent remembers" true under failure and at scale, rather than true only in the demo where nothing ever crashes mid-step and no conversation runs past a dozen turns.

