Pause a LangGraph Agent Mid-Run for Human Approval with interrupt and a Checkpointer
Most agent demos run start to finish with no human in the loop. That is fine until the agent is about to do something you cannot take back. Issuing a refund, sending an email, deleting a record. At that point you want th
Most agent demos run start to finish with no human in the loop. That is fine until the agent is about to do something you cannot take back. Issuing a refund, sending an email, deleting a record. At that point you want the agent to stop, ask a person, and only then continue.
LangGraph gives you two primitives that make this clean: interrupt() to pause a node, and a checkpointer to persist state so the run can resume later, even in a different process. In this tutorial I will walk through the technique using the shape of a small support-triage agent I built called RelayG, where a refund over a threshold has to be approved by a human before it goes out.
The goal is not to copy my project. It is to understand the mechanism so you can drop it into any graph.
The mental model
A LangGraph agent is a state machine. Nodes read a shared state dict and return partial updates. Edges decide what runs next. Normally you call graph.invoke(input, config) and it runs every node to the end.
interrupt() breaks that. When a node calls it, LangGraph:
- saves the current state to the checkpointer,
- stops the run,
- hands the interrupt payload back to your caller.
Nothing after the interrupt() call runs yet. Later you resume by invoking the graph again with a special Command(resume=...) value. LangGraph reloads the saved state, re-enters the same node, and this time interrupt() returns the value you resumed with. Execution continues from that point.
Two things make this work reliably. The checkpointer, which is the durable memory. And the thread_id, which is the key that ties a run to its saved state.
Step 1: a state and a graph
Start with a typed state and a plain three-node graph: classify the ticket, check policy, act on it.
import operator
from typing import Annotated, TypedDict
from langgraph.graph import START, END, StateGraph
from langgraph.checkpoint.base import BaseCheckpointSaver
class TicketState(TypedDict, total=False):
ticket_id: str
subject: str
body: str
policy_decision: str
approved: bool
actions: Annotated[list[dict], operator.add]
def build_graph(checkpointer: BaseCheckpointSaver | None = None):
builder = StateGraph(TicketState)
builder.add_node("classify", classify)
builder.add_node("policy_check", policy_check)
builder.add_node("act", act)
builder.add_edge(START, "classify")
builder.add_edge("classify", "policy_check")
builder.add_edge("policy_check", "act")
builder.add_edge("act", END)
return builder.compile(checkpointer=checkpointer)
Note the Annotated[list[dict], operator.add] reducer on actions. That tells LangGraph to append across updates instead of overwriting, which matters when a node runs, pauses, and runs again after resume.
The one thing to notice here: compile() takes the checkpointer. Without it, interrupt() has nowhere to save state and cannot resume. This is the single most common mistake.
Step 2: pause inside a node with interrupt()
The act node is where the human gate lives. When policy says a refund needs approval, we call interrupt() with a payload describing the decision, and we wait.
from langgraph.types import interrupt
def act(state: TicketState) -> dict:
ticket_id = state["ticket_id"]
decision = state["policy_decision"]
actions: list[dict] = []
update: dict = {}
if decision == "needs_approval":
amount = state.get("refund_amount", 0.0)
# Execution stops here. The payload goes back to the caller.
verdict = interrupt(
{
"ticket_id": ticket_id,
"question": f"Approve refund of ${amount:.2f}?",
"reason": state.get("policy_reason", ""),
}
)
# This line only runs AFTER a human resumes the graph.
approved = bool(verdict.get("approved"))
update["approved"] = approved
if approved:
actions.append(issue_refund(ticket_id, amount))
else:
actions.append(send_reply(ticket_id, "Refund declined after review."))
elif decision == "auto_approve":
amount = state.get("refund_amount", 0.0)
actions.append(issue_refund(ticket_id, amount))
update["actions"] = actions
return update
The mental trick is to read interrupt() as a function that returns twice. The first time the node runs, interrupt() does not return at all. It throws control back up to the caller. The second time, after you resume, interrupt() returns the resume value and the rest of the node runs normally.
Because everything before interrupt() runs again on resume, keep any irreversible side effects after the interrupt, not before it.
Step 3: a durable checkpointer
For a demo you could use MemorySaver, but that loses state when the process dies, which defeats the point. A real approval might take hours. Use the SQLite checkpointer so state survives restarts.
import sqlite3
from langgraph.checkpoint.sqlite import SqliteSaver
conn = sqlite3.connect("relayg_checkpoints.sqlite", check_same_thread=False)
checkpointer = SqliteSaver(conn)
graph = build_graph(checkpointer)
check_same_thread=False matters because LangGraph may touch the connection from a different thread than the one that opened it.
Step 4: run, detect the pause, resume
Every run needs a thread_id in the config. That id is the address of this ticket's saved state. Use the same id to resume.
from langgraph.types import Command
config = {"configurable": {"thread_id": "T-1002"}}
ticket = {
"ticket_id": "T-1002",
"subject": "Refund request",
"body": "The product broke after a week. I want a refund of $120.",
}
result = graph.invoke(ticket, config)
if "__interrupt__" in result:
payload = result["__interrupt__"][0].value
print("Graph paused, waiting on a human:", payload["question"])
# ... this can happen minutes or days later, in another process ...
verdict = {"approved": True, "note": "Verified purchase, within policy."}
result = graph.invoke(Command(resume=verdict), config)
print("actions:", result["actions"])
When the graph pauses, the return value carries an __interrupt__ key. Read result["__interrupt__"][0].value to get the payload you passed to interrupt(). That is what you show the reviewer.
To resume, invoke the graph again with the same config and a Command(resume=verdict) as the input. The dict you pass becomes the return value of interrupt() inside the node. Because SQLite persisted the state, you can restart your whole program between the pause and the resume, load the checkpointer again, invoke with the same thread_id, and it picks up exactly where it stopped. No re-classification, no lost context.
One honest caveat
Everything in a node before interrupt() runs a second time when you resume. LangGraph replays the node from the top; it does not freeze mid-function. So if you did something with a side effect before the interrupt call, like posting to Slack or charging a card, it can happen twice. Keep the code above interrupt() pure, do reads and computation there, and put the irreversible action after the interrupt returns. This is a design rule, not a bug you can configure away.
Wrapping up
That is the whole technique. Compile with a checkpointer, call interrupt() at the decision point, detect __interrupt__ in the result, and resume with Command(resume=...) on the same thread_id. The state machine and durable state do the heavy lifting, so a run can span a human's coffee break or a server restart without losing its place.
If you want to see it wired end to end, with a policy layer, a mock and a real classifier, and an audit trail of every action, the full support-triage version lives at github.com/AgentPostmortem/relayg. Clone it, run the demo with no API key, and watch a refund pause for approval and resume.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ full credit and traffic to the original publisher.