Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 6 min read

Build Persistent Memory for an AI Agent with Python

AI agents are good at responding to the conversation in front of them. The harder problem is remembering something after that conversation ends. A user might tell a support agent: their preferred contact method an acc

AI agents are good at responding to the conversation in front of them. The harder problem is remembering something after that conversation ends.

A user might tell a support agent:

  • their preferred contact method
  • an account identifier
  • a product configuration
  • the issue they previously reported

Without persistent memory, the next session begins from zero and the agent asks for the same information again.

In this tutorial, we’ll build a small Python CLI that sends a conversation to Telnyx Agent Memory, waits for the facts to be extracted, and then recalls those facts with a natural-language query.

The complete example is available here:

github.com/team-telnyx/telnyx-code-examples/tree/main/persistent-ai-agent-memory

What we’re building

The example follows a simple three-step workflow:

Conversation transcript
        |
        v
1. Ingest transcript
        |
        v
2. Poll async operation
        |
        v
3. Recall ranked facts

The CLI:

  1. Submits a support conversation for a profile.
  2. Receives an asynchronous operation ID.
  3. Polls until fact extraction finishes.
  4. Asks what the user’s preferred contact method is.
  5. Prints the matching memories in relevance order.

This pattern gives an agent continuity without placing every previous conversation into its prompt.

The Agent Memory model

Agent Memory organizes information using a few core concepts:

  • Namespace: an isolation boundary for an application or environment.
  • Profile: the person or entity the memories describe.
  • Source: the original session or fact submitted to the API.
  • Memory: an individual fact extracted from a source.
  • Operation: an asynchronous write job that can be tracked to completion.

The example uses the default namespace and a profile named user_123.

A real application could use an existing customer, account, or caller identifier as the profile ID.

The three API calls

Every endpoint is relative to:

https://api.telnyx.com/v2/ai/memory

The example uses these calls:

POST /namespaces/{namespace}/profiles/{profile_id}/ingest
GET  /namespaces/{namespace}/operations/{operation_id}
POST /namespaces/{namespace}/profiles/{profile_id}/recall

Let’s walk through each one.

1. Ingest a conversation

First, load the API key and prepare the authentication headers:

import os
import requests
from dotenv import load_dotenv

load_dotenv()

API_KEY = os.environ["TELNYX_API_KEY"]
BASE_URL = "https://api.telnyx.com/v2/ai/memory"

headers = {
    "Authorization": f"Bearer {API_KEY}",
    "Content-Type": "application/json",
    "Accept": "application/json",
}

Here is a shortened version of the support transcript:

transcript = [
    {
        "role": "user",
        "content": "Messages are failing for some international destinations.",
    },
    {
        "role": "assistant",
        "content": "Let me check your account configuration.",
    },
    {
        "role": "user",
        "content": "My preferred contact method is email at [email protected].",
    },
]

Submit the messages to a profile:

namespace = "default"
profile_id = "user_123"
session_id = "demo-session-001"

response = requests.post(
    f"{BASE_URL}/namespaces/{namespace}/profiles/{profile_id}/ingest",
    params={"session_id": session_id},
    headers=headers,
    json={"messages": transcript},
    timeout=30,
)

response.raise_for_status()
operation_id = response.json()["data"]["operation_id"]

print(f"Ingest accepted: {operation_id}")

The API returns 202 Accepted, not a completed memory.

A typical response looks like this:

{
  "data": {
    "operation_id": "op_abc123",
    "profile_id": "user_123",
    "session_id": "demo-session-001",
    "source_id": "src_xyz789"
  }
}

Fact extraction happens asynchronously. The operation_id is the handle we use to track it.

2. Poll until the write finishes

A memory cannot be recalled until its write operation has completed.

The sample checks the operation every two seconds and stops after reaching a terminal state:

import time

terminal_statuses = {"completed", "failed", "cancelled"}

while True:
    response = requests.get(
        f"{BASE_URL}/namespaces/{namespace}/operations/{operation_id}",
        headers=headers,
        timeout=30,
    )
    response.raise_for_status()

    status = response.json()["data"]["status"]
    print(f"Operation status: {status}")

    if status in terminal_statuses:
        break

    time.sleep(2)

The normal progression is:

pending -> processing -> completed

Production code should also enforce a timeout. The complete example stops polling after 60 seconds and raises an error rather than waiting forever.

This explicit operation lifecycle is useful because your application can distinguish between:

  • a write that is still processing
  • a write that completed
  • a write that failed
  • a write that was cancelled

If the client restarts while polling, it can continue checking the same operation instead of submitting the entire conversation again.

3. Recall relevant facts

Once ingestion completes, query the profile using natural language:

response = requests.post(
    f"{BASE_URL}/namespaces/{namespace}/profiles/{profile_id}/recall",
    headers=headers,
    json={
        "query": "What is the user's preferred contact method?",
        "top_k": 10,
    },
    timeout=30,
)

response.raise_for_status()
facts = response.json()["data"]

The results are ranked by relevance:

for fact in facts:
    print(f"[score={fact['score']}] {fact['text']}")

A result may look like this:

{
  "id": "mem_abc123",
  "text": "The user's preferred contact method is email at [email protected].",
  "recorded_at": "2026-10-05T12:00:05Z",
  "score": 0.92
}

The response contains the extracted fact rather than requiring the application to search the original transcript itself.

That fact can now be added to an agent’s context when a new conversation begins.

Run the complete example

Clone the repository:

git clone https://github.com/team-telnyx/telnyx-code-examples.git
cd telnyx-code-examples/persistent-ai-agent-memory

Create a .env file:

echo "TELNYX_API_KEY=your_telnyx_api_key_here" > .env

Install the dependencies:

pip install -r requirements.txt

Run the CLI:

python app.py

The example requires only a Telnyx API key. It does not require a phone number, messaging profile, or separate database.

Details worth handling in production

The sample includes a few useful safeguards that are easy to overlook.

Encode path identifiers

Namespace, profile, and operation identifiers should be encoded before being inserted into a URL:

from urllib.parse import quote

profile_id = quote(profile_id, safe="")

This prevents reserved characters in an existing customer identifier from changing the request path.

Validate inputs before sending them

The example enforces several API constraints:

if len(session_id) > 128:
    raise ValueError("session_id must be <= 128 characters")

if not query or len(query) > 4096:
    raise ValueError("query must be between 1 and 4096 characters")

if not 1 <= top_k <= 100:
    raise ValueError("top_k must be between 1 and 100")

Failing locally gives developers a clearer error than sending an invalid request and debugging it later.

Do not recall immediately after ingesting

An accepted write is not yet a recallable memory.

If recall returns an empty list immediately after ingestion, first confirm that the operation reached completed.

Choose isolation boundaries deliberately

Profiles are isolated from one another, and namespaces provide an additional boundary between applications or environments.

For example:

namespace: production-support
profile: customer_1024

You might use separate namespaces for development and production, while each customer or caller receives a unique profile inside that namespace.

Custom namespaces must be created before use. The default namespace is available without a separate provisioning step.

Plan for memory deletion

Persistent memory should come with a deletion strategy.

The Agent Memory API includes endpoints for deleting an individual source or an entire profile. That allows an application to remove one imported session or erase everything associated with a profile when required.

Where this pattern fits

The same ingest, poll, and recall workflow can support:

  • Customer support agents that remember previous issues and preferences
  • Sales assistants that maintain account context across conversations
  • Scheduling agents that recall availability or communication preferences
  • Voice agents that recognize returning callers
  • Internal copilots that retain user-specific working context

The communication channel can change while the memory model stays the same. SMS, voice, email, and browser chat can all resolve to the same profile identifier.

Final thoughts

A longer prompt is not the same thing as long-term memory.

Prompt context is temporary. Persistent memory gives an agent a durable place to store what it has learned and retrieve only the facts relevant to the next interaction.

The implementation here stays intentionally small:

ingest -> poll -> recall

That is enough to move an agent from isolated sessions toward real continuity.

Resources

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.