Building the WordPress Connector for Cognee
How I Built a WordPress Connector for Cognee: Giving AI Real Memory for Your Blog If you’ve ever tried building an AI chatbot or search tool for a WordPress site, you've probably noticed that standard search (and basic
How I Built a WordPress Connector for Cognee: Giving AI Real Memory for Your Blog
If you’ve ever tried building an AI chatbot or search tool for a WordPress site, you've probably noticed that standard search (and basic RAG) is pretty clunky:
It chops your articles into disconnected snippets, losing track of how posts, authors, and comments connect.
If you update a single typo in a blog post, traditional tools often re-read the entire site from scratch.
If you delete an article, the AI often keeps "remembering" it and making things up.
For the Cognee hackathon, I set out to fix this by building the official WordPress connector for Cognee (PR #221).
Here’s a quick breakdown of how it works in plain English and how you can use it.
What Does Cognee Do Differently?
Most AI search tools just dump your text into a vector database. Cognee works more like human memory: it organizes your content into an interconnected knowledge graph alongside vectors.
That means when someone asks a question, Cognee doesn't just match keywords—it traces relationships between concepts, writers, categories, and discussions.
4 Practical Problems I Had to Solve
Connecting WordPress to an AI memory engine sounds simple until you actually look at the data coming out of the WordPress API. Here’s what I built to handle the real-world messiness:
Safe, Hassle-Free Login
Nobody wants to hand their main admin password to an AI script. The connector uses WordPress Application Passwords (built into WordPress 5.6+). You generate a unique passcode right inside your WP Profile, pass it to the connector, and you can revoke it anytime with one click. It only makes read-only requests (GET), so it will never change or break anything on your site.-
Stripping the HTML Clutter
WordPress doesn't store plain text—it stores Gutenberg block comments,tags, and encoded symbols like & or –. Feeding raw HTML into an AI wastes tokens and confuses the model. I built a lightweight cleaner that strips the tags, converts entities back to normal characters, and hands clean plain text to Cognee.
Syncing Only What’s New (Incremental Sync)
If you have 500 articles and publish one new tutorial, you shouldn't have to re-process all 500. The connector remembers the timestamp of the last sync and tells WordPress: "Only send me posts modified after this date." This makes updates super fast.Forgetting What You Deleted ("Forget-on-Delete")
This was one of the biggest requirements. If you delete or unpublish an article, the AI shouldn't keep answering questions with outdated information.
Each sync, the connector does a quick, lightweight check of active post IDs. If a post that was there yesterday is gone today, it flags it as deleted. Cognee then automatically purges those nodes and vectors from its memory graph.
How to Use It in 5 Minutes
Here is how simple it is to pull your WordPress content into Cognee and ask questions:
Architecture Diagram
flowchart TD
subgraph WP ["1. WordPress Site (REST API v2)"]
A["Posts, Pages, Comments & Custom Types"]
end
A -->|"Application Passwords (HTTP Basic Auth)"| B
subgraph CONNECTOR ["2. Cognee WordPress Connector"]
B["wordpress_source()"]
B --> C["Lightweight ID Sweep\n(Detects deleted posts)"]
B --> D["Incremental Fetch\n(Only gets modified posts)"]
C --> E["Tombstone Emitter\n(_deleted = True)"]
D --> F["HTML Sanitizer\n(Strips tags & unescapes)"]
end
E -->|"dlt merge table"| G
F -->|"dlt merge table"| G
subgraph COGNEE ["3. Cognee Memory Engine"]
G["Document-Mode Routing"]
G --> H["Cognify Pipeline\n(Chunking & Entity Extraction)"]
H --> I["Knowledge Graph + Vector Store\n(Relationships & Embeddings)"]
end
I --> J["4. Search & Q&A\ncognee.search(..., GRAPH_COMPLETION)"]
python
import asyncio
import os
import cognee
from cognee_community_connector_wordpress import wordpress_source
async def main():
# 1. Connect to your site (posts, pages, and comments)
source = wordpress_source(
base_url="https://your-site.com",
username=os.getenv("WORDPRESS_USERNAME"),
app_password=os.getenv("WORDPRESS_APP_PASSWORD"),
content_types=["posts", "pages", "comments"],
)
# 2. Ingest into Cognee memory
await cognee.remember(
source,
dataset_name="my_blog",
primary_key="id",
write_disposition="merge",
max_rows_per_table=0,
)
# 3. Ask your blog anything!
answer = await cognee.search(
query_text="What are the main topics and recent tutorials covered on my site?",
query_type=cognee.SearchType.GRAPH_COMPLETION,
datasets=["my_blog"],
)
print("\nAnswer from Memory:\n", answer)
if __name__ == "__main__":
asyncio.run(main())
Testing & Quality
To make sure this works reliably in production, I wrote a test suite of 11 unit tests covering everything from URL formatting to transient error retries and deletion flags. Everything runs cleanly offline and passes all linter checks (pytest + ruff).
Wrapping Up
Building this connector showed me how powerful Cognee's cognitive graph approach is compared to standard vector RAG. It bridges the gap between the world's most popular CMS and modern AI memory.
Check out the code and discussion on GitHub:
Pull Request: topoteretes/cognee-community#221
Issue: topoteretes/cognee#4789
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.