Build a Zero-Trust Document Pipeline with OpenCV, Presidio, and LangGraph (Beginner Guide)
When building LLM applications, we often rush to pass user data directly into our prompts or RAG databases. But what happens when that data is a passport, a driver's license, or a medical record? Sending raw identity doc
When building LLM applications, we often rush to pass user data directly into our prompts or RAG databases. But what happens when that data is a passport, a driver's license, or a medical record? Sending raw identity documents across an external trust boundary to an LLM API is a massive privacy risk.
In this tutorial, we are building SecureKiosk AI, an open-source "edge gatekeeper." We will use local computer vision and NLP to intercept identity documents, extract the text, and deterministically mask Personally Identifiable Information (PII) before it ever reaches an LLM.
Here is the tech stack we will use:
- OpenCV & Tesseract (Computer Vision & OCR): To clean image noise and extract raw text locally.
- Microsoft Presidio & spaCy (NLP): To detect and redact sensitive PII in memory.
- LangGraph (Orchestration): To manage the pipeline as a resilient state machine.
- Streamlit: To build a clean, interactive frontend.
Pipeline Architecture
Rather than relying on linear, brittle Python scripts that crash when text extraction yields poor results, the pipeline operates as a state machine:
The Architecture Pipeline
Letβs get started!
Step 1: Setting Up the Environment
Because we are doing native computer vision and OCR, we need OS-level dependencies alongside our Python packages.
1. System Dependencies:
If you are on Ubuntu/Debian, install the C-libraries for vision and OCR:
sudo apt-get update
sudo apt-get install -y tesseract-ocr libgl1 libglib2.0-0
2. Python Environment:
Create a virtual environment (Python 3.11 is highly recommended for compatibility) and install the core packages.
Pro-Tip: We pin numpy<2.0.0 because OpenCV is currently compiled against NumPy 1.x. Upgrading to NumPy 2.0+ will cause your CV pipeline to crash!
# requirements.txt
numpy<2.0.0
opencv-python-headless==4.9.0.80
pytesseract==0.3.10
presidio-analyzer==2.2.355
presidio-anonymizer==2.2.355
langgraph>=0.0.60
streamlit>=1.35.0
https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.7.1/en_core_web_sm-3.7.1-py3-none-any.whl
Step 2: The "Eyes" (OpenCV & Tesseract)
Raw photos of IDs are messy, they have glare, watermarks, and bad lighting. If we feed raw images straight to an OCR engine, the text extraction will fail. We need to preprocess the image using OpenCV.
Create src/cv/processor.py and src/ocr/extractor.py. We will convert the image to grayscale, apply Gaussian blur, and use adaptive binarization to make the text pop.
import cv2
import numpy as np
import pytesseract
def preprocess_image(image_bytes):
# Convert bytes to numpy array
nparr = np.frombuffer(image_bytes, np.uint8)
img = cv2.imdecode(nparr, cv2.IMREAD_COLOR)
# Grayscale and binarize for crisp OCR
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
blur = cv2.GaussianBlur(gray, (5,5), 0)
binary = cv2.adaptiveThreshold(blur, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY, 11, 2)
return binary
def extract_text(processed_image):
# Extract text from the cleaned image
return pytesseract.image_to_string(processed_image)
Step 3: The "Shield" (Zero-Trust PII Redaction)
Now that we have the raw text, we must scrub it. We will use Microsoft Presidio, powered by spaCy's Named Entity Recognition (NER).
Why Presidio Instead of Regular Expressions?
Regex matches fixed patterns like email addresses or phone numbers, but fails on variable context (e.g., extracting person names, physical addresses, or organizational entities). Presidio combines regex pattern recognition with spaCy Named Entity Recognition (NER), context words, and checksum validation.
Example Transformation
INPUT:
Issued to: Alex Vance
Passport No: P9482014
Contact: +971 50 123 4567
Address: Building 4, Al Maryah Island, Abu Dhabi
OUTPUT:
Issued to: <PERSON>
Passport No: <ALPHANUMERIC_ID>
Contact: <PHONE_NUMBER>
Address: <LOCATION>
Architectural Note: We are specifically configuring Presidio to use the lightweight en_core_web_sm model (~12MB). If you default to the large model (~580MB), your app will likely suffer from Out-Of-Memory (OOM) crashes when deployed on free cloud tiers.
Create src/privacy/redactor.py:
from presidio_analyzer import AnalyzerEngine
from presidio_analyzer.nlp_engine import NlpEngineProvider
from presidio_anonymizer import AnonymizerEngine
class PIIRedactor:
def __init__(self):
# Force the lightweight spaCy model for edge deployment
nlp_config = {
"nlp_engine_name": "spacy",
"models": [{"lang_code": "en", "model_name": "en_core_web_sm"}],
}
provider = NlpEngineProvider(nlp_configuration=nlp_config)
self.analyzer = AnalyzerEngine(nlp_engine=provider.create_engine())
self.anonymizer = AnonymizerEngine()
def redact(self, text: str) -> str:
# Detect entities like PERSON, LOCATION, DATE_TIME
results = self.analyzer.analyze(text=text, entities=[], language='en')
# Mask them deterministically
anonymized = self.anonymizer.anonymize(text=text, analyzer_results=results)
return anonymized.text
β No raw PII sent to LLM
This is the core security guarantee of the pipeline:
all document processing and PII redaction occur locally before any downstream AI system receives the content.
Step 4: The Brain (LangGraph Orchestration)
Instead of a brittle, linear script, we will orchestrate these steps using LangGraph. A state machine provides a foundation for handling OCR failures, retries, and alternate processing paths as the pipeline evolves.
Create src/graph.py:
from typing import TypedDict
from langgraph.graph import StateGraph, START, END
# Define what data moves between our nodes
class DocumentState(TypedDict):
raw_image: bytes
processed_image: object
extracted_text: str
redacted_text: str
def process_vision(state: DocumentState):
state["processed_image"] = preprocess_image(state["raw_image"])
return state
def process_ocr(state: DocumentState):
state["extracted_text"] = extract_text(state["processed_image"])
return state
def process_redaction(state: DocumentState):
redactor = PIIRedactor()
state["redacted_text"] = redactor.redact(state["extracted_text"])
return state
# Build the pipeline
workflow = StateGraph(DocumentState)
workflow.add_node("vision", process_vision)
workflow.add_node("ocr", process_ocr)
workflow.add_node("redact", process_redaction)
workflow.add_edge(START, "vision")
workflow.add_edge("vision", "ocr")
workflow.add_edge("ocr", "redact")
workflow.add_edge("redact", END)
document_pipeline = workflow.compile()
Step 5: The Interface (Streamlit)
Finally, we wrap our LangGraph pipeline in a clean Streamlit interface. Create app.py:
import streamlit as st
from src.graph import document_pipeline
st.title("π‘οΈ SecureKiosk AI: Zero-Trust Pipeline")
st.write("Upload an ID to extract and redact PII locally.")
uploaded_file = st.file_uploader("Upload Document (PNG/JPG)", type=["png", "jpg", "jpeg"])
if uploaded_file:
st.image(uploaded_file, caption="Original Upload", use_column_width=True)
if st.button("Process Document"):
with st.spinner("Running localized edge pipeline..."):
# Pass the image into our LangGraph state machine
initial_state = {"raw_image": uploaded_file.read()}
final_state = document_pipeline.invoke(initial_state)
st.subheader("Raw Extracted Text (DANGER)")
st.warning(final_state["extracted_text"])
st.subheader("Redacted Text (SAFE FOR LLMs)")
st.success(final_state["redacted_text"])
Deploying to Streamlit Community Cloud
When you are ready to share this with the world, push your code to GitHub. To deploy on Streamlit Community Cloud, you must include a packages.txt file in your repository root so the server knows to install the C-libraries before running your Python code:
# packages.txt
tesseract-ocr
libgl1
libglib2.0-0
Select Python 3.11 in your deployment settings, hit deploy, and you now have a live, edge-executed zero-trust AI pipeline!
Performance Benchmarks & Edge Footprint
The primary design goal was local in-memory execution on lightweight compute without out-of-memory (OOM) failures:
- RAM Footprint: ~145 MB steady state (by avoiding heavier transformer checkpoints and using en_core_web_sm).
- Processing Latency: 320 ms to 650 ms per page on standard 2-core CPU environments.
- Network Egress: 0 bytes prior to PII stripping.
These benchmarks were collected on a standard 2-core CPU environment without GPU acceleration.
As privacy regulations become stricter and organizations increasingly adopt LLMs, architectures that minimize data exposure before inference will become essential rather than optional.
What features do you think are most important for privacy-first AI pipelines moving forward?
Drop your thoughts below, or check out the full source code on GitHub.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.
