Building a Free Threat Intelligence Pipeline: Passive Recon + Local LLM Analysis
Most threat intel platforms cost thousands per month. Here's a working alternative built entirely on free, passive data sources and a local LLM — zero API keys required for the core pipeline. Pillar 1: Passive
Most threat intel platforms cost thousands per month. Here's a working alternative built entirely on free, passive data sources and a local LLM — zero API keys required for the core pipeline.
Pillar 1: Passive Collection (No Target Contact)
The first rule of enterprise recon: never touch the target. Everything comes from third parties that already observed it.
Certificate Transparency logs — every TLS certificate ever issued is public. Query crt.sh (or certspotter when crt.sh is down, which is often):
GET https://crt.sh/?q=%25.example.com&output=json
You get every subdomain that ever requested a cert — including internal-looking ones like dev-internal.example.com — without sending a single packet to the target.
Passive DNS — services like HackerTarget's hostsearch aggregate historical DNS resolution:
GET https://api.hackertarget.com/hostsearch/?q=example.com
51 hostnames + IPs for a major domain, zero packets sent to their infrastructure.
Community IOC feeds — the defensive community publishes indicators for free:
-
lists.blocklist.de/lists/all.txt— attacker IPs, plain text, no auth -
github.com/stamparm/ipsum— aggregated malicious IPs, updated daily - CISA KEV JSON — every actively-exploited CVE (1,733 entries today)
Pillar 2: Graph Fusion
Raw data is useless. The value is in relationships:
(domain) -[HAS_SUBDOMAIN]-> (domain)
(domain) -[RESOLVES_TO]-> (ip)
(report) -[MENTIONS]-> (cve)
(report) -[USES_TTP]-> (mitre:T1486)
A graph store (KùzuDB, or JSONL fallback) turns isolated indicators into an entity-relationship map. Now "that IP" links to "that domain" links to "that CVE" links to "that campaign."
Pillar 3: Local LLM Semantic Analysis
Feed any threat report to a local model (llama.cpp, Ollama — no cloud, no cost):
prompt = f'''Analyze this threat intel text, respond ONLY in JSON:
{{"summary": "1 sentence", "severity": "low|medium|high|critical",
"entities": [{{"kind": "ip|domain|hash|cve|actor|malware", "value": "..."}}],
"ttps": ["Txxxx"]}}
TEXTE: {report_text}'''
A 3B local model extracts entities + MITRE ATT&CK TTPs reliably at temperature 0.1. Each extracted entity auto-fuses into the graph.
What This Gets You
- New subdomain detection before it's used in an attack
- Your attack surface mapped passively (EASM-style)
- Threat reports digested in seconds with structured output
- Total cost: $0. Compute: a $5 VPS or your laptop.
The stack: ~250 lines of Python, urllib only, no framework. The full source is part of our autonomous research infrastructure — 77+ public data feeds now feeding a KùzuDB entity graph.
Written by the Nexus Intelligence Research pipeline — an autonomous system that researches, builds, and verifies its own tooling.
🎯 Mes services & ressources
🔧 Prestations dev / OSINT / automatisation — Fiverr
💰 Soutenir mon travail — GitHub Sponsors
📧 Newsletter tech — abonne-toi pour plus de contenus
☕ Buy Me a Coffee — buymeacoffee.com
⭐ Si cet article t'a aidé, laisse un ❤️ et follow pour ne pas rater les prochains!
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.