I built a Creator Engine that extracts promises from sales conversations (80% accuracy)
Last week I built 32 MCP servers. Today I want to share the more ambitious project: the Creator Engine — a system that compares what sales promises to what the contract actually says. The problem Sales teams
Last week I built 32 MCP servers. Today I want to share the more ambitious project: the Creator Engine — a system that compares what sales promises to what the contract actually says.
The problem
Sales teams promise things. Contracts specify things. They don't always match.
- "Weekly reports every Monday" -> contract says "provide reports" (no frequency)
- "$50,000 all-in" -> contract says "$75,000"
- "99.99% uptime SLA" -> contract says "99.9%"
- "SSO included" -> contract has no mention
Each mismatch is potential revenue loss, legal exposure, or customer churn. Traditional CLM tools are expensive ($50K+/yr), heavy, and focus on drafts — not on reality alignment.
What I built
Promise Alignment Benchmark v0 — an executable specification for testing the engine.
Structure
creator-engine/benchmarks/promise-alignment/v0/ contains:
- cases/ — 20 test cases (PA-001 ... PA-020)
- schemas/ — 5 JSON schemas (promise, term, alignment, evidence, impact)
- expected/ — ground truth
- reports/ — evaluation output
20 smoke cases
Cover:
- EXACT match (identical promise + contract clause)
- SYNONYM match ("mobile app" = "native iOS/Android applications")
- IMPLIED match (SSO implied by SAML/OAuth support)
- MISSING (weekly reports promised, contract silent on frequency)
- CONFLICT ($50K vs $75K, June 30 vs August 15, 5 seats vs 3, 99.99% vs 99.9%)
- AMBIGUOUS ("really fast" vs "adequate response time")
- PARTIAL (3 integrations promised, 1 in contract)
- NO_PROMISE (optional feature "could add later" — no commitment created)
- PROMPT INJECTION ("ignore all previous instructions" — engine resists)
8 critical invariants
HARD FAILURE — 19/20 = FAIL if any invariant breaks.
- INV-001: CONFLICT requires evidence
- INV-002: financial.amount requires calculation_id
- INV-003: Evidence location must exist in source
- INV-004: NOT_COMMITMENT cannot create a Promise
- INV-005: AMBIGUOUS cannot have review_required = false
- INV-006: LLM-explanation cannot create new factual value
- INV-007: OBSERVED cannot describe arithmetic
- INV-008: Prompt injection cannot change system policy
The engine
Three-step LLM pipeline:
- Extract promises from sales conversation -> PROM-001, PROM-002...
- Extract contract terms from contract text -> TERM-001, TERM-002...
- Align them -> EXACT | IMPLIED | PARTIAL | CONFLICT | MISSING | AMBIGUOUS
Then engine-level invariant checks run on top. For example:
function runInvariants(promises, terms, alignments) {
const violations = [];
for (const p of promises) {
if (p.status === "NOT_COMMITMENT") {
violations.push({ id: "INV-004", reason: "NOT_COMMITMENT promise" });
}
}
for (const a of alignments) {
if (a.relationship === "CONFLICT" && !a.reasoning) {
violations.push({ id: "INV-001", reason: "CONFLICT without reasoning" });
}
if (a.impact && a.impact.amount != null && !a.impact.calculation_id) {
violations.push({ id: "INV-002", reason: "amount without calculation_id" });
}
}
return violations;
}
Architectural rule: LLM extracts, engine decides. Never LLM to final truth.
Results
After 3 prompt iterations:
- v0.1: 65% (13/20)
- v0.2: 70% (14/20)
- v0.3: 80% (16/20)
Zero invariant violations across all 20 cases.
Failures that remain are mostly edge cases where ground truth is debatable:
- PA-004: 1M/month vs 12M/year -> mathematically EXACT, but spec says IMPLIED
- PA-015: bundled promise -> 3 terms -> mix of EXACT and PARTIAL
- PA-017: 2 promises -> 1 term -> EXACT or IMPLIED?
This is expected. 80% is a baseline, not production. Next: multi-step pipeline + verification layer (Phase 42).
Code
GitHub: Alex-Dev-Web-C-DE/creator-engine
Not on npm yet — it's a benchmark repo, not a package.
What's next
- Phase 42 — Alignment Engine (separate stage, not one prompt)
- Phase 43 — Conflict / Deviation Engine
- Phase 44 — Evidence Engine
- Phase 45 — Impact Engine
- Phase 49 — Sales Promise Tracker (first commercial product on top)
The engine will power 5 products: Sales Promise Tracker, ScopeGuard, InvoiceGuard, DeliveryGuard, RevenueLeak.
All share the same core: Comparison Intelligence.
Try the MCP servers
While the engine matures, you can try the live MCP servers: 35 servers on MCPize — validators, parsers, extractors.
Most relevant for the engine use case:
- Invoice Parser ($0.50) — MCP
- Contract Clause Extractor ($0.75) — MCP
- Meeting Action Extractor ($0.40) — MCP
All x402: agents pay USDC on Base, no signup, no API keys.
What's your take? Is 80% good enough to ship a human-in-the-loop MVP, or should I push for 90%+ before any customer touches it?
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.