From Job Scraper to Production AI System: Building ApplyLens AI
Job searching started feeling like a second job. I was jumping between different job sites and ATS platforms, seeing duplicate or stale postings, comparing multiple resume versions, and still manually trying to answer t
Job searching started feeling like a second job.
I was jumping between different job sites and ATS platforms, seeing duplicate or stale postings, comparing multiple resume versions, and still manually trying to answer the questions that actually mattered:
Is this role really worth applying to? Which resume should I use? What am I missing? Should I tailor first โ or move on?
There are already strong products that solve valuable parts of this workflow: resume analysis, job tracking, autofill, matching, and application organization.
I wasn't trying to claim that those problems had never been solved.
The engineering question that interested me was different:
What would happen if I treated the entire workflow as one AI systems problem instead of a collection of separate tools?
So I started with something much simpler: a job scraper.
Somewhere between ATS integrations, deduplication, ranking, LLM evaluation, resume matching, provider routing, retries, observability, agentic workflows, persistence, and production deployment... things got slightly out of hand. ๐
That project became ApplyLens AI โ a production-deployed AI job intelligence and application planning platform.
Live deployment: applylensjobs.com (access-controlled)
Source code: GitHub โ ApplyLens AI
What ApplyLens does
At a high level, ApplyLens combines job acquisition, deterministic filtering, AI evaluation, resume intelligence, application prioritization, evidence-grounded tailoring, retrieval, and human review within one production system.
Three numbers give a useful sense of the current scope:
- 11 acquisition adapters across ATS platforms and job sources
- a 16-stage observable processing pipeline
- shared production job acquisition running automatically every 6 hours
The acquisition layer currently includes:
Workday, Greenhouse, Lever, Ashby, Workable, Jobvite, Recruitee, SmartRecruiters, Built In, USAJobs, and Himalayas.
But the number of sources is not the part of the architecture I find most interesting.
The important part is what happens after the jobs arrive.
The architecture in one view
The system can be simplified into this flow:
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Job / ATS Sources โ
โ 11 acquisition adapters โ
โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Shared Job Intelligence โ
โ PostgreSQL corpus โ
โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโ
โ
โโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโ
โ โ
โผ โผ
Deterministic Processing User-Specific Projection
filtering / dedupe / rank preferences / resumes /
seen state / AI routing
โ โ
โโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโ
โผ
JD Intelligence
โ
โผ
Bounded LLM Evaluation
โ
โผ
Resume Matching
โ
โผ
Final Application Scoring
โ
โผ
Planning / Tailoring / Review
โ
โผ
Human Decision
There are several design decisions behind this that became much more important than the original scraping problem.
1. Collect once. Personalize independently.
One of the first architectural decisions was separating global acquisition from user-specific intelligence.
A naive multi-user design could effectively repeat acquisition work for every user.
I didn't want that.
Instead, ApplyLens maintains a shared job corpus and then projects that data through an authenticated user's own workflow.
The shared layer handles job acquisition.
The personalized layer owns things such as:
- role and seniority preferences
- location preferences
- excluded terms
- preferred skills
- seen-job state
- resume variants
- ranking context
- AI credentials
- provider/model routing
- application planning state
In other words:
Collect once. Personalize independently.
This also creates a useful operational boundary.
The production job acquisition schedule runs every six hours in global-acquisition-only mode. It refreshes the shared job corpus without silently launching user-specific planning, tailoring, or application decisions.
Personalization happens separately.
That separation became one of the most important pieces of the system.
2. I deliberately did not let one LLM make the whole decision
It is very easy to design an AI application like this:
Resume + Job Description
โ
LLM
โ
Fit Score
It is simple.
It is also not the architecture I wanted.
ApplyLens separates three responsibilities:
Deterministic relevance filtering
โ
Bounded LLM evaluation
โ
Final application scoring
Those are deliberately different stages with different authority.
Deterministic prefiltering
Before an expensive model needs to interpret anything, deterministic logic can evaluate things such as:
- role/title relevance
- seniority
- location policy
- freshness
- exclusions
- duplicate identity
- previously seen jobs
- early evidence rules
There is no reason to ask an LLM whether a posting is three weeks old or whether a title violates an explicit exclusion rule.
Those are deterministic questions.
LLM evaluation
The LLM becomes useful later, where interpretation matters.
For example:
- interpreting job-description requirements
- reasoning over less explicit fit signals
- enriching structured intelligence
- producing grounded tailoring assistance
But the model operates on a bounded candidate set rather than becoming the first and only filter.
Final application scoring
Final application-priority scoring remains a separate responsibility.
This matters because an LLM evaluation is not automatically equivalent to the final decision the application should make.
A relevance score, model interpretation, resume-match score, optimization score, and final application-priority score represent different things.
I wanted the architecture to preserve those distinctions instead of collapsing them into one mysterious number.
3. Different AI workloads should not automatically use the same model
Another thing I wanted to avoid was this:
MODEL = "whatever-model-is-currently-best"
followed by sending every AI task through it.
ApplyLens treats AI work as workload-specific.
Different workloads include areas such as:
- skill extraction
- job-fit evaluation
- JD intelligence
- grounded RAG answers
- resume fallback ranking
- ambiguous resume adjudication
- critic evaluation
- tailoring generation
- tailoring refinement
- tailoring judging
- manual Scan phrase generation
The application maintains explicit provider/model qualification and routing rather than assuming that one preferred provider should own every task.
For authenticated workflows, the effective route is resolved for the specific workload.
A user's generic provider preference does not silently override a workload-specific qualified route.
And if required routing, qualification, or credentials are unavailable, sensitive AI paths can fail closed rather than silently switching behavior.
That was an important lesson for me:
Multi-model support is not the same thing as model routing.
Supporting several APIs is easy.
Defining which workload is allowed to use which model, how that decision is resolved, and what happens when that route is unavailable is the actual systems problem.
4. Resume tailoring has an evidence boundary
Resume optimization creates another tempting failure mode for generative AI.
If the system's only goal is "make this resume look more relevant," a model can produce wording that sounds excellent while quietly introducing experience that never existed.
I didn't want that behavior.
The tailoring path is built around resume and job evidence.
Conceptually:
Resume evidence + JD evidence
โ
Rewrite direction
โ
Candidate generation
โ
Validation
โ
Replacement selection
โ
Human review/export
Generated changes are expected to remain grounded in supplied evidence.
Unsupported tools, skills, metrics, methods, domains, or responsibilities are rejected or kept as directional guidance instead of being silently inserted into the resume.
And the source resume itself is not overwritten by generation.
That distinction is important to me because a resume assistant should help express real evidence better โ not manufacture new evidence.
5. The human remains the authority for consequential actions
This eventually became one of the core design principles behind ApplyLens.
AI can:
- analyze
- interpret
- recommend
- prioritize
- explain
- draft
- critique
- assist with tailoring
But ApplyLens does not give the AI unrestricted authority over consequential actions.
It does not silently:
- submit applications to an ATS
- mark jobs as applied
- message recruiters
- overwrite source resumes
- turn an AI recommendation into application approval
Human review and application execution remain separate.
Even optional adjudication is intentionally constrained: commentary can be added without silently overriding the authoritative resume winner, final score, ranking, queue, or action.
I think this boundary becomes more important as AI systems become more agentic.
The interesting question is no longer only:
"Can the model do this?"
It is also:
"Should the model have authority to do this?"
Those are very different questions.
6. Agentic does not have to mean autonomous
I also experimented with agentic orchestration and LangGraph.
But I did not want to rebuild the entire application around an "agent owns everything" architecture just because agent frameworks became available.
ApplyLens keeps its existing deterministic and AI owners, while selected stages can be routed through bounded, explicitly gated LangGraph paths.
Those guarded paths cover areas such as:
- deterministic prefilter/deduplication
- JD intelligence
- semantic evaluation
- final scoring
- prioritization
- tailoring decisions
- tailoring generation
- conditional operator review
The important word there is bounded.
These paths are not permission for an agent to invent a new application state or bypass an existing authority.
Agent recommendations, trace persistence, human checkpoints, and authoritative application actions remain separate concerns.
Many of these capabilities are also deliberately default-off unless explicitly enabled.
That may sound less exciting than saying "fully autonomous AI agent."
I think it is better engineering.
7. Observability became part of the product
Once a pipeline has enough moving pieces, "it failed" is no longer useful information.
I wanted to be able to answer:
- Which stage is currently running?
- Which stages completed?
- How many jobs survived each stage?
- Which source degraded?
- Was a failure transport-related, parsing-related, or pagination-related?
- Did an AI workload hit cache?
- Which provider/model route was resolved?
- Did parsing fail?
- Was there a retry?
- What artifacts were produced?
- What happened in an agentic workflow?
- What is the scheduler doing?
- Did a persisted operation actually succeed?
That led to several operational surfaces inside the application:
- Pipeline Dashboard
- Advanced Diagnostics
- Agentic Operations
- Run-scoped Agentic Review
- Scheduler Health
- run/artifact inspection
- notification and operational state
- admin and Super User controls
The runtime itself tracks a 16-stage status model:
startup
โ scraping
โ filtering
โ dedupe
โ ranking
โ cache_filter
โ details
โ intelligence
โ ai_evaluation_filter
โ embedding_prefilter
โ ai_evaluation
โ resume_matching
โ application_priority
โ rag_export
โ planning
โ finalization
Different runtime modes can substitute or bypass work where appropriate โ for example, an authenticated shared-corpus projection uses shared input rather than scraping โ but the stage model provides a common observability contract.
The lesson here was simple:
For production AI, observability is not an admin afterthought. It is part of the architecture.
8. Acquisition needed its own reliability model
Scraping multiple ATS platforms is not simply "send HTTP requests in a loop."
Different sources have different:
- APIs
- pagination rules
- rate limits
- HTML/data structures
- detail endpoints
- completeness characteristics
- failure modes
So acquisition has bounded retry, timeout, pagination, and concurrency behavior.
Source health distinguishes states such as:
SUCCESSEMPTYPARTIALFAILED
Failures are classified rather than flattened into one generic exception.
The system also tracks data-quality signals such as URL, timestamp, and description completeness.
Why does that matter?
Because "this source returned zero relevant jobs" and "this source silently failed halfway through pagination" are not the same operational event.
If downstream AI consumes bad or incomplete acquisition data, the problem may look like a model-quality problem even though the failure happened much earlier.
9. PostgreSQL is authoritative; Redis is not
Another production decision was making state ownership explicit.
In the deployed architecture:
PostgreSQL 18 is authoritative for persistent application state.
It stores areas such as:
- authentication/session state
- profile resumes
- onboarding preferences
- AI settings
- pipeline runs
- seen-job state
- saved scans
- application actions
- operator decisions
- scheduler history
- job/RAG documents
- discovery state
- metrics
- agent traces and related operational state
Redis 7 is used for acceleration and coordination where appropriate โ caching, invalidation, and locking โ but it is not treated as authoritative application storage.
I wanted the system to remain understandable when caches disappear.
That sounds basic, but clear state ownership prevents a surprising number of production problems.
10. RAG is useful, but it still needs boundaries
ApplyLens also includes retrieval over the job corpus for search and grounded question answering.
The retrieval layer supports things such as:
- job corpus search
- query filtering
- lexical retrieval
- semantic retrieval where configured
- result deduplication/ranking
- grounded answer generation
Owner-facing retrieval is constrained to allowed job identities before evidence reaches the answerer.
An empty allowed set fails closed instead of silently widening into the entire shared corpus.
That is another example of a pattern that appears throughout the project:
Retrieval convenience should not override identity boundaries.
11. Production engineering was most of the work nobody sees
The visible AI features are only one part of the application.
A large amount of development time went into things that are much less impressive in a screenshot:
- retry behavior
- rate limiting
- timeout bounds
- pagination limits
- source-health monitoring
- cache behavior
- deduplication
- owner isolation
- authentication
- role-based access
- registration approval
- persisted run identity
- concurrency controls
- scheduler safety
- process liveness
- artifact tracking
- backups
- rollback procedures
- deployment health checks
The production topology currently looks roughly like this:
Internet
โ
Caddy
โ
FastAPI / Python web application
โ
โโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโ
โ โ โ
โผ โผ โผ
PostgreSQL 18 Redis 7 Run/artifact storage
(authority) (cache/locks)
The application is containerized with Docker Compose.
Production scheduling is handled through systemd, including the six-hour shared acquisition schedule.
Deployment has an explicit backup and rollback path rather than treating git pull && restart as a deployment strategy.
This project started as a scraper.
It ended up teaching me much more about operating an AI system than about scraping itself.
12. The frontend also had to grow with the architecture
ApplyLens is not only a pipeline with a thin HTML wrapper.
The application has a hybrid frontend architecture:
- FastAPI server-rendered surfaces
- classic JavaScript for several workflow-heavy pages
- React 18 + TypeScript + Vite for dashboard/operational surfaces
The product now includes areas for:
- executive overview
- planning
- decisions
- application tracking
- pipeline execution
- optimization review
- tailoring
- scheduler health
- agentic operations
- advanced diagnostics
- profile/resume management
- AI provider/model settings
- saved scans
- admin workflows
Building the UI changed the backend design too.
Once users can inspect a score, save a decision, rerun a pipeline, or review an AI suggestion, concepts like identity, persistence, dirty state, idempotency, and write verification stop being backend implementation details.
They become product behavior.
What makes ApplyLens different for me
I don't think the interesting claim is:
"I built a job app with AI."
There are already many good products in that category.
The part I wanted to explore was the architecture underneath the experience:
Can shared acquisition, deterministic processing, bounded model reasoning, workload-specific routing, evidence-grounded generation, retrieval, agentic orchestration, observability, and human authority coexist in one system without collapsing everything into an LLM call?
ApplyLens became my attempt to answer that question.
And the answer I arrived at is:
Yes โ but only if the boundaries are designed deliberately.
What I learned
A few lessons became much clearer while building this.
1. The biggest AI decision is often where not to use AI
Freshness, explicit exclusions, identity, access control, deterministic scoring boundaries, and state transitions do not become better merely because a language model can participate in them.
2. Expensive intelligence should come after cheap certainty
Filter and deduplicate first.
Use model reasoning when the remaining ambiguity actually benefits from it.
3. Model routing is an application architecture problem
Provider/model selection, qualification, credentials, failure behavior, caching, and fallback rules need explicit ownership.
4. Agentic systems need authority boundaries
A traceable recommendation is not the same thing as permission to mutate production state.
5. Observability is part of AI quality
If you cannot tell whether the source failed, the cache was stale, parsing broke, the provider route changed, or the model produced poor output, every failure starts looking like "the AI is bad."
6. Human-in-the-loop should be an architectural property
It should not be a disclaimer added after autonomous behavior is already designed.
The stack
The main technology stack currently includes:
Python ยท FastAPI ยท PostgreSQL ยท Redis ยท React ยท TypeScript ยท Vite ยท Docker ยท LangGraph ยท LLM/GenAI workflows ยท RAG ยท systemd ยท Caddy
But the most valuable part of this project for me has not been learning another framework.
It has been learning how to decide which component should own which decision.
Explore the project
If you'd like to look deeper:
๐ Access-controlled live deployment: https://applylensjobs.com
sriram-hariharan
/
job-scraper
Production AI job intelligence & application planning platform with multi-source acquisition, LLM evaluation, resume intelligence, RAG, and observability.
ApplyLens AI
AI-powered job discovery, application planning, resume scanning, and tailoring workspace.
Scrape jobs from modern ATS platforms, rank opportunities against saved resumes, review AI optimization guidance, generate tailored drafts, and track the full application workflow from one local operator app
Table of Contents
- What This App Does
- Core Workflows
- Dashboard Pages
- Pipeline Capabilities
- AI Optimize Scan
- Resume Tailoring Workspace
- Scheduler Operations
- Storage and Persistence
- Supported ATS Sources
- Project Structure
- Local Setup
- Running the App
- Running Pipelines from the CLI
- Environment Variables
- Testing and Validation
- Portfolio and Demo Docs
- Deployment Notes
- Roadmap Ideas
What This App Does
ApplyLens AI is a job-search operating system for serious application workflows. It combines job scraping, resume intelligence, AI-assisted review, and decision tracking into a single FastAPI web application.
The app is designed around a practical application loop:
- Discover and scrape jobs from multiple ATS platforms.
- Filter, deduplicate, and rank jobs against your savedโฆ
I'm especially interested in conversations around AI engineering, ML engineering, production LLM systems, model routing, evaluation, RAG, agentic workflows, and human-in-the-loop system design.
If you've built something similar โ or disagree with any of these architectural choices โ I'd genuinely like to hear how you approached it.
AI disclosure: I used ChatGPT to help research, structure, and edit this article. The ApplyLens project, implementation, architecture decisions, and technical claims are my own, and I reviewed the final content against the project repository before publishing.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ full credit and traffic to the original publisher.