How We Built Our Own AI Interview Grader (No GPT)
Most "AI interview" tools are a thin wrapper around GPT. We did the opposite — we trained our own generative model, from scratch, to run the whole interview: ask the questions, grade your answers, and explain them. Here'
Most "AI interview" tools are a thin wrapper around GPT. We did the opposite — we trained our own generative model, from scratch, to run the whole interview: ask the questions, grade your answers, and explain them. Here's how it works, what went wrong on the way, and why the harder road was worth it.
When you finish an answer in Peakblick, something has to decide: did that actually cover what the question was testing, or did it just sound confident? Doing that well — on a timer, for every answer, consistently — is the whole product. Everything else is a frontend. So we spent our time on the thing that's actually hard: the model that runs the conversation and judges what you say.
Why not just call GPT?
It's the obvious shortcut, and plenty of "AI interview" products take it: forward your answer to a big API, ask it to score, print the result. We didn't, for three reasons.
Consistency. A grader built on someone else's API quietly changes every time that model is updated. The same answer can score a 6 one week and an 8 the next, and you have no way to know why. For a tool whose entire job is to give you a trustworthy read on where you stand, that drift is disqualifying.
Cost. Grading every answer of every mock interview through a frontier API doesn't scale to something people use daily without the price landing back on the user. We wanted practice to be cheap enough to do constantly.
Independence. We wanted a system we own, can measure end to end, and can improve on our own schedule — not rent from a provider who can change the terms, the price, or the behaviour overnight. A side effect we like: your answers are handled entirely by our own model and never shipped off to a third party.
So we trained one ourselves — a compact, from-scratch generative model, with no distillation from GPT or any other external model. Every improvement is ours, and so was every problem. That constraint made what came next much harder, and much more interesting.
One generative model, the whole interview
Peakblick runs on a single model we built and trained. It is fully generative — the same model that generates the questions also generates a model answer and writes your feedback. There's no external API anywhere in the loop, and nothing hard-coded from a question bank: it produces the interview the way a good interviewer improvises one, shaped by the role, seniority and topics you picked.
That matters because a real interview isn't a quiz with fixed answers. A strong senior backend question looks nothing like a junior one on the same topic, and a good interviewer reads your answer and decides what's worth probing next. We wanted a model that behaves like that — not one that reads lines off a card.
Generating the questions
The first job is asking well. Given a role and level, the model plans the ground to cover and then generates each question fresh — phrased naturally, pitched at the right difficulty, and varied so you're not drilling the same three prompts every session. Generating questions is the part a generative model is genuinely great at, and it's where the "no fixed bank" promise pays off: the space of questions is effectively unlimited, so practice stays useful past the first few runs.
The hard part: judging without hallucinating
Asking is the easy half. The hard half is judgement — and it's where most of our work went.
A compact model asked to grade cold can do exactly what a nervous candidate does: sound fluent and confident while being wrong. Early versions of our grader had a specific, dangerous failure mode — faced with an answer that was actually incorrect, the model would sometimes write a confident, well-structured explanation of why it was right. It talked itself into approving bad answers. And that's the one failure a grader can't have: tell someone a wrong answer was good, and you've done worse than not grading at all. You've taught them the wrong thing right before the real interview.
Getting that reliable — not just good on average, but trustworthy on the answers that matter — was the real engineering problem behind Peakblick.
Our approach: let the model answer first
The fix that made judgement reliable is simple, and it plays directly to what a generative model is best at. Before it ever looks at your answer, the model writes its own ideal answer to that exact question — what a strong response should actually cover, generated fresh for this question rather than pulled from a list.
Only then does it read your answer, and measure it against that reference: point by point, what you covered, what you touched lightly, what you missed, and where you said something that contradicts a key idea. Grounding the grade in the model's own generated answer — instead of asking it to pass judgement out of thin air — is what turned a shaky grader into one we trust. The model is doing what it's good at (writing a strong answer, then comparing) instead of conjuring a verdict from nothing. It's the same reason a human grader with the model solution in front of them is fairer than one marking from memory.
A score you can read, not a mystery number
Because the grade is built up point by point, the final 1–10 isn't a vibe — it's the sum of concrete judgements you can see. A 6 means something specific: these points landed, this one was thin, that one was missing. That's what lets the feedback be useful instead of generic. "Good effort, keep practising" helps no one; "you explained the mechanism but missed the one trade-off the interviewer was really after" is something you can actually fix before Monday.
Feedback that's specific, not generic
Feedback is only worth anything if it names the real gap. Ours is grounded — tied to what the question was actually testing and to the model answer it laid out — so it tells you exactly where you fell short and why it mattered, instead of hand-waving. It never pads with invented criticism or "good effort, keep practising" filler.
Every line points at something concrete you can fix: the trade-off you skipped, the edge case you didn't mention, the term you reached for that wasn't quite right. That specificity is the whole difference between feedback you nod along to and feedback you actually act on — and acting on it, before the real interview, is the entire point.
Teaching it the whole stack
A grader is only as good as what it knows. A model can't fairly judge an answer about database indexes, or Git internals, or how a framework handles requests, if it doesn't deeply understand those things itself. So a large and ongoing part of the work isn't the model's wiring at all — it's its education.
We train it on a broad, carefully-sourced body of real technical material across the stack: multiple languages, the frameworks people actually ship with, databases and SQL, system design, version control, and the fundamentals underneath all of it. As we widen that knowledge, the grading gets fairer and holds up across more of the topics a real interview might wander into. The architecture stays the same; the knowledge keeps growing — and that's a dial we can keep turning for as long as the product lives.
What this means — and what's next
The headline isn't "we use AI." Everyone does. It's that the thing running your interview is a generative model we built, own, measure, and keep improving — not a wrapper whose behaviour we can't see or steer. That ownership is why we can promise consistency, keep practice cheap, and keep raising the ceiling instead of waiting for someone else's roadmap.
Next, we're broadening both ends: more topics and seniorities on the question side, and deeper knowledge on the grading side — including interviews that dig into real code. The goal is simple and a little ambitious: an interviewer in your pocket that's good enough to make the real one feel easy.
Try it on your own answers. Peakblick asks you real interview questions on a timer and scores every answer 1–10 with specific feedback — so you find the weak spots before the real thing. Free to try, no card.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.