I Scanned Dozens of Real Backends into OpenAPI. Here's Why the Naive Approaches Fail.
You inherit a backend. Forty endpoints, three frameworks across two services, zero documentation. Someone asks, "Can you just generate the OpenAPI spec?" There are three tempting answers, and all three bite you: Hand
You inherit a backend. Forty endpoints, three frameworks across two services,
zero documentation. Someone asks, "Can you just generate the OpenAPI spec?"
There are three tempting answers, and all three bite you:
- Hand-write it. Accurate for a week, then a field gets renamed and the spec starts lying.
- Ask an LLM to read the repo. It returns a beautiful, confident document with routes that don't exist, fields that were removed last quarter, and response types it hallucinated from a variable name.
-
Grep for routes. A regex that matches
app.get(...)also matchescache.get(...), misses every mounted sub-router, and has no idea what the handler returns.
I spent a lot of time on this problem while building
@powerduck/code-to-openapi,
an MIT-licensed scanner that turns a running codebase into a validated
OpenAPI 3.2 document. This is what "generate OpenAPI from code" actually
requires — and where the shortcuts quietly fail.
1. A route is not a string match
Here is the trap in one snippet:
const value = await cache.get(`session:${id}`); // not a route
app.get("/orders/:id", getOrder); // a route
router.get("/health", () => "ok"); // a route, only if mounted
A regex cannot tell these apart. It either floods your spec with junk paths or
makes you hand-filter the output, which defeats the point.
The deterministic approach is framework instance tracing. The engine
follows the import/export graph to prove that the receiver of .get(...) is
the real express() app or an express.Router() that is actually mounted.
Then:
-
cache.get(...)is ignored because the receiver is not a framework object. - A router that is created but never mounted is reported as unreachable, not emitted as a path.
- Middleware arrays, chained
Router().use()composition, and CommonJSmodule.exports = routerare traced the same way as ESM.
This generalizes. In Go it means following r.GET(...) on the actual
*gin.Engine; in Spring it means resolving @RestController beans; in Laravel
it means following Route::controller()->group() and array-callables. The
framework already knows your routes. The job is to read its registration graph,
not to guess from tokens.
2. The hard part is never the path. It's the contract
Finding POST /orders is maybe 10% of the work. The other 90% is:
- What is the request body shape?
- Which query parameters does it actually read?
- What does it return, nested relations included, and what is nullable?
- Which status codes can it produce?
That information lives behind generics, DTOs, constructors, serializers and
helper functions. A few real examples the engine resolves:
// generics on both sides
app.get("/users", async (req: Request, res: Response<User[]>) => { ... });
// Zod is the contract in many modern stacks
const Body = z.object({ email: z.string().email(), age: z.number().int().optional() });
// Go: the response is built in a constructor in another file
h.JSON(200, service.NewOrderListResponse(orders))
For TypeScript the scanner uses the TypeScript compiler API to resolve
generics (Response<User[]>, Request<Params, ResBody, ReqBody, Query>),
named interfaces, enums, utility types (Partial/Pick/Omit) and Zod
schemas. Named declarations become reusable components.schemas with $refs;
anonymous shapes stay inline. Python, Go, Java, C#, Rust and PHP are parsed
through tree-sitter WASM, so no language toolchain has to be installed.
Crucially, it follows the data across files: constructor return structs in Go,
service calls behind Gin's c.JSON, Spring ResponseEntity/Page envelopes,
Laravel API Resources and transformers.
3. "I don't know" is a feature, not a bug
This is the principle that separates a tool you can trust from a confident
liar. Every parameter, request body and response is classified as exactly one
of:
- proven — with evidence in the source;
- proven absent — the code demonstrably never sets it;
- a gap — explicitly tagged with a machine-readable code.
A route is never emitted as a bare URL with empty contracts. Gap codes include
query-unknown, body-schema-unknown, response-unknown, auth-unknown and
sse-events-unknown.
When I ran this against real open-source backends, the honest gaps were often
the most informative part of the report:
| Backend | Scanned | Operations | What stayed a gap, honestly |
|---|---|---|---|
| Express (RealWorld) | 28 files | 20 | 6 untyped request bodies |
FastAPI (docs_src) |
514 files | 434 | 188 responses with no declared response_model
|
| Spring (real app) | 1,593 files | 432 | 12 WebSocket/SSE event payloads |
| ASP.NET MVC (RealWorld) | 65 files | 19 | 10 responses not statically traced |
A FastAPI handler that returns a raw dict gets response-unknown, not a
schema invented from the function name. An ASP.NET Minimal API using
Results<T> keeps an explicit gap instead of a guessed body. That restraint is
deliberate: a fabricated schema is worse than a flagged one because it looks
finished.
4. Where AI actually helps — and where it must be banned from
Naive "throw the repo at an LLM" generation fails because the model is allowed
to invent routes. But AI is genuinely useful for one narrow job: filling the
specific gaps the deterministic engine could not prove.
The scanner never calls a model vendor itself. It ships a prompt contract and a
strict validator; the host performs the HTTP call. The resolver receives only a
small per-handler slice for routes that actually have gaps — never whole
files — and its answer is clamped to a safe JSON Schema subset (no $ref,
bounded depth and property counts).
import {
scanProject,
buildGapMessages,
parseGapResolution,
gapCacheKey,
GAP_PROMPT_VERSION,
type GapResolver,
} from "@powerduck/code-to-openapi";
const cache = new Map<string, unknown>();
const resolver: GapResolver = {
id: "openai-compatible-host",
async resolve(request) {
const key = gapCacheKey(request, GAP_PROMPT_VERSION);
const hit = cache.get(key);
if (hit) return hit as never;
const res = await fetch("https://your-model-host/v1/chat/completions", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${process.env.MODEL_API_KEY}`,
},
body: JSON.stringify({
model: "your-model",
messages: buildGapMessages(request),
temperature: 0,
response_format: { type: "json_object" },
}),
});
if (!res.ok) return null; // a failed fill is never fatal to the scan
const data = await res.json();
const parsed = parseGapResolution(
data.choices?.[0]?.message?.content ?? "",
); // clamps to the safe subset
if (parsed) cache.set(key, parsed);
return parsed; // null leaves the gap visible, never invents a route
},
};
const result = await scanProject({ root: "./api", gapResolver: resolver });
The model can fill a query parameter, a header, a body field or a status-keyed
response schema. It cannot invent a route, method or path. When the static
engine is certain, the model is never consulted. When it's uncertain, the gap
stays visible and the human stays in control. That is the correct division of
labor: deterministic first, AI only for the residue, with the boundary
explicit.
5. Try it in under a minute
npm install @powerduck/code-to-openapi
npx tsx node_modules/@powerduck/code-to-openapi/examples/basic.ts ./my-api
Or from code:
import { scanProject } from "@powerduck/code-to-openapi";
const result = await scanProject({ root: "/path/to/api" });
console.log(
`${result.report.routesConfirmed} confirmed, ` +
`${result.report.routesPartial} partial routes`,
);
const { document, documentValid } = await result.convert();
console.log("OpenAPI 3.2 valid:", documentValid);
Today there are 28 framework packs across 8 languages — Express, Fastify,
NestJS, Hono, Koa, Next.js and Elysia; FastAPI, Flask, DRF, Starlette and
SQLModel; Gin, Chi, net/http, gorilla/mux, Echo and Fiber; Spring, JAX-RS and
Micronaut; ASP.NET Core and FastEndpoints; Axum, actix and Rocket; Laravel,
Symfony and Slim. SSE endpoints are emitted with a canonical
x-protocol: "sse" extension.
Two details matter for real adoption:
-
Incremental rescans. A
.powerduck/discovery.jsonsidecar fingerprints files and routes, so rescans diff added/changed/removed routes. A three-way merge preserves your manual edits; removed routes are flagged for review, never silently deleted. -
Monorepo leaves. A root with no server probes
packages/*,apps/*andservices/*and aggregates supported leaves into one document.
The bar is "honest," not "impressive"
A generator that always returns a complete-looking document is not smart; it's
just willing to lie. The useful behavior is boring: prove what the code does,
mark what it can't prove, and let AI fill only the marked residue under tight
constraints.
If you've been maintaining an OpenAPI document by hand for an existing
codebase — or worse, trusting a model to summarize it — give the scanner a run
on one service and read the gaps report before you read the pretty document.
The gaps tell you more about your API than the routes do.
- Source and docs: github.com/powerducklab/code-to-openapi
- Or scan a folder visually (with per-route AI gap review) in the free web app: powerduck.com/app
What's the worst case you've seen from an "AI-generated" API spec — a route that
didn't exist, or a field type that was flat-out wrong? Drop it in the comments;
those failure modes are exactly what the completeness gate is built to prevent.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.