Dev.to AI 🤖 Ai 👁 0 📖 2 min read

A contract.pdf that was secretly a DOCX broke my file-type auto-detection

When my PDF, DOCX, XLSX, PPTX and EPUB converters ran as separate endpoints, agents kept picking the wrong one. A universal "send anything, we'll figure out the type" endpoint was the obvious fix. My first detection logi

When my PDF, DOCX, XLSX, PPTX and EPUB converters ran as separate endpoints, agents kept picking the wrong one. A universal "send anything, we'll figure out the type" endpoint was the obvious fix. My first detection logic: file extension first, Content-Type header as backup. It survived about a day of real traffic.

The file that broke it: a contract named invoice.pdf. My PDF parser opened it and found no %PDF header at all — it was a DOCX somebody had renamed during a mail-merge export. And I couldn't lean on names even in general, because a large share of my uploads arrive as raw base64 with no filename, and most clients stamp anything ambiguous as application/octet-stream. Extensions and headers are rumors. Bytes are facts.

So I rebuilt detection from the bytes up:

  • Starts with %PDF- — easy, it's a PDF.
  • Starts with PK — it's a zip container, unzip and peek inside. A word/document.xml means DOCX, xl/workbook.xml means XLSX, ppt/presentation.xml means PPTX, and a mimetype entry containing application/epub+zip means EPUB. Four of my formats turned out to be zip files; stopping at "it's a zip" would have collapsed four converters into one confidently wrong answer.
  • Otherwise heuristics: a BOM, a doctype, tag density suggests HTML; mostly printable bytes suggests plain text.

Extensions and MIME types got demoted to tiebreakers, not evidence.

Honest residuals: an HTML file stripped of its tags still fools my printable-ratio check occasionally, and sniffing is a 99% solved problem, not 100%. That's also why the per-format converters still exist — sometimes you know the answer better than the bytes do.

The lesson generalizes past file parsing: any pipeline that keys identity off names or headers will eventually meet a renamed file. The first few bytes can't lie, because nobody thinks to edit them.

I ended up packaging this as the Universal Document-to-Markdown API at https://x402.freeq.one/tools/document_to_markdown.html — the endpoint itself is mostly boring glue over those sniffing rules, which is where all the interesting bugs lived.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.