Dev.to AI 🤖 Ai 👁 0 📖 1 min read

DOCX, PPTX, XLSX and EPUB all start with the same magic bytes — auto-detect means asking the ZIP what it is

When I added auto-detect to my universal document converter, I assumed file extensions and Content-Type headers would carry the weight. Both lied to me within the first week. A contract arrived named contract.pdf. The e

When I added auto-detect to my universal document converter, I assumed file extensions and Content-Type headers would carry the weight. Both lied to me within the first week.

A contract arrived named contract.pdf. The extension said PDF, the upstream header said application/pdf. My PDF parser choked on byte zero: the file was actually a DOCX. Someone had renamed an export, and a proxy in the delivery chain had stamped the header based on the filename. Nothing on the outside matched what was inside.

Magic bytes seemed like the obvious fix, and PDFs declare %PDF, so that part worked immediately. Then an uncomfortable fact surfaced: DOCX, PPTX, XLSX, and EPUB are all ZIP archives. They share the exact same PK\x03\x04 signature — one magic number, four formats.

The disambiguation lives inside the container:

  • EPUB is required to store a file literally named mimetype as the first entry, uncompressed, containing application/epub+zip. Two reads and you're done.
  • Office formats carry [Content_Types].xml at the archive root. Look at the declared paths: word/ means DOCX, ppt/ means PPTX, xl/ means XLSX.

Plain HTML and text are the awkward ones — no container to ask, so those fall back to conservative heuristics that fail loudly instead of silently guessing wrong.

Two lessons stuck with me. First: never trust the filename, and only half-trust the header — proxies happily "correct" one from the other, so neither is independent evidence. Second: when formats share a signature, the answer usually isn't a cleverer fingerprint on the raw bytes. It's opening the archive and reading its manifest. The metadata you need was there all along, one layer deeper than where you were looking.

I packaged all of this into one auto-detecting endpoint at https://x402.freeq.one/tools/document_to_markdown.html — bytes or URL in, detected format handled, Markdown out, instead of dispatching between five separate converters.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.