My HTML-to-Markdown converter broke paragraphs on newlines — HTML whitespace rules aren't Markdown rules
While building my HTML-to-Markdown API, I hit a bug that embarrassed me precisely because it looked harmless. I ingested a 40-page frontend tutorial into my own RAG test pipeline, and retrieval kept returning sentence fr
While building my HTML-to-Markdown API, I hit a bug that embarrassed me precisely because it looked harmless. I ingested a 40-page frontend tutorial into my own RAG test pipeline, and retrieval kept returning sentence fragments. The output looked fine at a glance — until I diffed it against the browser.
The source HTML was formatted for humans:
html
A component is a reusable piece of UI.
My converter preserved the newlines, producing Markdown that any HTML purist would call correct. But Markdown isn't HTML. CommonMark treats a single newline as a soft break, yet plenty of GFM-flavored renderers, chat clients, and text extractors treat it as a hard break. Worse for me: my chunker split on newlines, so that paragraph became two chunks — "A component is a reusable" went into one retrieval bucket and "piece of UI." into another. Sentences chopped mid-thought, then ranked individually as if they were complete facts.
The HTML spec says any run of whitespace between block elements collapses to a single space. Browsers have done this since the nineties; my token-to-AST pass never did. The fix wasn't a global reflow at the end — that mangles <pre> blocks. It was collapsing whitespace at the inline-token level, before block structure is decided, and only inside paragraphs, list items, and headings. Fenced code content keeps its bytes exactly as they arrived.
Measured after the fix: about 35% of paragraphs in that tutorial corpus contained mid-paragraph source newlines, and the fragment-as-top-hit cases in my eval set dropped to zero. One edge case actually got slightly worse — definition lists with line-significant formatting — which I now document as a known tradeoff rather than hide.
Lesson: when bridging two formats, map the whitespace semantics, not just the syntax. I ended up packaging the corrected pipeline as https://x402.freeq.one/tools/html_to_markdown.html, and that tutorial corpus is now the first regression test I run after any tokenizer change.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.