Markdown to HTML Conversion: 7 Common Pitfalls and How to Fix Them
You write documentation in Markdown. Beautiful formatting, code blocks, tables.
You convert to HTML. Result:
- โ Code blocks lose syntax highlighting
- โ Tables misaligned
- โ Special characters become
<gibberish - โ Links with parentheses broken
- โ User input creates XSS vulnerability
The problem: Markdown parsers vary wildly. What works in GitHub doesn't work in your tool.
This guide shows you the exact differences between Markdown flavors and how to convert without breaking formatting, code blocks, or security.
Pitfall 1: Choosing the Wrong Markdown Flavor
The Markdown Fragmentation Problem
Original Markdown (2004):
- By John Gruber
- Ambiguous spec (many edge cases undefined)
- No tables, task lists, or strikethrough
Modern flavors:
| Flavor | Tables | Task Lists | Strikethrough | Fenced Code | Autolinks |
|---|---|---|---|---|---|
| CommonMark | โ | โ | โ | โ | โ |
| GitHub (GFM) | โ | โ | โ | โ | โ |
| Markdown Extra | โ | โ | โ | โ | โ |
| kramdown | โ | โ | โ | โ | โ |
Real-world impact:
Your Markdown has a table. CommonMark parser renders it as plain text (tables not supported). GFM parser renders as proper HTML table โ
How to Choose
Use GitHub Flavored Markdown (GFM) if:
- โ You need tables
- โ You need task lists
- โ You need strikethrough
- โ You want automatic URL linking
Use CommonMark if:
- โ You need strict spec compliance
- โ You're building a parser
Most documentation: Use GFM.
Pitfall 2: Code Blocks Lose Syntax Highlighting
The Problem
Your Markdown contains code with syntax highlighting markers.
Basic HTML conversion:
<pre><code>function hello() {
console.log("Hello");
}
</code></pre>
Problem: No syntax highlighting. Just plain monospace text.
Solution 1: Use a Markdown Parser with Highlighting
Popular libraries:
| Library | Language | Highlighting | GFM Support |
|---|---|---|---|
| marked.js | JavaScript | โ (needs plugin) | โ |
| markdown-it | JavaScript | โ (with plugin) | โ |
| Python-Markdown | Python | โ (with extension) | โ ๏ธ Partial |
| commonmark.js | JavaScript | โ | โ |
Solution 2: Post-Process with Highlight.js
Step 1: Convert Markdown to HTML
import { marked } from 'marked';
// Example: converting markdown with code blocks
const markdown = "# Your markdown here\n" +
"```javascript\n" +
"function hello() {\n" +
" console.log('Hello');\n" +
"}\n" +
"```";
const html = marked.parse(markdown);
Step 2: Apply syntax highlighting
<link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/highlight.js/styles/github.css">
<script src="https://cdn.jsdelivr.net/npm/highlight.js/lib/core.min.js"></script>
<script>
document.querySelectorAll('pre code').forEach((block) => {
hljs.highlightElement(block);
});
</script>
Result: Code blocks with syntax highlighting applied.
Solution 3: Server-Side Highlighting
Using markdown-it + highlight.js (Node.js):
const md = require('markdown-it')({
highlight: function (str, lang) {
if (lang && hljs.getLanguage(lang)) {
return hljs.highlight(str, { language: lang }).value;
}
return '';
}
});
const html = md.render(yourMarkdown);
Benefit: HTML already has highlighting (no client-side JavaScript needed).
Pitfall 3: Tables Render Incorrectly
The Problem
Your Markdown:
| Left | Center | Right |
|:-----|:------:|------:|
| L | C | R |
Expected HTML: Table with proper alignment attributes.
What goes wrong:
- CommonMark parser: Renders as text (no table support)
- Basic GFM parser: No alignment attributes
Solution: Use GFM-Compatible Parser
JavaScript (marked with GFM):
import { marked } from 'marked';
import { gfmHeadingId } from 'marked-gfm-heading-id';
marked.use(gfmHeadingId());
const html = marked.parse(markdown, { gfm: true });
Python (markdown with tables extension):
import markdown
html = markdown.markdown(text, extensions=['tables'])
Pitfall 4: Special Characters Break
The Problem
Your Markdown:
Use <div> and in HTML.
Bad conversion:
Double-escaped: &amp;nbsp;
Expected:
Properly escaped: &nbsp;
Problem: Double-escaping.
Why It Happens
- Markdown parser escapes:
<โ< - You manually escape again:
<โ&lt;
Result: Double-escaped gibberish.
Solution: Let the Parser Handle Escaping
DON'T do this:
// โ Manual escaping before parsing
const escaped = markdown.replace(/</g, '<');
const html = marked.parse(escaped);
DO this:
// โ
Let marked handle it
const html = marked.parse(markdown);
Markdown parsers automatically escape HTML entities in code blocks and inline code.
Pitfall 5: Links with Parentheses Break
The Problem
Your Markdown:
[Wikipedia](https://en.wikipedia.org/wiki/Markdown_(language))
Bad HTML: Link breaks at the first closing parenthesis.
Why: Parser sees ) inside URL as end of link.
Solution: Percent-Encode Parentheses
Option 1: Encode manually
[Wikipedia](https://en.wikipedia.org/wiki/Markdown_%28language%29)
Option 2: Use angle brackets
[Wikipedia](<https://en.wikipedia.org/wiki/Markdown_(language)>)
Angle brackets tell parser to include everything inside.
Pitfall 6: User Input Creates XSS Vulnerability
The Problem
User submits Markdown:
<script>alert('XSS')</script>
Your converter:
const html = marked.parse(userInput);
document.body.innerHTML = html; // ๐ฅ XSS!
Result: Script executes.
Solution 1: Sanitize HTML Output
Use DOMPurify:
import { marked } from 'marked';
import DOMPurify from 'dompurify';
const dirty = marked.parse(userInput);
const clean = DOMPurify.sanitize(dirty);
document.body.innerHTML = clean; // โ
Safe
What DOMPurify does:
- Removes
<script>tags - Removes
onclick,onerrorattributes - Removes
javascript:URLs - Keeps safe tags
Solution 2: Disable Raw HTML in Markdown
Marked.js:
const html = marked.parse(userInput, {
sanitize: true,
});
Result: HTML tags rendered as text (harmless).
Which to Use?
Sanitize (DOMPurify):
- โ Allow safe HTML
- โ Remove dangerous HTML
- Use for: User comments, blog posts
Escape all HTML:
- โ Simpler
- โ Users can't use any HTML
- Use for: Strict Markdown-only
Pitfall 7: Line Breaks Don't Work
The Problem
Your Markdown:
Line 1
Line 2
HTML output:
<p>Line 1 Line 2</p>
Expected:
<p>Line 1<br>Line 2</p>
Why It Happens
Markdown spec: Single line breaks are ignored.
To force a line break:
- End line with two spaces
- Or use a blank line (new paragraph)
Solution 1: Use Two Spaces
End line with two spaces (invisible in editors).
Solution 2: Enable breaks Option (GFM)
Marked.js:
const html = marked.parse(markdown, {
breaks: true,
});
Result: Every newline becomes <br>.
Use with caution: May insert unwanted line breaks.
Conversion Tools Comparison
| Tool | GFM | Highlighting | XSS Protection | Speed |
|---|---|---|---|---|
| marked.js | โ | โ ๏ธ Plugin | โ (use DOMPurify) | โกโก Very fast |
| markdown-it | โ | โ Plugin | โ ๏ธ Partial | โก Fast |
| showdown.js | โ | โ | โ | โก Fast |
| Python-Markdown | โ ๏ธ | โ Extension | โ ๏ธ Partial | โก Fast |
| Pandoc | โ | โ | โ | โ ๏ธ Slower (CLI) |
Recommendation:
- JavaScript: marked.js + DOMPurify + highlight.js
- Python: markdown + bleach
- CLI: Pandoc
Step-by-Step: Safe Markdown to HTML
Goal: Convert user-submitted Markdown to HTML safely (no XSS, syntax highlighting, tables).
Step 1: Install dependencies
npm install marked dompurify highlight.js
Step 2: Convert safely
import { marked } from 'marked';
import DOMPurify from 'dompurify';
import hljs from 'highlight.js';
marked.setOptions({
gfm: true,
breaks: false,
highlight: function(code, lang) {
if (lang && hljs.getLanguage(lang)) {
return hljs.highlight(code, { language: lang }).value;
}
return code;
}
});
function convertMarkdown(userInput) {
const dirty = marked.parse(userInput);
const clean = DOMPurify.sanitize(dirty, {
ALLOWED_TAGS: ['p', 'br', 'strong', 'em', 'code', 'pre',
'a', 'ul', 'ol', 'li', 'table', 'thead',
'tbody', 'tr', 'th', 'td', 'h1', 'h2',
'h3', 'h4', 'h5', 'h6'],
ALLOWED_ATTR: ['href', 'class', 'align']
});
return clean;
}
const html = convertMarkdown(userInput);
document.getElementById('output').innerHTML = html;
Result:
- โ Tables render correctly
- โ Code blocks have syntax highlighting
- โ XSS attempts blocked
- โ Links work
FAQ
Q: Which Markdown flavor should I use?
A: GitHub Flavored Markdown (GFM) for 90% of use cases.
Q: How do I preserve HTML in Markdown?
A: By default, most parsers allow HTML passthrough. To sanitize user input, use DOMPurify.
Q: Why is my code block not highlighted?
A: Your parser may not support syntax highlighting. Use markdown-it with highlight.js plugin.
Q: Can I convert HTML back to Markdown?
A: Yes, using libraries like turndown (JavaScript) or html2text (Python). But conversion is lossy.
Q: What's the fastest Markdown parser?
A: marked.js (JavaScript) and mistune (Python) are among the fastest.
Conclusion
Markdown to HTML conversion is not a one-size-fits-all process. Avoid these 7 pitfalls:
- Choose GFM (not basic Markdown)
- Add syntax highlighting (highlight.js or parser plugin)
- Use GFM-compatible parser for tables
- Let the parser escape special characters
- Encode parentheses in URLs or use angle brackets
- Sanitize user input (DOMPurify) to prevent XSS
- Understand line break rules
Safe conversion workflow:
- Use marked.js with gfm: true
- Add highlight.js for code blocks
- Sanitize output with DOMPurify
- Test with edge cases
Result: Clean, safe, properly-formatted HTML from Markdown.