ENES
pdf-mergeEngineering Guide

Merge PDF Files Without Losing Bookmarks and Hyperlinks: 4 Tools Compared

AC
Alex ChenΒ·Lead Systems Architect
Published on 2026-08-31Β·8 min readΒ·Daily Toolbox Engineering

Merge PDF Files Without Losing Bookmarks and Hyperlinks: 4 Tools Compared

You have 5 PDF chapters. Each has:

  • Table of contents with bookmarks
  • Internal links (click "see Section 3.2" β†’ jumps to page 15)
  • Hyperlinks to external resources

You merge them with an online tool. Result:

  • ❌ All bookmarks gone (flat 200-page document, no navigation)
  • ❌ Internal links broken (click "Section 3.2" β†’ nothing happens)
  • ❌ File size ballooned from 5MB β†’ 12MB

The problem: Most "merge PDF" tools simply concatenate pages without preserving document structure.

This guide shows you how to properly merge PDFs while keeping bookmarks, hyperlinks, form fields, and metadata intact.


Why PDF Merging Breaks Things

What Gets Lost

Common casualties:

Element Lost by Most Tools Why It Matters
Bookmarks (TOC) βœ… Often lost Navigation in long documents
Internal links βœ… Often broken Cross-references between sections
External hyperlinks ⚠️ Sometimes broken Citations, resource links
Form fields βœ… Often lost Interactive forms
Metadata ⚠️ Partially kept Author, title, keywords (SEO)
Layers (OCR) ⚠️ Sometimes flattened Searchable text layer

Example:

  • PDF 1: Pages 1-50, bookmark "Chapter 1" β†’ page 10
  • PDF 2: Pages 1-30, bookmark "Chapter 2" β†’ page 5

Naive merge:

  • Combined PDF: 80 pages
  • PDF 1 bookmarks: Still point to page 10 βœ…
  • PDF 2 bookmarks: Still point to page 5 ❌ (should be page 55)

Smart merge:

  • PDF 2 bookmarks: Adjusted to point to page 55 (50 + 5) βœ…

Method 1: pdftk (Best for Preserving Structure)

What is pdftk?

PDF Toolkit (pdftk):

  • Command-line tool
  • Free and open-source
  • Best at preserving bookmarks, links, forms
  • Fast (processes 1000 pages in seconds)

Basic Merge

Merge 3 PDFs:

pdftk file1.pdf file2.pdf file3.pdf cat output merged.pdf

What it preserves: βœ… Bookmarks (auto-adjusted page numbers)
βœ… Internal links (recalculated)
βœ… External hyperlinks
βœ… Form fields
βœ… Metadata (from first file)

What it doesn't do: ❌ Compress (output may be larger than inputs)
❌ Optimize (no deduplication of embedded fonts/images)


Advanced: Selective Pages

Merge pages 1-10 from file1, pages 5-15 from file2:

pdftk A=file1.pdf B=file2.pdf cat A1-10 B5-15 output merged.pdf

Merge all files in a folder:

pdftk *.pdf cat output merged.pdf

Preserving Metadata

By default, pdftk keeps metadata from the first file.

Merge with custom metadata:

# Step 1: Merge
pdftk file1.pdf file2.pdf cat output temp.pdf

# Step 2: Extract metadata template
pdftk temp.pdf dump_data output metadata.txt

# Step 3: Edit metadata.txt (change title, author, etc.)

# Step 4: Update metadata
pdftk temp.pdf update_info metadata.txt output final.pdf

Method 2: Ghostscript (Best for Compression)

When to Use Ghostscript

Use Ghostscript if:

  • βœ… You need to compress while merging (large scanned PDFs)
  • βœ… You want to standardize format (all PDFs β†’ PDF/A)
  • ⚠️ You can tolerate minor bookmark/link issues

Don't use Ghostscript if:

  • ❌ You need perfect bookmark preservation (use pdftk)
  • ❌ You need form fields intact (Ghostscript flattens forms)

Basic Merge with Compression

gs -sDEVICE=pdfwrite -dCompatibilityLevel=1.4    -dPDFSETTINGS=/ebook    -dNOPAUSE -dQUIET -dBATCH    -sOutputFile=merged.pdf file1.pdf file2.pdf file3.pdf

Result:

  • 3 PDFs merged
  • Images downsampled to 150 DPI
  • File size: Typically 60-90% smaller than naive merge

Bookmarks: ⚠️ Partially preserved (simple bookmarks work, complex nested structures may break)


Merge Without Compression

gs -sDEVICE=pdfwrite -dNOPAUSE -dQUIET -dBATCH    -sOutputFile=merged.pdf file1.pdf file2.pdf

(Omit -dPDFSETTINGS to avoid downsampling)


Method 3: PyPDF2 / pikepdf (Python, Programmable)

Best for Automation

Use Python libraries if:

  • βœ… You need to merge 100+ PDFs (batch processing)
  • βœ… You need conditional logic (merge only PDFs with "draft" in filename)
  • βœ… You're building a web app (merge user-uploaded PDFs)

PyPDF2 Example

from PyPDF2 import PdfMerger

merger = PdfMerger()

# Add PDFs
merger.append('file1.pdf')
merger.append('file2.pdf')
merger.append('file3.pdf')

# Merge
merger.write('merged.pdf')
merger.close()

What it preserves: βœ… Bookmarks (with manual page offset adjustment)
βœ… Hyperlinks
⚠️ Form fields (basic only, complex forms may break)


pikepdf Example (Better Preservation)

import pikepdf

pdf = pikepdf.Pdf.new()

# Merge PDFs
for file in ['file1.pdf', 'file2.pdf', 'file3.pdf']:
    src = pikepdf.open(file)
    pdf.pages.extend(src.pages)

pdf.save('merged.pdf')

pikepdf advantages:

  • Better bookmark handling than PyPDF2
  • Faster
  • Actively maintained

Method 4: Browser-Based (Privacy-Safe, No Install)

Client-Side PDF Merging

How it works:

  1. Select multiple PDF files
  2. Browser loads them into memory
  3. JavaScript merges using PDF.js library
  4. Download merged file

Advantages: βœ… No upload (files never leave your device)
βœ… No installation
βœ… Works offline (after first load)
βœ… Cross-platform

Limitations: ⚠️ Large PDFs (100+ MB) may crash browser
⚠️ Advanced features (form fields, signatures) may not preserve


When to Use

Good for:

  • Quick merges (2-5 files, < 50MB total)
  • Sensitive documents (no cloud upload)
  • Devices where you can't install software

Not good for:

  • Large batches (100+ files)
  • Files with complex forms/signatures
  • Very large files (browser memory limits)

Comparison Table

Tool Bookmarks Links Forms Compression Speed Privacy
pdftk βœ… Perfect βœ… Yes βœ… Yes ❌ No ⚑ Fast βœ… Local
Ghostscript ⚠️ Partial βœ… Yes ❌ Flattens βœ… Yes ⚑ Fast βœ… Local
PyPDF2 ⚠️ Manual βœ… Yes ⚠️ Basic ❌ No ⚑ Fast βœ… Local
pikepdf βœ… Good βœ… Yes βœ… Good ❌ No ⚑⚑ Very fast βœ… Local
Browser-based ⚠️ Varies βœ… Yes ⚠️ Basic ❌ No ⚠️ Slow (large files) βœ… Client-side
Adobe Acrobat βœ… Perfect βœ… Yes βœ… Yes βœ… Yes ⚑ Fast βœ… Local ($$$)

Recommendation:

  • Best overall: pdftk (free, perfect preservation)
  • Best for compression: Ghostscript (if bookmarks aren't critical)
  • Best for automation: pikepdf (Python)
  • Best for privacy + no install: Browser-based tool

Real-World Use Cases

Case 1: Merge Contract Chapters

Scenario:

  • 5 chapters (10-20 pages each)
  • Each has bookmarks (Chapter 1, Section 1.1, etc.)
  • Internal references ("see Section 3.2 on page 45")
  • External links to legal citations

Best tool: pdftk

Command:

pdftk chapter1.pdf chapter2.pdf chapter3.pdf chapter4.pdf chapter5.pdf   cat output full_contract.pdf

Result:

  • All bookmarks preserved and adjusted
  • Internal links recalculated ("Section 3.2" now points to correct page in merged doc)
  • External links intact
  • File size: Sum of inputs (no compression)

Case 2: Merge Scanned Documents

Scenario:

  • 20 scanned invoices (each 2-5 pages, 600 DPI)
  • No bookmarks
  • No internal links
  • Total size: 150MB

Best tool: Ghostscript (merge + compress)

Command:

gs -sDEVICE=pdfwrite -dPDFSETTINGS=/ebook    -dNOPAUSE -dQUIET -dBATCH    -sOutputFile=all_invoices.pdf invoice*.pdf

Result:

  • All invoices merged
  • Downsampled to 150 DPI
  • File size: 150MB β†’ 12MB (92% reduction)
  • No bookmarks to lose (scans didn't have them)

Case 3: Merge User-Uploaded PDFs (Web App)

Scenario:

  • Users upload multiple PDFs
  • Server must merge them
  • No manual intervention

Best tool: pikepdf (Python)

Code:

import pikepdf

def merge_pdfs(file_list, output_path):
    pdf = pikepdf.Pdf.new()
    
    for file in file_list:
        src = pikepdf.open(file)
        pdf.pages.extend(src.pages)
    
    pdf.save(output_path)

# Usage
merge_pdfs(['upload1.pdf', 'upload2.pdf'], 'result.pdf')

Common Mistakes

Mistake 1: Using Naive Concatenation

Problem:

# DON'T DO THIS
cat file1.pdf file2.pdf > merged.pdf

Why it breaks:

  • PDFs aren't plain text
  • Binary concatenation corrupts structure
  • Result: Unreadable file

Correct approach: Use PDF-aware tools (pdftk, Ghostscript, libraries).


Mistake 2: Ignoring Page Number Offsets

Problem:

  • PDF 1: Bookmark "Chapter 1" β†’ page 5
  • PDF 2: Bookmark "Chapter 2" β†’ page 3
  • Naive merge: Both bookmarks still point to their original pages (broken)

Solution: Tools like pdftk automatically adjust. If using Python, manually calculate offsets:

# After merging, adjust PDF 2 bookmarks:
# Original page 3 β†’ new page (50 + 3) = 53

Mistake 3: Merging Different PDF Versions

Problem:

  • PDF 1: PDF 1.4 (old)
  • PDF 2: PDF 2.0 (new, with advanced features)
  • Merged PDF: Downgraded to 1.4, features lost

Solution: Specify compatibility level:

gs -dCompatibilityLevel=1.7 ...

Or use pdftk (preserves highest version).


Step-by-Step: Merge 10 PDFs with Bookmarks

Goal: Merge 10 chapters, preserve all bookmarks, keep file size reasonable.

Step 1: Install pdftk

# Ubuntu/Debian
sudo apt install pdftk

# macOS
brew install pdftk-java

# Windows
# Download from: https://www.pdflabs.com/tools/pdftk-the-pdf-toolkit/

Step 2: Merge

pdftk chapter*.pdf cat output book.pdf

(Assumes files named chapter1.pdf, chapter2.pdf, etc.)

Step 3: Verify bookmarks Open book.pdf in a PDF reader (Adobe, Evince, Preview). Check:

  • βœ… All bookmarks present
  • βœ… Clicking bookmarks jumps to correct pages

Step 4: Compress (if needed) If file is too large:

gs -sDEVICE=pdfwrite -dPDFSETTINGS=/ebook    -dNOPAUSE -dQUIET -dBATCH    -sOutputFile=book_compressed.pdf book.pdf

Result:

  • 10 chapters merged
  • All bookmarks working
  • File compressed to 150 DPI (good for screen reading)

FAQ

Q: Will merging PDFs reduce quality?
A: Not by default (pdftk, PyPDF2). Only if you use Ghostscript with compression settings.

Q: Can I merge password-protected PDFs?
A: Yes, but you must provide the password:

pdftk protected.pdf input_pw PASSWORD cat output unlocked.pdf

Q: Do merged PDFs keep form fields editable?
A: Yes with pdftk and pikepdf. No with Ghostscript (flattens forms).

Q: Can I reorder pages while merging?
A: Yes:

pdftk A=file1.pdf B=file2.pdf cat B A output merged.pdf

(Merges file2 first, then file1)

Q: Why is my merged PDF larger than the sum of inputs?
A: Possible causes:

  • Duplicate embedded fonts (each PDF embeds the same font)
  • Duplicate images (same logo on every page of every PDF)
  • No compression

Solution: Use Ghostscript to deduplicate and compress.


Conclusion

Merging PDFs isn't just appending pagesβ€”it's preserving document structure.

Quick reference:

Your Need Best Tool Command
Preserve everything pdftk pdftk *.pdf cat output merged.pdf
Merge + compress Ghostscript gs -sDEVICE=pdfwrite -dPDFSETTINGS=/ebook ... *.pdf
Automation pikepdf (Python script)
No install Browser tool (upload files)

The safest workflow:

  1. Use pdftk to merge (preserves everything)
  2. If file too large, use Ghostscript to compress afterward
  3. Verify bookmarks and links in result

Result: Merged PDFs with intact navigation, cross-references, and metadata.

Start with pdftk for 99% of use cases.

#pdf-merge#pdftk#combine-pdf#bookmarks
AC
Written by Alex ChenLead Architect

Alex Chen is a distributed systems engineer and core maintainer at Daily Toolbox with over 10 years of experience in client-side web technologies, RFC standards compliance, and cryptographic protocols. He specializes in zero-knowledge client architectures and WebAssembly-accelerated algorithms.

Try the free tools mentioned above

⚑ Open PDF Merger β†’