Skip to content
DevToolKit

PDF Repair

Repair corrupted or damaged PDF files online — rebuild broken cross-reference tables and trailers with qpdf. 100% client-side processing, no uploads.

pdf

Drop your damaged PDF here, or click to browse

Repairs run entirely in your browser — the file is never uploaded

Processed locally
Was this tool helpful?

How to Use

Repair a corrupted or damaged PDF in four steps:

  1. Upload the damaged PDF -- Drag and drop the file that refuses to open, or click the dropzone to browse. The tool verifies the %PDF- signature before accepting the file. Nothing is uploaded to any server.
  2. Choose a repair level -- "Standard rebuild" rewrites the document and reconstructs its cross-reference table and trailer, which resolves the majority of corruption cases. "Aggressive recovery" additionally ignores broken xref streams, unpacks object streams, and preserves raw stream bytes to salvage more content from severely damaged files.
  3. Review the diagnostic log -- Click "Repair PDF" to run a structural health check followed by the rebuild. The monospace log shows exactly what the engine found: the PDF version, damage warnings such as missing xref tables or malformed objects, and a verification pass on the repaired output.
  4. Download the repaired file -- The summary reports how many issues were found and fixed. If the primary engine cannot recover the file, a fallback engine performs a last-resort extraction and the result is clearly marked as a partial recovery.

The entire pipeline runs in your browser using a WebAssembly build of qpdf, the industry-standard PDF structural library. Your document never leaves your device, which makes this tool safe for confidential contracts, medical records, and legal filings.

About This Tool

A PDF file is more than a sequence of pages. At the end of every PDF sits a cross-reference table -- an index mapping each indirect object to its exact byte offset -- followed by a trailer dictionary and a startxref pointer telling the reader where that index begins. This machinery lets viewers jump directly to any object without scanning the whole file, but it is also the most fragile part of the format. A truncated download, a failed email attachment, a bad USB copy, or a buggy export tool can leave the offsets wrong, the trailer missing, or the pointer pointing at garbage. Most viewers then refuse to open the file at all, even though every page of content inside is perfectly intact.

This tool repairs that structural damage by rebuilding rather than patching. It first runs a diagnostic pass equivalent to qpdf --check, which reports the PDF version, encryption status, and every structural problem it encounters -- damaged xref entries, unexpected bytes where an object should be, missing startxref markers. Those raw diagnostics appear in the repair log so you can see precisely what was wrong with the file.

The repair pass then rewrites the document. When qpdf opens a damaged file it walks the byte stream looking for every recognizable indirect object, discards the broken index entirely, and constructs a brand-new cross-reference table and trailer from the recovered objects. Because the output is a freshly serialized document rather than a patched original, the result is often cleaner than the input even when no corruption was present -- incremental-update bloat and malformed object numbering are normalized away.

Aggressive recovery goes further. The --ignore-xref-streams flag tells the engine to disregard cross-reference streams entirely -- the most compact and most corruption-prone indexing format -- and reconstruct object locations by scanning. Object streams are unpacked so every object is individually addressable, and stream data is preserved verbatim so a partially damaged image or font stream does not abort the whole recovery. When even that fails, a second engine based on pdf-lib performs a last-resort extraction of whatever objects still parse, and the result is honestly flagged as a partial recovery.

Finally, a verification pass re-checks the repaired output and counts remaining problems, so the "issues fixed" number in the report is measured, not assumed. A document that reports zero remaining issues has a structurally valid xref, trailer, and object graph that any standards-compliant viewer can open.

Why Use This Tool

PDF corruption is common and usually unrecoverable by simply re-opening the file. This tool addresses the situations where re-downloading or re-exporting is not an option:

  • Interrupted downloads and transfers -- A PDF that stopped at 95% of a download is missing its trailer entirely. Rebuilding the xref from the surviving objects typically restores the full document.
  • Email and scanner corruption -- Scanners, multifunction printers, and email gateways routinely emit PDFs with malformed xref entries that strict viewers reject but tolerant engines can rebuild.
  • Forensic and archival work -- Recovered files from failing drives or old backups often have damaged structure but intact content. The diagnostic log doubles as a damage report for documentation purposes.
  • Buggy PDF generators -- Some tools write non-compliant files that strict validators flag. A normalize-and-rebuild pass produces a clean file that downstream processors accept.
  • Privacy-sensitive documents -- Online repair services require uploading the file to a server. Here, repair runs in WebAssembly inside your browser -- appropriate for contracts, medical records, tax documents, and anything you cannot send to a third party.
  • No cost or signup -- Desktop repair utilities and commercial recovery services charge for what is fundamentally a structural rebuild. This tool is free with no file-size limits imposed by a server.

Not every file can be saved -- if the page content itself is destroyed rather than merely unindexed, no tool can reconstruct it. When recovery is only partial, this tool says so plainly instead of silently handing back an incomplete file. Related tools: PDF Compress shrinks a healthy file after repair. PDF Unlock removes passwords from files that open but are restricted. PDF Sanitize strips scripts and hidden data from recovered documents. PDF Extract Pages pulls surviving pages into a fresh document. PDF Flatten bakes form fields and annotations into static page content.

FAQ

What kinds of PDF damage can this tool fix?
It repairs structural corruption: broken or missing cross-reference (xref) tables, damaged trailers, bad object offsets, missing startxref markers, and malformed object headers — the most common reasons a PDF refuses to open.
How does PDF repair work?
The tool runs a qpdf diagnostic check, then rewrites the document. During the write, qpdf scans for every valid object, discards the corrupted xref table, and builds a fresh one plus a new trailer — producing a clean, openable PDF.
What is the difference between standard and aggressive recovery?
Standard rebuild rewrites the file and reconstructs xref/trailer data. Aggressive recovery additionally ignores xref streams, disables object streams, and preserves raw stream bytes to salvage more content from severely damaged files.
Will I lose any content from a repaired PDF?
A successful qpdf rebuild preserves all recoverable pages, fonts, and images. If only the fallback engine succeeds, the result is flagged as partial recovery — some objects or metadata may be missing.
Is my PDF uploaded to a server?
No. All diagnostics and repair run locally in your browser via WebAssembly. Your file never leaves your device, so it is safe for confidential documents.