To convert a physical book to an ebook when no digital files survive, you rebuild the book in four stages: scan every page, run OCR software to turn the page images into editable text, proofread that text against the printed copy, and build a new ebook file from the corrected text. A scanned PDF on its own is not an ebook.
- Do You Have the Right to Republish the Book?
- How Is the Book Scanned, and Will You Get It Back?
- What Does OCR Do, and Which Tools Can You Use?
- How Accurate Is OCR, and Why Does the Text Need Proofreading?
- How Is the Corrected Text Turned Into a Real Ebook?
- How Do You Avoid Ever Needing This Rescue?
- Frequently Asked Questions
Do You Have the Right to Republish the Book?
A book can only be republished as an ebook by someone who holds the ebook rights, so the paperwork comes before any scanning or conversion work. Authors usually need a print-to-ebook rescue for one of two reasons. Some published with a conventional publisher years ago, and the book has since gone out of print. Others self-published, but the files are gone: a computer died, a backup failed, or the company that produced the book closed down. In the second case the rights were always yours, and you can move straight on to scanning.
For a conventionally published book, “out of print” does not automatically return your rights. The publishing contract controls what happens, usually through a reversion clause. The Authors Guild describes the typical clause as having four parts: a definition of “out of print”, a written notice from the author, a period for the publisher to respond, and reversion if the publisher does nothing. Get the reversion confirmed in writing and keep that document, because retailers can ask for proof that you hold the rights.
Rights revert; files usually do not. The publisher may no longer hold usable production files, and few contracts oblige them to hand anything over. Check the extras too: the old cover art, photographs, maps, or a foreword by another writer may have been licensed separately, so they need clearing or replacing before the new edition uses them.
How Is the Book Scanned, and Will You Get It Back?
A printed book is scanned in one of two ways. Destructive scanning slices the spine off, so the loose pages feed through a document scanner like ordinary sheets of paper. Non-destructive scanning keeps the book intact and photographs it page by page, usually with an overhead camera and a cradle that supports the binding.

Destructive scanning is typically cheaper and produces cleaner images, because flat loose pages have no page curve and no shadow near the spine. The trade-off is obvious: the book does not survive. Non-destructive scanning costs more and needs software correction for curve and shadow, but you get the copy back. With an out-of-print title, the copy you send may be one of very few left. If you want it returned, tell the vendor up front and get that confirmed before work starts; do not assume the book comes back by default.
As a guide to cost, US vendor prices checked in August 2026 ran from around $0.19 to $0.32 per page for destructive scanning and from around $0.26 to $0.77 per page for non-destructive work, sometimes with a base fee on top. Treat those as illustrations rather than a market rate, because quotes vary with page count, condition, and resolution. For resolution, aim for about 300 DPI (dots per inch, a measure of scan detail) in grayscale (plain shades of grey rather than colour) for ordinary text; OCR software makers recommend 300 DPI as the practical baseline. Scanning at home with a flatbed scanner or a phone scanning app is possible too. Whichever route you take, keep the master page images, so later stages can be redone without scanning the book twice.
Need help with the technical aspects of self-publishing? Get a free quote.
Thank you. Your enquiry is on our desk.
We've received your message and will get back to you by email, usually within one working day. If it's urgent, you can also reach us via the contact page.
What Does OCR Do, and Which Tools Can You Use?
OCR, short for optical character recognition, is software that reads a picture of a page and produces editable text from it. The scan gives you photographs of your pages. OCR gives you words you can correct, restyle, and rebuild. Every print-to-ebook rescue passes through this stage, because no retailer sells page photographs as an ebook.
The main tools are familiar ones. Adobe Acrobat’s paid plans include OCR and can export the recognised text to Word. ABBYY FineReader is a dedicated OCR application, listed at $99 a year in the US (£84 in the UK) for the Standard edition in August 2026. Tesseract is free and open source and supports more than 100 languages, but it has no point-and-click interface, so it suits readers comfortable with technical setup. Phone apps such as Adobe Scan capture and OCR pages in one step, though the free tier limits OCR to 25 pages per file, which matters for a full-length book.
For a clean modern paperback, the honest difference between these tools is smaller than the difference in what comes next. Whichever engine reads the pages, the output is a draft, and the cleanup effort is where the real time goes.
How Accurate Is OCR, and Why Does the Text Need Proofreading?
Good OCR software recognises roughly 96–99% of characters correctly on clean printed text; that range comes from ABBYY’s own engine documentation, and it is a vendor figure for favourable conditions, not a guarantee. The arithmetic is what catches authors out. A typical paperback page holds around 1,500 characters. At 99% accuracy, up to 15 of them can be wrong; at 96%, up to 60. Across a 250-page book, even a good result leaves thousands of characters to check.
OCR errors also follow patterns rather than falling randomly. Letter pairs with similar shapes swap: “rn” becomes “m”, so “modern” turns into “modem”. Capital I, lowercase l, and the digit 1 trade places, as do the letter O and the digit 0. Words that print hyphenated at the end of a line keep the hyphen in the middle of a sentence. And the running heads and page numbers (the book title and numbering printed at the top or foot of every print page) get read as ordinary text and dropped into the middle of paragraphs.
Many of these mistakes form real words, and “modem” sails through a spellcheck. Because of that, the OCR draft needs a full proofread against the printed copy, page by page, not a spelling pass. Budget more time for this stage than for the scanning and OCR combined; it is the slowest part of the rescue and the one that decides whether the finished ebook reads like the original book.
How Is the Corrected Text Turned Into a Real Ebook?
An ebook is not a picture of a page. It is reflowable text: text that re-wraps itself to fit any screen size and any font size the reader chooses. Once the recognised text has been proofread against the printed book, it is built into a new EPUB file (EPUB is the standard ebook format retailers accept), exactly as if the book were being produced for the first time. The page scans were the rescue vehicle, not the product; nothing from them survives into the finished ebook except the corrected words and any images the new edition keeps.
Print furniture (the headers, footers, and page numbers that repeat on every printed page) goes first. Reflowable text has no fixed pages, so there is nothing for that furniture to attach to. KDP’s formatting guidance says page numbers, headers, and footers do not apply to reflowable ebooks, and IngramSpark’s ebook specifications bar page-number references anywhere in the file, including the contents list. Manual line breaks left over from the print layout go too, along with the blank left-hand pages print books sometimes need. We have covered why running headers and page numbers exist only in print in more detail.
Then the structure is rebuilt rather than copied. Chapter headings are tagged as headings, a linked table of contents replaces the page-numbered print one, and notes become tappable links. EPUB is the format every major platform accepts; IngramSpark, for example, takes reflowable EPUB 2 or 3 files up to 100MB. Before upload, the file should be checked and previewed; we have a separate guide to validating your EPUB before uploading it to a retailer.
How Do You Avoid Ever Needing This Rescue?
Keep your files. Whenever you pay someone to produce your book (a formatter, a cover designer, a full-service company), the handover at the end of the job should include the final output files (the print-ready PDF and the ebook file) and, where the service supplies them, the application files: the working Word or InDesign documents those outputs were built from. Ask for both at the start, so it is part of the agreement rather than a favour later.
Then store them in two places, such as a cloud account plus a local drive. Dead computers, failed migrations, and providers that go out of business are exactly how book files disappear. If your book was conventionally published and the rights revert, ask the publisher for the final PDF at reversion time. Some cannot or will not share it, but it costs nothing to ask, and a digital PDF, while it still needs the same rebuild, skips the scanning stage and most of the correction work.
Rescue jobs of the kind this article describes are rare; at ebookpbook we might see one a year. But each one starts the same way, as a filing problem rather than a publishing problem, and the fix costs nothing at the moment the book is first produced.
If the files are already gone, all is not lost. Confirm the rights, scan the pages, let OCR produce the draft, proofread it against the printed copy, and have the ebook built fresh from the corrected text. The result is not a replica of the old paperback; it is a new, retail-ready edition of a book that was out of reach.
Frequently Asked Questions
Can you sell a scanned PDF of your book as an ebook?
No. A scan is a set of page photographs, and ebook retailers need reflowable text that adapts to each reader’s screen. We have covered why converting a PDF straight to EPUB produces broken ebooks; a scanned PDF has the same problem with more force, because it contains no text at all until OCR runs.
Can you reuse the paperback’s ISBN for the ebook?
No. An ISBN (the 13-digit number that identifies one specific edition of a book) can only be used for one format. KDP, IngramSpark, and Draft2Digital all apply the same rule, so the ebook needs its own identifier even though the text is the same.
Do you have to buy an ISBN for the new ebook?
Not every platform requires one. KDP assigns its own catalogue number (an ASIN) instead of requiring an ISBN, while IngramSpark requires an ebook ISBN because it supplies many stores and library systems at once. IngramSpark and Draft2Digital offer free ISBNs that list the platform rather than you as the publisher of record. A purchased ISBN cost about US$125, £93 in the UK, and A$44 in Australia in August 2026.
Can you convert a physical book to an ebook for free?
You can convert a physical book to an ebook almost entirely with free tools, if your time is free. A phone scanning app captures the pages, Tesseract turns them into text at no cost, and any word processor handles the correction. What free tools do not remove is the labour: the page-by-page proofread and the ebook build still have to be done, by you or by someone you hire.
What happens to photos and illustrations inside the book?
Printed photos and illustrations can be rescanned, but a scan of a printed photo is a copy of a copy and loses quality, so use the original digital images wherever they survive. Check the rights as well: images licensed from external rights-holders (a photographer, illustrator, or archive) may have been cleared for print only, and the ebook edition needs its own permission.