
Digitization
Book Digitization Services: How to Choose a Vendor for Large Backlist Projects
Book Digitization Is Not Just Scanning
A publisher with 500, 5,000 or 50,000 legacy titles does not simply need somebody who owns scanners.
The real requirement is usually much larger:
Physical or legacy source → clean digital master → structured content → searchable/accessible formats → publishing-ready output
A digitization partner may therefore need to handle several connected processes:
- Source assessment
- Scanning
- Image cleanup
- OCR
- OCR correction
- Structural reconstruction
- Tables and footnotes
- Image extraction
- Metadata
- XML
- EPUB
- Searchable PDF
- POD files
- Accessibility requirements
- Quality assurance
Current digitization providers increasingly market workflows that extend from OCR into EPUB, XML and other structured outputs rather than treating OCR as the final deliverable.
Direct Answer: What Should Publishers Look for in Book Digitization Services?
For a large backlist, evaluate a vendor across seven areas:
Source handling → Scanning → OCR → Structure → Output formats → QA → Production capacity
Do not choose primarily on page price or an advertised OCR percentage.
A better procurement test is:
Can this supplier reliably transform representative titles from our actual collection into the digital assets we need downstream?
Ask for a representative pilot containing difficult pages before approving full production.
What Are Book Digitization Services?
Book digitization converts physical or legacy publications into usable digital assets.
Depending on the collection, this may include:
Scanning
Capturing each page as a digital image.
Image Processing
Correcting:
- Skew
- Cropping
- Noise
- Contrast
- Orientation
- Page alignment
OCR
Optical character recognition converts page images into machine-readable text.
OCR Cleanup
Machine recognition is not the same as publication-ready text.
Cleanup may need to correct:
- Character substitutions
- Broken words
- Hyphenation
- Paragraph boundaries
- Headers and footers
- Column order
- Numbers
- Tables
Current OCR-cleanup vendors explicitly identify these as common problems in legacy scanned books.
Structural Conversion
Content may then be reconstructed into:
- XML
- HTML
- EPUB
- Searchable PDF
- Word
- Other structured formats
This is the stage where a collection moves from merely digitized pages toward reusable publishing content.
Scanning vs OCR vs Structured Conversion
These terms should not be treated as synonyms.
Process | What It Produces | Main Purpose |
|---|---|---|
Scanning | Page images | Preserve visual page |
OCR | Machine-readable text | Search/extraction |
OCR cleanup | Corrected text | Improve reliability |
Structural conversion | Semantic content | Reuse and publishing |
EPUB conversion | Digital publication | Ebook distribution |
Accessible conversion | Structured accessible content | Inclusive reading |
POD preparation | Print-ready assets | On-demand printing |
A scanned PDF can look excellent while containing poor OCR.
Likewise, accurate OCR does not automatically provide:
- Heading hierarchy
- Semantic tables
- Lists
- Footnotes
- Navigation
- EPUB structure
- Accessibility
That distinction matters enormously when comparing quotes.
Why OCR Accuracy Percentages Can Mislead Buyers
A vendor may advertise:
99% OCR accuracy
That number sounds excellent.
But imagine a 100,000-word academic title.
Even a seemingly small error rate could leave many recognition errors.
More importantly, a single aggregate percentage tells you little about where errors occur.
Mistakes in:
- ISBNs
- Mathematical expressions
- Dates
- Proper names
- Footnotes
- Tables
- Bibliographies
can matter much more than ordinary prose errors.
Historical fonts, degraded paper, skewed scans, complex columns and multilingual material also make recognition harder.
Recent archival research continues to identify reliability, transparency, bias and accountability as significant OCR/AI concerns rather than treating automated extraction as infallible.
Better question
Instead of asking only:
“What is your OCR accuracy?”
ask:
“How is accuracy measured, sampled, corrected and reported for our content types?”
10 Questions to Ask a Book Digitization Vendor
1. What source material can you handle?
Ask about:
- Bound books
- Loose pages
- Fragile originals
- Existing scanned PDFs
- Image files
- Legacy digital files
The workflow should fit the source rather than force every title through the same process.
2. Is scanning destructive or non-destructive?
For archives, libraries and valuable books, preservation may matter as much as throughput.
Some digitization workflows use overhead/book scanners specifically to avoid unbinding originals.
3. What happens after OCR?
This is one of the most important questions.
Does the vendor deliver:
Raw OCR
or:
Reviewed, structured publishing content?
Ask exactly what correction and structural reconstruction are included.
4. How are complex pages handled?
Your pilot should include:
- Tables
- Footnotes
- Multi-column pages
- Illustrations
- Captions
- Indexes
- Bibliographies
- Special characters
- Equations where applicable
A perfect sample consisting only of plain paragraphs proves very little.
5. What output formats are available?
Do not plan only for today's requirement.
A publisher may eventually need:
- Searchable PDF
- EPUB
- XML
- HTML
- POD
- Accessible PDF
Structured content creates more reuse opportunities than an image-only archive.
6. How is QA performed?
Ask whether the process includes:
- Automated checks
- Human review
- Source comparison
- Sampling
- Validation
- Correction loops
Competitor workflows increasingly advertise combinations of OCR and dedicated human/manual quality review, which means buyers should expect vendors to explain their QA model rather than merely naming an OCR engine.
7. Can the vendor maintain structure?
A book is not simply a sequence of characters.
It contains relationships.
For example:
Chapter → Heading → Paragraph → Figure → Caption → Footnote
Those relationships become particularly important when content is destined for EPUB, XML, accessibility or digital platforms.
8. What production capacity can they sustain?
Ask for:
- Daily capacity
- Peak capacity
- Team structure
- Parallel workflow capability
- Escalation process
- Turnaround assumptions
Large-scale digitization is fundamentally different from processing ten books.
9. How are exceptions managed?
Every serious backlist contains exceptions.
Ask what happens when production encounters:
- Missing pages
- Damaged scans
- Unusual fonts
- Foldouts
- Handwritten annotations
- Complex tables
- Unsupported characters
A mature workflow needs an exception path.
10. Can they run a representative pilot?
Never test only the easiest title.
Choose:
Easy + average + difficult + unusual
That gives procurement teams a much more realistic picture of quality and effort.
Book Digitization Vendor Evaluation Scorecard
A useful procurement scorecard could look like this:
Evaluation Area | Suggested Weight |
|---|---|
Output quality | 25% |
OCR + manual QA process | 20% |
Structural conversion capability | 15% |
Capacity/scalability | 15% |
Required output formats | 10% |
Complex-content handling | 10% |
Communication/reporting | 5% |
Price still matters.
But choosing the lowest per-page quote can become expensive if another team later has to reconstruct:
- Broken OCR
- Tables
- Metadata
- EPUB structure
- Accessibility
Evaluate total downstream cost, not only conversion cost.
When Should a Publisher Convert Directly to EPUB?
If the primary goal is digital distribution, it can be inefficient to think:
Scan → OCR → finished.
Instead:
Scan → OCR → clean structure → EPUB → validation → QA
For older titles, this also creates an opportunity to improve:
- Navigation
- Semantic headings
- Lists
- Tables
- Image alternatives
- Metadata
This becomes especially important when accessibility is part of the future publishing strategy.
Digitization and Accessibility Should Not Be Separate Afterthoughts
One strategic mistake is:
- Digitize a backlist.
- Produce inaccessible digital files.
- Later remediate everything again.
Where project requirements allow, accessibility should be considered when defining the target structure.
For example, good semantic reconstruction can support both:
- EPUB production
- Accessibility
That does not mean OCR automatically creates accessible content.
It means good source structure reduces downstream rework.
A Better Large-Scale Digitization Workflow
Stage 1 — Collection Assessment
Classify:
- Source condition
- Languages
- Layout types
- Complexity
- Required outputs
Stage 2 — Representative Pilot
Test real collection material.
Stage 3 — Capture
Scan and process images.
Stage 4 — OCR
Extract machine-readable content.
Stage 5 — Human Correction
Review recognition against source.
Stage 6 — Structural Conversion
Create the required semantic format.
Stage 7 — Output Production
Generate:
- EPUB
- XML
- POD
- Other required deliverables
Stage 8 — Validation and QA
Validate both technical format and content quality.
Stage 9 — Batch Reporting
Track:
- Received
- In production
- QA
- Rework
- Approved
- Delivered
The workflow becomes:
Assess → Pilot → Capture → OCR → Correct → Structure → Convert → Validate → Deliver
Where Gentize Fits
Gentize has a verified large-scale digitization proof point:
35,000+ title large-scale academic digitization project, delivering scanning, OCR, EPUB, POD, and accessible PDF.
Gentize also has a 1,000-books-per-day digitization capacity and a 200+ member production team.
These are particularly relevant for publishers, libraries and digitization partners evaluating high-volume backlist work because the requirements extend beyond OCR into multiple publishing outputs.
Internal verification before publication: Link the 35,000+ project claim to the exact Gentize case-study/service page currently approved for public use. Also confirm whether the phrase “1,000 books per day” is currently published externally; if it is internal-only, retain it for sales collateral but remove it from the public blog.
Book Digitization Procurement Checklist
Before signing a large project, confirm:
- Source assessment completed
- Representative sample approved
- Scan specification agreed
- OCR methodology documented
- OCR QA defined
- Complex pages tested
- Structural requirements agreed
- Output formats defined
- Metadata requirements documented
- Accessibility requirements documented
- Exception workflow defined
- Capacity validated
- Reporting cadence agreed
- Rework criteria documented
- Final acceptance criteria signed off
A strong statement of work should define what “complete” means before production begins.
Frequently asked questions
What is book digitization?
Book digitization converts physical or legacy publications into digital assets through processes such as scanning, OCR, text correction, structural conversion and digital-format production.
Is OCR the same as book digitization?
No. OCR is one stage. Complete digitization can also include scanning, cleanup, structural tagging, metadata, EPUB/XML conversion and QA.
How do publishers choose a book digitization company?
Evaluate real sample quality, OCR correction, structural capability, output formats, complex-content handling, QA, capacity and reporting—not only price.
Can scanned books be converted to EPUB?
Yes. Scans can be OCR-processed, corrected, structured and converted to EPUB, although complexity and source quality affect the workflow.
Should publishers request a sample first?
For large collections, yes. Use representative material containing both ordinary and difficult pages so quality and workflow can be evaluated before scaling.
Keep reading
Accessibility
PDF/UA-1 vs PDF/UA-2: Which Standard Should Your Accessible PDF Project Target?
Accessibility
Accessible EPUB Backlist Remediation: A Publisher’s Guide to Scaling Legacy Titles
Accessibility
Section 508 Document Remediation Services: A Procurement Guide for Accessible PDFs and Digital Content
Related services
Want this kind of work shipped on your project?
Brief the studio
Got something on your desk that needs this kind of attention?
Tell us the rough outline. We reply within a working day with a scoped response from the practice lead — not a sales person.
