PDF where text can be found and copied, often via OCR.
May contain original text or OCR text layers.
Key outcome for scanned-document modernization.
Use this concept when optimizing, converting, validating, or troubleshooting document workflows.
Optical Character Recognition that turns scan images into machine-readable text.
TXT is plain text without embedded rich formatting semantics.
PDF variant combining multiple structural and compatibility characteristics.
Machine-readable text content embedded alongside scanned page imagery.
Mapping from glyphs to Unicode code points, often via ToUnicode data in PDFs.