Advertisement
PDF Classifier
Instantly detect whether your PDF is text-based, scanned, image-based, or mixed. Get confidence scores and identify which pages need OCR.
๐ Your files never leave your browser
Loading PDF engine...
Drop your PDF here
or click to browse
๐ PDF files only๐ Max 750MB
Processing your PDF...
This happens entirely in your browser
Something went wrong
Error message here
Advertisement
More Free PDF Tools
Frequently Asked Questions
Four types: TextBased (has extractable text), Scanned (images of text), ImageBased (contains only images), and Mixed (combination of text and scanned pages).
The confidence score (0.0 to 1.0) indicates how certain the classification is. Higher scores mean more reliable detection.
These are specific pages in your PDF that contain only images and no extractable text. They would need Optical Character Recognition (OCR) to extract text.
Classification typically completes in 10-50ms, even for documents with hundreds of pages. The engine samples content streams without loading the full document.
Knowing your PDF type helps you choose the right processing pipeline. Text-based PDFs can be extracted directly, while scanned PDFs need OCR โ saving you time and money.