pdf-inspector
A layout-aware native parser that returns Markdown with page markers.
markdown · no OCR · MIT
The parser roster
Text extraction, Markdown conversion and hosted OCR solve different problems. Compare the configured adapters and inspect their output before choosing.
A layout-aware native parser that returns Markdown with page markers.
markdown · no OCR · MIT
LlamaIndex's parser, exposed here through its Markdown output and page metadata.
markdown · OCR capable · Apache-2.0
A JVM-based PDF parser with structured output, running in the parsing engine.
markdown · OCR capable · Apache-2.0
A JavaScript PDF-to-Markdown converter available directly in the web application.
markdown · no OCR · MIT
A JavaScript text extractor that provides a plain-text baseline for PDF workflows.
text · no OCR · MIT
A JavaScript PDF parser used here as a text-extraction baseline.
text · no OCR · MIT
A hosted, finance-specialised OCR service for document extraction.
text · OCR capable · LicenseRef-Proprietary