parse_file.py parses one real local document through Harbor’s public parser
interfaces. It verifies file access, automatic registry routing or an explicit
PDF selection, and non-empty extracted content without pytest or test doubles.
Read the shared smoke-test safety and exit-code guidance before using sensitive documents or engines that download models.
uv sync --package harborrag-adapters --extra parsers
For optional PDF profiles and exact backends, install the PDF engines as well:
uv sync --package harborrag-adapters --extra parsers --extra pdf
Some PDF engines require model downloads, native libraries, or suitable CPU/GPU resources. Installation alone does not guarantee that every backend is runnable on the current machine.
parse_file.py does not load dotenv files. If a backend needs credentials,
model-cache controls, or OCR settings, use
env-example/.env.parser.example as a reference
and export the selected variables in the process environment before running the
command.
Pass a real document path and let HarborParser select the parser:
python packages/harborrag-adapters/tests/parsers/smoke/parse_file.py samples/report.docx
python packages/harborrag-adapters/tests/parsers/smoke/parse_file.py samples/data.xlsx
python packages/harborrag-adapters/tests/parsers/smoke/parse_file.py samples/report.pdf
Maintain a representative corpus with at least one DOCX, PPTX, XLSX, CSV, HTML, EPUB, image, Markdown/text, JSON, text PDF, and scanned PDF.
For PDFs, choose either a Harbor profile or one exact backend. The options are mutually exclusive.
python packages/harborrag-adapters/tests/parsers/smoke/parse_file.py \
samples/report.pdf --pdf-profile fast
python packages/harborrag-adapters/tests/parsers/smoke/parse_file.py \
samples/scan.pdf --pdf-profile ocr
python packages/harborrag-adapters/tests/parsers/smoke/parse_file.py \
samples/report.pdf --pdf-backend pymupdf
python packages/harborrag-adapters/tests/parsers/smoke/parse_file.py \
samples/report.pdf --pdf-backend docling
Supported profiles come from PdfParserProfile. Exact backend choices are
docling, liteparse, mineru, paddleocr, and pymupdf. PDF options used
with a non-PDF input return exit code 1.
The check passes only when parsing completes and extracted content is non-empty. It prints only:
It does not print or persist extracted content.
| Code | Meaning |
|---|---|
0 |
Parsing succeeded with non-empty content |
1 |
Selection was invalid, the parser failed, or content was empty |
2 |
The requested real input file does not exist |
parsers extra is installed.pdf extra and check native,
model-cache, memory, CPU/GPU, and network requirements for that backend.