HarborRAG

Connector smoke checks

These scripts perform real connector discovery and load operations through HarborConnector, using the same declarative sources as the application: config/connectors.yaml for connector settings and config/parsers.yaml for attachment/document parsing. They verify authentication, source scoping, API/filesystem access, mapping into SourceRecord, and real content parsing. Confluence and JIRA also repeat the load with attachment processing enabled; Local parses the discovered file directly (PDF, DOCX, images, and everything else HarborParser supports).

They are manual checks, not pytest tests. Read the shared smoke-test safety and exit-code guidance before using real credentials or content.

Prerequisites

The base connector clients are installed with the adapter:

uv sync --package harborrag-adapters

Real parsing needs the parser dependencies, plus PDF/OCR engines. These scripts parse PDFs with Docling and images with RapidOCR by default:

uv sync --package harborrag-adapters --extra parsers --extra pdf

The pdf extra explicitly installs both rapidocr and its default CPU inference runtime, onnxruntime. Docling acceleration (CPU/CUDA/MPS/XPU) is configured through config/parsers.yaml.

Configuration

Connector settings live in config/connectors.yaml; credentials live in env/.env.connector. Both fall back to their .example counterparts (config/connectors.example.yaml, env-example/.env.connector.example) when the real file doesn’t exist yet, so copy the example and fill in your values:

cp config/connectors.example.yaml config/connectors.yaml
cp env-example/.env.connector.example env/.env.connector

Each key under connectors is an application-level connection_id, not a provider name. For example, harborrag-workspace may use the local provider and jira-main may use the jira provider. Smoke commands select these IDs so multiple connections can be configured for the same provider.

Connector Required environment variables Optional
Local LOCAL_SOURCE_PATH None
GitHub Environment variables referenced by the selected connection, normally GITHUB_REPOSITORY_URL, GITHUB_TOKEN Provider settings such as branch/root path stay in YAML
Confluence CONFLUENCE_BASE_URL, CONFLUENCE_SPACE_KEY, CONFLUENCE_TOKEN CONFLUENCE_EMAIL for Cloud
JIRA JIRA_BASE_URL and either JIRA_TOKEN or JIRA_API_TOKEN JIRA_EMAIL for Cloud
SharePoint Environment variables referenced by the selected connection, normally Microsoft tenant/client credentials and SHAREPOINT_SITE_URL Drive/root path stay in YAML

Everything else - content type filters, attachment limits, pagination, JIRA project_keys scoping, and so on - is a literal setting in config/connectors.yaml. Edit that file directly instead of adding more environment variables. Relative LOCAL_SOURCE_PATH values are resolved from the repository root.

GitHub and SharePoint now use the same catalog resolution and environment references as every other smoke check.

Run a connector

python packages/harborrag-adapters/tests/connectors/smoke/run.py --connector harborrag-workspace
python packages/harborrag-adapters/tests/connectors/smoke/run.py --connector confluence-main
python packages/harborrag-adapters/tests/connectors/smoke/run.py --connector jira-main

For convenience, a provider name such as --connector jira is accepted when exactly one enabled connection uses it. Pass the connection ID when multiple connections use the same provider. The provider modules remain thin, no-argument entry points; they also require a unique enabled connection for their provider.

After individual checks pass, run every enabled configured connection:

python packages/harborrag-adapters/tests/connectors/smoke/run_all.py

run_all.py dispatches by each definition’s provider, fails when any enabled connection fails, and returns 2 when no enabled connection has a smoke runner.

Save parsed output

By default nothing is written to disk. Pass --output txt or --output md to save the parsed content under tests/connectors/smoke/output/ (override with --output-dir):

python packages/harborrag-adapters/tests/connectors/smoke/run.py --connector jira-main --output txt
python packages/harborrag-adapters/tests/connectors/smoke/run.py --connector jira-main --output md

Saved output is currently supported by Local, Confluence, and JIRA runners; GitHub and SharePoint reject --output explicitly.

txt saves a flat concatenation of the body and every parsed attachment’s text. md saves a structured Markdown document instead: a # title, a short header list (source, content type), a ## Metadata section with a curated set of provider-specific fields (Jira: issue key, status, assignee, priority, labels, …; Confluence: space, version, author, labels, breadcrumb, …; Local: parser, page count, OCR settings, figures extracted, warnings - whichever fields have a value), the body, and each parsed attachment under its own ### heading.

md output also makes images actually viewable:

txt output has no such folder - it’s OCR text only, since plain text can’t reference a file.

For Confluence/JIRA this covers the page/issue body plus every parsed attachment’s text. For Local it saves the real parsed content of the discovered file (not raw bytes). Use --limit to change how many records are discovered and processed - each gets its own output file (default: 3).

What each check verifies

Target Discovery limit Required result
Local 5 (3 via run.py default) At least one record and a successful real parse of the first file
GitHub 3 At least one repository file and a successful blob load
Confluence 3 First page loads without attachments; also loads with attachments if include_attachments: true in config
JIRA 3 First issue loads without attachments; also loads with attachments if include_attachments: true in config
SharePoint 3 At least one drive item and a successful first-file download

The attachment pass only runs when the connector’s include_attachments setting in config/connectors.yaml is true; when it’s false, the check prints a skip message and passes without touching attachments. When the pass does run, Confluence and JIRA fail if an attempted attachment ends in failed or unsupported. A source with no attachments can still pass.

Local fails only on a genuine parse error, or on a parser returning empty content for a file that itself has non-blank bytes. A source file that is itself empty or whitespace-only still passes, matching Confluence/JIRA’s tolerance of blank page/issue bodies.

Parser selection

PDF and image parsing come from config/parsers.yaml (falling back to config/parsers.example.yaml), the same catalog the application uses. The shipped default enables pdf-docling, which parses PDFs with Docling and OCRs scanned pages with RapidOCR. Plain image attachments and local image files always OCR through RapidOCR - that routing isn’t expressible in the declarative parser catalog, so the smoke bootstrap wires it directly.

HARBOR_SMOKE_PDF_BACKEND=docling \
  python packages/harborrag-adapters/tests/connectors/smoke/confluence.py

Docling defaults to auto, asks Docling’s accelerator resolver for the best available device, and prints both the requested and resolved values before the smoke check. Override it with auto, cpu, cuda, cuda:N, mps, or xpu:

HARBOR_SMOKE_PDF_BACKEND=docling \
HARBOR_SMOKE_DOCLING_DEVICE=xpu \
  python packages/harborrag-adapters/tests/connectors/smoke/confluence.py

CUDA and XPU require a matching accelerator-enabled PyTorch build; MPS requires supported Apple hardware. Keep auto for portable configuration and CPU fallback.

Set HARBOR_SMOKE_IMAGE_BACKEND=rapidocr to OCR image attachments. Selecting Docling as the PDF backend also selects RapidOCR for images unless explicitly overridden. On first use the smoke helper reports the ONNX Runtime providers it can see and reuses one loaded RapidOCR engine for all attachments.

Output and troubleshooting

Successful output includes discovered IDs, media types, character counts, and attachment status/count information. Full provider content is not printed unless HARBOR_SMOKE_VERBOSE=1 is set (bounded, redacted previews; disabled in CI).