HarborRAG

Extending HarborRAG

Choose the package that owns the behavior, implement against its public contract, and test the real implementation with deterministic fake dependencies.

Connectors

Add providers under packages/harborrag-adapters/src/harborrag_adapters/connectors/<provider>/.

from collections.abc import Iterator

from harborrag_adapters.connectors.base import BaseConnector
from harborrag_adapters.connectors.schemas import ConnectorQuery
from harborrag_core.domain.raw_document import RawDocument
from harborrag_core.domain.source import SourceRecord


class ExampleConnector(BaseConnector):
    provider_name = "example"

    def discover(self, query: ConnectorQuery | None = None) -> Iterator[SourceRecord]:
        ...

    def load(self, record: SourceRecord) -> RawDocument:
        ...

A connector contribution should include typed config, client/session lifecycle, filtering, pagination, retry/rate-limit behavior, size/collection caps, permission/provenance mapping, public exports, registry metadata, and tests for failures and security boundaries. Register one ConnectorProviderDefinition containing the connector class, typed config factory, aliases, constructor dependency names, config path fields, and default document kind. Runtime configuration and construction derive their behavior from this definition; do not add parallel provider-name switches.

Use fakes in tests; do not ship a provider mock.py as a substitute for production behavior. Add an opt-in smoke script only when a real-system check adds value.

Parsers

Add a complete HarborParser family under parsers/<family>/. Add an individual provider below that family’s engines/<provider>/ directory and implement the family-specific engine contract, such as HarborPDFEngine.

Parsers should:

Register complete families through HarborParserFactory; provider engines do not belong in the root MIME/extension registry. Add provider metadata to the runtime parser provider table when it should be YAML-configurable.

Model providers and behavior

Model code lives under models/chat, models/embed, models/rerank, and provider runtime capabilities under models/runtime. Core request, response, and error shapes live in the corresponding harborrag_core schema modules.

When adding provider support:

Do not emulate embeddings or reranking through chat. Keep embedding fallbacks compatible in space and dimension.

Repositories

Add a backend below the matching family:

repositories/vector/<provider>/
repositories/graph/<provider>/
repositories/cache/<provider>/
repositories/object_store/<provider>/
repositories/database/<provider>/
repositories/state/<provider>/

Implement the family’s Harbor contract and plugin/config pattern. Repository requirements include:

Use repositories/, not a new stores/ family. Do not return raw provider responses by default.

Engine stages

Put provider-independent RAG orchestration in harborrag-engine:

Inject connector/parser/model/repository contracts. Production stages should preserve provenance and permissions, thread tenant/request context, expose bounded concurrency, and make partial progress observable. Do not import a concrete provider subpackage into an engine stage.

Runtime services

Put configuration loading, provider composition, checkpoint coordination, and durable workflow implementation in harborrag-runtime. Reuse the core job, repository, lifecycle, and observation ports instead of adding runtime-owned copies.

The existing connector/parser catalogs demonstrate strict versioning and environment references. A unified composition must retain explicit construction and avoid importing optional providers until selected.

For source-specific canonical behavior, implement ConnectorDocumentTransform inside the provider’s connector package and set document_transform_factory on its ConnectorProviderDefinition. Runtime discovers registered transforms and stays provider-neutral. Use SourceDocumentNormalizerBuilder only for application-local overrides that do not belong to a connector package. For source-specific chunk behavior, provide a ChunkingConfig source profile and an additional ChunkStrategy when building IngestionRuntimeBuilder. Keep provider validation on the strategy’s optional record_validator hook instead of branching in the shared chunk validator.

Temporal SDK integration belongs in harborrag-runtime.temporal. It must not become a core, adapter, or engine dependency.

Application and MCP surfaces

CLI and HTTP code should call BaseAppService; MCP tools should call runtime/service interfaces. Neither surface should construct raw provider clients in handlers.

Stored chat prompts belong under harborrag_runtime/chat/prompts/templates/. To add a public prompt, add its UTF-8 Markdown template, add a stable name and filename mapping to ChatPrompt/PromptCatalog, and test selection through every transport that exposes it. Keep model selection, credentials, endpoints, tenant data, and runtime interpolation out of templates.

Production interfaces also need stable schemas, exit/error mapping, identity and tenant context, permission enforcement, capability budgets, safe observability, lifecycle handling, and audit recording. Update user documentation only after the command, route, transport, or tool is actually wired.

Public exports

Add an export to the harborrag meta-package only when it is stable, implemented, tested, and documented. Package-local public exports should likewise be intentional and included in import smoke tests.

Before a pull request

uv run make lint
uv run make typecheck
uv run make deps-check
uv run make compile
uv run make coverage

See Testing and CONTRIBUTING.md.