Skip to content

Parquet output from the ingestion tooling (tracks linkml-map#325) #21

Description

@amc-corey-cox

The metadata index is published as Parquet on object storage (see ARCHITECTURE.md, #12), but the tooling that produces it emits DuckDB native database files. Parquet output does not exist yet.

Tracking here because it blocks our storage design, even though the work happens upstream:

The reason we need Parquet rather than native files is version coupling. DuckDB's native storage format is backward compatible but only best-effort forward compatible, and our producer (Python ingestion tooling) and consumer (service tier) upgrade independently. A producer running a newer DuckDB can write an artifact the consumer cannot open, and it fails at the consumer's deploy rather than at build. Parquet has no such coupling, and it keeps the index readable by the wider ecosystem rather than by one engine.

Until this lands, the Parquet-on-S3 read pattern in #14 has to be exercised against files produced some other way — hand-converted, or written directly with DuckDB's COPY ... TO ... (FORMAT PARQUET). That is fine for the spike and shouldn't block it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Metadata Source of TruthLinkML metadata, semantic bindings, query engineTrackingTracking issues for development and reporting

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions