Skip to content

Fix obograph_source prefix-map destruction and predicate mapping - #548

Merged
sierra-moxon merged 2 commits into
masterfrom
improve-obograph-gene-properties
May 28, 2026
Merged

sierra-moxon merged 2 commits into
masterfrom
improve-obograph-gene-properties

Conversation

@kevinschaper

Copy link
Copy Markdown
Collaborator

Summary

Two related bugs in the obograph TSV pipeline that together caused obojson IRIs to land as URIs (instead of CURIEs) in TSV output and well-mapped RO relations like RO:0002162 (in_taxon) to be demoted to biolink:related_to.

  • TsvSource.set_prefix_map and SssomSource.set_prefix_map now use update_prefix_map (additive) rather than set_prefix_map (destructive). The base PrefixManager loads ~600 default JSON-LD prefix mappings in __init__; the transformer's source.set_prefix_map({}) call was wiping them, leaving only the three hardcoded biolink/owlstar/MONARCH entries that set_prefix_map re-adds.
  • ObographSource.read_edge passes the already-computed CURIE (not the IRI) to bmt's get_element_by_mapping. bmt only recognizes CURIE form, so the IRI lookup silently returned None and edges fell through to the catch-all biolink:related_to. Six common RO:* relations get proper biolink slots after this fix (in_taxon, causes, has_participant, disrupts, caused_by, expresses).

Closes #546.
Closes #547.

Test plan

  • Two new unit tests in tests/unit/test_source/test_obograph_source.py (one per bug), backed by a small obograph fixture (tests/resources/obograph_curie_and_predicate.json).
  • Adjusted test_read_obograph1/test_read_jsonl2 expected edge count from 205 to 206. Previously, two distinct relations between the same node pair (BFO:0000050 and RO:0002211 both between GO:0007165 and GO:0008150) were silently demoted to biolink:related_to and produced colliding deterministic edge ids ({sub}-{predicate}-{obj}) that conflated in the test's dict. With BFO:0000050 now correctly mapped to biolink:part_of, all three edges are kept distinct. A code comment in the test explains this.
  • Full unit test suite passes (335 passed, 8 skipped — the skipped ones require external services like neo4j/arango).

Two related bugs in the obograph TSV pipeline that together caused obojson
IRIs to land as URIs (instead of CURIEs) in TSV output and well-mapped RO
relations like RO:0002162 (in_taxon) to be demoted to biolink:related_to.

- TsvSource and SssomSource set_prefix_map now use update_prefix_map
  (additive) rather than set_prefix_map (destructive). The base
  PrefixManager loads ~600 default JSON-LD prefix mappings in __init__;
  the transformer's source.set_prefix_map({}) call was wiping them.
- ObographSource.read_edge passes the already-computed CURIE (not the IRI)
  to bmt's get_element_by_mapping, so RO:* relations with biolink slot
  mappings resolve correctly instead of falling through to related_to.

Adjusted test_read_obograph1/test_read_jsonl2 expected edge count from 205
to 206: previously, two distinct relations between the same node pair
(BFO:0000050 and RO:0002211) were both demoted to biolink:related_to and
produced colliding deterministic edge ids that conflated in the test's dict.
With BFO:0000050 now mapped correctly to biolink:part_of, all three edges
are kept distinct.

Closes #546.
Closes #547.
kevinschaper added a commit to Knowledge-Graph-Hub/kg-phenio that referenced this pull request May 9, 2026
Three small follow-ups surfaced when running the new gene-taxon materializer
against a real phenio.json with a kgx checkout that has biolink/kgx#548
applied (the prefix-map fix that finally lets identifiers.org/{hgnc,ncbigene}/
URIs contract to CURIEs through ObographSource):

- sources.py: add NCBIGene to EDGE_SOURCES and NODE_SOURCES with
  ("ncbi-gene", "Gene"). Previously NCBIGene URIs were always dropped by
  BAD_PREFIXES on the subject side, so the maps never had to know about
  them. With kgx contracting them now, phenio_node_sources.py's
  category_sources/infores_sources lookups would KeyError otherwise.

- phenio_node_sources.py: switch the category_sources / infores_sources
  lookups to .get() with safe defaults. The kgx fix lets several previously-
  filtered prefixes flow through (e.g. MAXO, dct, prov, GOREL, UBERON_CORE);
  unknown ones now keep whatever category kgx assigned and fall back to a
  lower-cased prefix as the infores name rather than crashing the whole
  transform.

- phenio_transform.py: the gene-taxon materializer's subject filter accepts
  both CURIE and URI form (HGNC:, NCBIGene:, http://identifiers.org/{hgnc,
  ncbigene}/) so it works whether the upstream kgx has the contraction fix
  or not. Also always emit in_taxon / in_taxon_label columns even when no
  rows have a value, since phenio_node_sources.yaml declares them and
  Koza checks the header on every run.

Verified end-to-end: PhenioTransform on the current phenio.json plus this
materializer attaches in_taxon to 753 of 769 NCBIGene nodes, with
NCBIGene:698782 (macaque MLH1) coming through as
"biolink:Gene | MLH1 | infores:ncbi-gene | NCBITaxon:9544 | Macaca mulatta".
HGNC nodes pass through cleanly but without taxon (they need the SPARQL
update from monarch-initiative/phenio#130 to ship in a phenio rebuild
before the materializer has anything to fold for them).
With the get_element_by_mapping(curie) fix, BFO:0000050 maps to
biolink:part_of instead of falling through to biolink:related_to.
The recovered predicate distinguishes an edge that previously
collided on edge-id with another related_to edge between the same
nodes, so the post-fix graph correctly preserves one more edge.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants