Skip to main content

Reuse: Surveying and Aligning Existing Ontologies

📍 Where we are: Part II · Reuse — the second phase of the lifecycle, governed by the NeOn reuse scenario and the reuse-first authoring discipline of LOT. The specification is written and the running example is loadable. Before we draw a single new class, we go shopping: we survey what the world already publishes, decide what to borrow, and write the borrowing down as one file.

The cheapest class in an ontology is the one you never had to author. The specification fixed what the model must answer; the reuse phase fixes how much of the answer already exists. For a CHO (Chinese Hamster Ovary) cell line — the industry-standard host for making monoclonal antibodies, the targeted protein drugs grown in living cells — the answer is: almost all of it. The Chinese hamster the line is built from is a node on the public tree of life; the IgG complex the cells secrete (the antibody molecule itself), the Protein A capture step that pulls it out of the broth (the liquid the cells grow in), the chromatography medium that step runs on (the packed material that separates the antibody from everything else), and the "process parameter" a feed rate is an instance of are each a published class somewhere with a stable IRI (an Internationalized Resource Identifier — a global web address, like a URL, that names exactly one thing so two systems can confirm they mean the same entity) and a community maintaining it. The physical meaning of each manufacturing step named here — Protein A capture, viral inactivation and filtration, polishing, UF/DF (ultrafiltration/diafiltration), and fill-finish — is walked end to end in Book 1's bioprocessing overview; here we reuse only their names and ontology layers. Reuse-first is not frugality for its own sake — it is the only way a private campaign graph becomes interoperable with the rest of biomedicine and manufacturing the moment it is loaded, rather than after a retrofit nobody budgets for [4].

The simple version

Before building a kitchen, you do not forge your own screws. You buy standard screws, standard pipe fittings, a standard electrical socket — because the dishwasher you install next year expects exactly those, and a homemade socket means a homemade dishwasher forever. Reusing an ontology is buying standard parts: the screw is NCBITaxon for "Chinese hamster," the pipe fitting is iof:CaptureStep for "Protein A capture." This chapter is the trip to the hardware store — surveying the aisles, choosing the parts, and writing the parts list — done before any custom welding begins.

Start from the questions

Reuse is the book's one non-functional requirement (a quality the whole model must have, not a single question it must answer) made operational: the ORSD (Ontology Requirements Specification Document) lists reuse-first — "align to public ontologies, mint only what is genuinely local" — as a quality the model must have regardless of any single competency question. But it also directly underwrites one functional question. CQ-20 asks what host organism, by stable NCBI Taxon IRI, does the line express its product in? — and the only honest answer for a CHO line is Cricetulus griseus, named not by the ambiguous three-letter string "CHO" but by obo:NCBITaxon_10029. That question is unanswerable unless the model has already reused NCBI Taxonomy: you cannot return a stable taxon IRI you never borrowed. And the question matters because the host is the deepest fact in the genealogy — the cell line is alive, mutates, and drifts, and naming its species by a stable identifier shared with the public databases is the first defence against a misidentified root propagating, with full confidence, to every batch beneath it. Reuse is therefore not a stylistic preference; it is a precondition for a specific safety-critical test passing, which is exactly why the lifecycle puts it before conceptualization rather than treating it as polish.

Conceptualize the reuse: a layered survey, grounded in one batch

Ontologies stack, and they stack for a manufacturing reason: a single batch of antibody is not one kind of thing but several, and each kind is best described by a different layer. Take BATCH-2026-001. The batch material — the cells, the broth, and later the purified protein — is one thing; the production bioreactor that held it is a different thing, because equipment persists across many batches and must not be fused into the material; the fermentation that grew the cells is a third thing that both participate in; the 98.611 % monomer purity (the fraction of the product that is intact single antibody molecules rather than clumped-together aggregate) is a fourth, a quality that inheres in — belongs to and cannot exist apart from — one specific material lot and nowhere else; the vessel's standing as "this campaign's production reactor" is a fifth, a role it bears and could shed (the same steel tank is still the same tank if it is reassigned, so its production-reactor role is a separate thing from the vessel that holds it); and the master batch record that prescribed the whole run is a sixth, copyable information distinct from any printout. Six entities, one batch — and the NeOn reuse scenario tells us to look for reusable resources at every level that family of entities spans, because borrowing high up buys the most compatibility per import. Surveying the campaign top-down yields four layers.

  • Foundational. One upper ontology — a domain-neutral vocabulary of the most general kinds of thing (object, quality, process, role) that every field's terms hang under — grounds everything: BFO, the Basic Formal Ontology, standardized as ISO/IEC 21838-2 [1]. It contributes no domain terms — only material entity, quality, process, role — but it is precisely what lets the six things in that one batch each find the right kind of home. The vessel and the broth are both material entities (obo:BFO_0000040) yet stay separate individuals; the fermentation is a process (obo:BFO_0000015) the model never confuses with the material it outputs; the purity is a quality (obo:BFO_0000019) that cannot be accidentally asserted of the empty tank; the production-reactor standing is a role (obo:BFO_0000023). This continuant/occurrent split — things that persist versus things that happen — is the spine the upper-spine chapter built [2], and it is what buys batch traceability: because the vessel and the material are distinct continuants, the graph can later say a second batch ran in the same BR-101 without contradiction; because the fermentation is an occurrent both participate in, "when did this happen?" has somewhere to attach.
  • Mid-level. The IOF Core, published by the Industrial Ontologies Foundry, sits between BFO and the plant: material artifact, manufacturing process, piece of equipment, requirement specification [5]. This is the rung that makes a bp:Batch an IOF material artifact and a fermentation an IOF manufacturing process — the generic manufacturing vocabulary a 98.611 %-purity lot needs before any antibody-specific term is drawn. IOF was modeled deliberately on the OBO Foundry, the life-sciences community that proved coordinated, principle-based ontology building works at scale [3].
  • Domain. Two domain families meet here. The IOF biopharma modules supply the manufacturing interior the campaign runs through: bioreactor, chromatography column and medium, membrane filter, and the unit-operation chain itself — capture, viral-clearance, polishing, UF/DF (ultrafiltration/diafiltration), and formulation (each physical step is built in Book 1 — see capture, viral filtration, UF/DF, and formulation and fill-finish) — plus the cell line and cloned cell line, and the QbD (Quality by Design) parameter/range classes that link a feed rate to a purity. The OBO members supply the biology and the science under those operations: the Gene Ontology (GO) for the IgG molecular function [6], the Protein Ontology (PRO) for the antibody's target [7], NCBI Taxonomy for the Chinese-hamster host [9], the Cell Line Ontology (CLO) for the CHO line [10], a Human Disease Ontology for the indication the product concept names [8], OBI for assays and investigations [11], ChEBI for the formulation excipients, NCIt for the product and mechanism-of-action concepts, and IAO for information artifacts such as the master batch record. And the viral-clearance rung is not one step but two orthogonal barriers the IOF terms keep distinct: a low-pH-hold inactivation (iof:ViralInactivation, validated LRV 4.5 — a log-reduction value, where each log is a tenfold reduction in virus; see viral filtration — enveloped-virus inactivation) and a virus-retentive nanofiltration (iof:ViralFiltration, validated LRV 4.2, size-based removal), whose log-reduction values sum to a claimed total clearance of 8.7 logs. Borrowing two different IOF process classes — rather than collapsing both into "viral step" — is what lets the graph add only across genuinely orthogonal mechanisms (the two steps must remove virus by independent means for their logs to be summed), exactly the additivity rule ICH Q5A — the international regulatory guideline on viral safety of biologics — imposes (the sum is checked by CQ-12 in the validation chapter).
  • Cross-cutting. Orthogonal vocabularies that any domain reuses (here "orthogonal" means independent of the manufacturing domain — they describe units, provenance, and metadata that any field needs): QUDT for units (so the purity, the pH of a viral-inactivation hold, and a fill volume never travel as bare numbers), PROV-O for activity provenance, SOSA for the bioreactor sensors, SKOS for definitions and mappings, Dublin Core for metadata, and the Allotrope Foundation Ontologies (AFO) for the chromatography devices and the analytical results that decide release.

In Turtle — the plain-text notation for writing down these graphs — each @prefix line just gives a short nickname for a long IRI base, so a term can be written obo:NCBITaxon_10029 instead of the full web address:

# align.ttl — the four reuse layers, each a real prefix bound to a verified host.
@prefix obo: <http://purl.obolibrary.org/obo/> . # BFO, RO, OBI, GO, CLO, NCIT, NCBITaxon, PR, DOID, IAO, PATO
@prefix iof: <https://spec.industrialontologies.org/ontology/construct/> . # IOF Core (classes mint under /construct/)
@prefix af-p: <http://purl.allotrope.org/ontologies/process#> .
@prefix qudt: <http://qudt.org/schema/qudt/> .
@prefix prov: <http://www.w3.org/ns/prov#> .
@prefix sosa: <http://www.w3.org/ns/sosa/> .
@prefix dct: <http://purl.org/dc/terms/> .

That single prefix block is the survey's result: every external world the model touches, named once, before a domain class is drawn — and each prefix earns its place because one of the six entities inside a real batch needed it.

How candidates are found: the registries, and what they cannot tell you

A term you cannot locate you cannot reuse, so the survey is a search across the field's registries. Each ontology in the layers above was found, not guessed, and the finding leaves a trail of where to look:

  • OLS4 — the EBI (European Bioinformatics Institute) Ontology Lookup Service — is the front door for the OBO (Open Biological and Biomedical Ontologies) world. Every BFO, RO, OBI, GO, CLO, NCIT, NCBITaxon, PR, DOID, IAO, and PATO IRI in this model was confirmed there, against the live term page — including the obo:NCBITaxon_10029 that names the Chinese hamster and the obo:CLO_0002421 "CHO cell" that names the line.
  • The published IOF release is the front door for IOF, because OLS does not host IOF. The IOF Core and biopharma terms — the bioreactor, the capture step, the chromatography medium, the drug-product formulation process — were verified against github.com/iofoundry/ontology at Release_202602; both the term IRIs and the module ontology IRIs are live and dereferenceable.
  • BioPortal and OBO Foundry round out the OBO search — OBO Foundry's principles page tells you which ontology is the canonical home for a given kind of thing (the target protein → PRO, the indication → a disease ontology) so two antibody programs reach for the same one [3].
  • LOV (Linked Open Vocabularies) and the publishers' own sites cover the cross-cutting layer: QUDT at qudt.org, PROV-O and SOSA at the W3C, AFO at the Allotrope release.
  • The unit registry has a second job here: it is where values shed their bare-number ambiguity at the hardware boundary. QUDT (and the UCUM code system it carries) is not surveyed only so the purity reads 98.611 %; it is surveyed because the *_to_rdf.py loaders read the UCUM (Unified Code for Units of Measure) string straight off the wire — the unit slot a controller exposes via OPC UA (the standard industrial protocol for live plant data), its EUInformation field, and the UnitOfMeasure element of B2MML (the XML the manufacturing-execution system exchanges) — and turn it into qudt:ucumCode / qudt:hasUnit on a qudt:QuantityValue. The reuse decision is what makes a pH or a feed rate cross from the PLC (the floor controller running the equipment) into the graph already wearing its unit, instead of arriving as a float that some later script has to re-annotate by guesswork (the mechanism the from-the-wire-to-the-graph chapter traces end to end).

Crucially, the survey's verdict is reproducible without trusting the prose. The three files the offline checker loads — bioproc.ttl, align.ttl, and instances.ttl — are plain Turtle a reader opens in Protégé (the standard open-source ontology editor) or parses with RDFLib (the Python RDF library the knowledge-graph chapter builds the whole graph on); the per-CQ acceptance harness is validate.py reading cq-catalog.json, the executable ORSD in which each competency question is a PASS/FAIL test (requirements == tests). Because that harness reasons only over the three local files and fetches nothing external, anyone can re-run the reuse checks — CQ-20's stable host IRI, CQ-01's lineage walk — on a laptop with no network and confirm the alignment holds, rather than taking the chapter's word for it. The offline scope is a deliberate reproducibility choice, not a limitation: the borrow is provable where it can be proved.

What the registries cannot tell you is the domain knowledge needed to pick the right hit — and that barrier is real in biopharma. A registry will return dozens of "cell line" terms; only someone who knows that a manufacturing cell line is a managed production input with a research-master-working banking discipline will know that the manufacturing sense lives in IOF biopharma while the biological sense (a CHO cell, a node in the cell-line tree) lives in CLO. A registry lists "chromatography medium" without knowing that for a Protein A capture step the medium is a persisting consumable with a usage history: the running example's resin lot RESIN-PrA-07 sits at cycle 38 of a validated cycle-lifetime limit of 200 (the clean-in-place cycle and the resin-reuse claim are built in the capture-chromatography chapter), so once bp:ChromatographyResin is typed as iof:ChromatographyMedium a partner's tool can ask "how many cycles has this resin seen, and is carryover still within the validated reuse claim?" — the kind of resin-reuse question an inspector raises under the cell-substrate and process-validation guidance. The borrowed term is the hook the cycle count hangs on. The honest record of where each IRI was checked is itself a reuse artifact — written into the header of the alignment file — because a borrowed term whose source you cannot name is a term you cannot defend in a review, and a regulator reviewing a CHO process will ask exactly that.

Select what to borrow: the criteria, and why a near-miss is dangerous

Finding a candidate is not the same as adopting it. Each survey hit is judged against a fixed set of selection criteria before it earns an edge — and in a regulated antibody campaign the cost of a wrong adoption is not abstract:

How each reuse-vs-mint decision was made. The criteria below are not a checklist read in the abstract; they are a short, repeatable procedure run once per local term, and it is worth doing in the open on two real cases that land on opposite verdicts — bp:Specification (reused) and bp:FillFinishProcess (minted local):

  1. Name the local thing precisely. A specification is the released acceptance brief a lot is judged against; a fill-finish process is the aseptic step where bulk antibody becomes sealed, sterile vials. Saying exactly what the term means locally is what makes the next step a fair search rather than a keyword lottery.
  2. Search the registries for a settled external term. For the specification, the published IOF release (Release_202602) offers iof:RequirementSpecification — and notably no bare iof:Specification class. For fill-finish, the same release has a drug-product formulation process but no fill-finish, aseptic-fill, or lyophilization class at all.
  3. Test the best hit against coverage and the other criteria. Does the external term mean exactly the local thing? iof:RequirementSpecification does — a release specification is a requirement specification — and it is Released, BFO-grounded, and freely reusable, so it passes. For fill-finish there is no hit to test: the closest term (formulation) means a different step, so a borrow would be a near-miss that asserts a falsehood.
  4. Reuse by rdfs:subClassOf, or mint local and flag — never fake the borrow. A term that passes earns one upward edge to the real IRI — bp:Specification rdfs:subClassOf iof:RequirementSpecification — and never an owl:equivalentClass or owl:sameAs (the one-way claim "my term is a kind of theirs" is the only safe one). A term with no faithful hit stays a local bp: class with the gap recorded in align.ttl's header — bp:FillFinishProcess does exactly that — so a reader can see it was a deliberate, documented gap rather than an oversight.

Run on every local term, this procedure is what populates the criteria below; the criteria are the tests a candidate passes at step 3, and the cost of failing each is why the verdict matters:

  • Coverage — does the term mean exactly the local thing, or only roughly? A near-miss is worse than a local mint, because it asserts a falsehood that travels. (This is why bp:Specification aligns to iof:RequirementSpecification — there is no bare iof:Specification class to borrow.)
  • Maturity — is the term Released, or provisional and likely to move? A re-audit of the IOF biopharma release found its bioreactor, chromatography-column, and master-recipe classes were all Released, not the Provisional an earlier pass had assumed — which is what let the campaign's equipment and recipe be adopted with confidence rather than left as flagged local mints.
  • Adoption — is the ontology actually used by the databases and partners you want to interoperate with? GO and PRO win on this alone; public databases annotate to them by default, so the antibody's target already has a PRO IRI, and that target protein carries a GO molecular-function annotation — both waiting in the public databases [6][7].
  • Logical compatibility — does it sit on the same BFO spine, so importing it cannot make the model inconsistent? A CLO cell type and an IOF cell line both descend from BFO, which is the only reason one bp:CellLine can wear both at once.
  • License and import cost — is it freely reusable, and how heavy is the file you would pull in?

When a candidate fails coverage and no faithful 1:1 external term exists, the discipline is to keep the class local and flagged, not to fake a borrow — and the biggest such gap in this campaign is at fill-finish. The IOF biopharma release has a drug product formulation process but no fill-finish, aseptic-fill, or lyophilization class at all, so bp:FillFinishProcess, bp:ContainerClosureSystem, and bp:fillsInto stay flagged local classes. The temptation is to bend some near-term onto them; the cost of doing so is concrete: a downstream importer assembling a regulatory submission expects the real meaning of "aseptic fill," and a faked edge would silently mislabel the single step where the bulk antibody becomes the sterile, sealed vials a patient receives. Charge variants, aggregate variants, GS1 packaging tiers, and the governance records are kept local for the same reason — a documented gap is honest, an unfaithful edge is a latent error that surfaces at the worst time.

Formalize the borrowing: align.ttl, and the import gap

The reuse phase produces exactly one artifact: align.ttl, a file that does nothing but assert rdfs:subClassOf (and rdfs:subPropertyOf) edges from each local bp: term up to a real, dereferenceable external IRI. Here is the biology-and-manufacturing core of it — the host borrowed from NCBI Taxonomy, the line's biology from CLO, its manufacturing role from IOF biopharma, the IgG itself from GO, and the capture consumable from IOF:

# align.ttl — the host is borrowed (NCBI Taxonomy), the line's biology (CLO) and role (IOF) too.
bp:HostOrganism rdfs:subClassOf obo:NCBITaxon_10029 . # NCBI Taxonomy 'Cricetulus griseus' (Chinese hamster, verified via OLS4)
bp:WorkingCellBank rdfs:subClassOf obo:CLO_0002421 . # CLO 'CHO cell' (Chinese hamster ovary cell line)
bp:CellLine rdfs:subClassOf iof:CellLine . # IOF biopharma 'cell line' (Released, Release_202602)
bp:Clone rdfs:subClassOf iof:ClonedCellLine . # IOF biopharma 'cloned cell line'
bp:Antibody rdfs:subClassOf obo:GO_0071735 . # GO 'IgG immunoglobulin complex'
bp:Excipient rdfs:subClassOf obo:CHEBI_24431 . # ChEBI 'chemical entity' — the generic anchor; specific excipients (polysorbate, histidine, sucrose) would use ChEBI leaves
bp:CaptureChromatography rdfs:subClassOf iof:CaptureStep . # IOF 'capture step' (Protein A capture)
bp:ChromatographyResin rdfs:subClassOf iof:ChromatographyMedium . # IOF biopharma 'chromatography medium' (Released)
bp:derivedFrom rdfs:subPropertyOf obo:RO_0001000 . # RO 'derives from'; ALSO owl:TransitiveProperty in bioproc.ttl — this is what powers CQ-01's (bp:derivedFrom)+ walk
bp:bindsTo rdfs:subPropertyOf obo:RO_0002436 . # RO 'molecularly interacts with' (antibody-target binding, verified via OLS4)

Reuse is not only classes: the last two lines align relations. bp:derivedFrom borrows the OBO Relation Ontology's derives from (and is declared owl:TransitiveProperty in bioproc.ttl — meaning if A derives from B and B from C, then A derives from C automatically), which is exactly what lets CQ-01 resolve a genealogy of any depth with one (bp:derivedFrom)+ property path (a query pattern that follows the derivedFrom edge one-or-more hops); bp:bindsTo borrows RO's molecularly interacts with so the antibody-target binding reads as a standard molecular interaction. A reused relation is what makes the edges between borrowed classes interoperable, not just the nodes.

Each line carries a manufacturing fact, not just a vocabulary edge. bp:HostOrganism rdfs:subClassOf obo:NCBITaxon_10029 is what lets the cell-bank root name its species by a stable identifier instead of a string — the genealogy's first line of defence against misidentification. bp:Excipient rdfs:subClassOf obo:CHEBI_24431 anchors the polysorbate-80, histidine, and sucrose that keep the formulated antibody stable to ChEBI's generic chemical entity; the dataset deliberately stops at the generic anchor and flags the specific leaf IRIs as illustrative, exactly the honest-gap discipline above. And bp:CaptureChromatography rdfs:subClassOf iof:CaptureStep types the Protein A step every batch passes through as a shared manufacturing concept a partner's tool already understands.

Now the most important sentence in the chapter, because it is the gap most adoptions trip on. align.ttl contains no owl:imports. Asserting rdfs:subClassOf iof:CaptureStep gives a real, shared IRI another team's IOF-aware tool can line up with — but on its own it does not make a reasoner (the software that derives new facts from the stated ones) conclude that a bp:CaptureChromatography is a BFO process, because the chain from iof:CaptureStep up through IOF to BFO lives inside files the validator does not load. The validator loads only bioproc.ttl + align.ttl + instances.ttl, so the OWL-RL closure (the full set of facts the reasoner can deduce — every implied edge made explicit) gets transitive rdfs:subClassOf over the edges asserted here, plus the local axioms (the rules stated in those files) — but not IOF's own chain and not BFO's disjointness (the rule that, say, a process can never also be a material thing). The result is lexical interoperability (shared, real, dereferenceable IRIs — addresses you can look up to get the published definition back), not cross-ontology entailment (the reasoner actually deducing that your term is a BFO process).

What does that buy in the antibody world, concretely? A contract manufacturer or a regulator whose tools speak IOF can line up your bp:CaptureChromatography batches with their notion of a capture step without ever comparing English prose; a public database that annotates proteins to PRO can join your antibody's target to its records; a partner querying "which lots came from a Chinese-hamster line?" gets your batches back because the host carries the same obo:NCBITaxon_10029 everyone else uses. If they could not align — if the host were the string "CHO" and the capture step a private code — every one of those joins becomes a hand-built translation table maintained forever. To turn the assertion into an inference you owl:imports IOF Core (which itself imports BFO 2020), and the reasoner then classifies your terms through the whole stack. The running example does both halves honestly: align.ttl reuses the IOF terms it could confirm, and a companion bioproc-imports.ttl carries the real owl:imports declarations — while the offline validator stays scoped to what it can prove without fetching the external stack. The one-line summary: alignment buys shared vocabulary for free; cross-ontology reasoning only once you import.

One deliberate non-choice deserves a name: every edge in align.ttl is rdfs:subClassOf (the one-way claim "my class is a kind of theirs") or rdfs:subPropertyOf, and not one is owl:equivalentClass (a two-way claim that the two classes are exactly the same set) or owl:sameAs (a claim that two names denote the very same individual). The difference is not stylistic. Equivalence is bidirectional: assert bp:CaptureChromatography owl:equivalentClass iof:CaptureStep and you have promised that every IOF capture step is one of yours and vice versa — and the moment you owl:imports IOF, its domain, range, and disjointness axioms flow back onto your local class and can make the model inconsistent over facts you never authored. Subsumption is the safe, one-way claim the reuse phase actually wants: my term is a kind of theirs. Aligning up by subClassOf is precisely the discipline that lets the import in the companion file be a no-risk reasoning upgrade rather than a landmine — and it sidesteps the single most common alignment defect (sameAs/equivalentClass where subClassOf was meant, the OOPS! pitfall scanner's equivalent-classes/relationships not explicitly declared and merging different concepts family), which this file avoids by construction.

The two ways of borrowing line up cleanly:

align.ttl alone (what the validator loads)+ bioproc-imports.ttl (owl:imports IOF Core, which imports BFO 2020)
What is assertedbp:CaptureChromatography rdfs:subClassOf iof:CaptureStepthe same edge, plus IOF's own chain and BFO's axioms
What a reasoner concludesthe asserted edge and its transitive closure over edges in this filebp:CaptureChromatography is an IOF manufacturing process and a BFO process
What a partner's tool getsa shared, dereferenceable IRI to line up against — lexical interoperabilityfull cross-ontology entailment through the stack
Costone file of edges; nothing external fetchedthe external ontologies pulled in and reasoned over

The two halves the running example keeps honest: align.ttl earns shared vocabulary without fetching anything, and bioproc-imports.ttl carries the real owl:imports for teams that want the reasoner to classify local terms through the whole IOF-to-BFO chain.

A top band of four borrowed reuse layers (BFO 2020 foundational, IOF Core plus biopharma, the OBO domain ontologies, and the cross-cutting vocabularies), and below them a band of local bp classes grouped into an IOF-aligned panel and an OBO-aligned panel, each connected by one labeled rdfs or rdfs arrow pointing up to its borrowed parent column, with a rose callout at the bottom stating the honest import gap: align.ttl asserts the subClassOf edges but contains no owl, so a reasoner gets lexical interoperability, not cross-ontology entailment. How align.ttl borrows: each local bp: term asserts one rdfs:subClassOf or rdfs:subPropertyOf edge up to a real external IRI — the host to NCBI Taxonomy, the capture step and cell line to IOF, the IgG and the derives-from relation to OBO — and the rose callout names the honest gap: the file carries no owl:imports, so the result is shared, dereferenceable vocabulary, not cross-ontology entailment. Original diagram by the authors, created with AI assistance.

The OBO–IOF seam: a crosswalk you author, running through a living cell

Reuse here is not a single clean import; it straddles a fault line, and the fault line runs straight through the most important node in the whole graph — the cell line. The ontologies that best describe a cell line's biology — CLO, NCBI Taxonomy, GO, PRO — grew up in the OBO Foundry. The ontologies that best describe making the antibody — IOF Core and IOF biopharma — grew up in the Industrial Ontologies Foundry. Both descend from BFO, which is precisely why a single bp:CellLine can sit under both obo:CLO_0002421 (its biology, the CHO cell) and iof:CellLine (its manufacturing role, a managed production input) at once. But "can meet" is not "have met": there is no published bridge that says how CLO's CHO cell relates to IOF's cell line, so that crosswalk is author-curated — a mapping the program writes, reviews, and maintains, not a feature it imports.

The seam matters here more than anywhere because of why cell banks need genealogy in the first place. A cell line is alive: a population of cells that mutate and drift over generations, so the working bank is not genetically identical to the master, and a culture grown too long can lose productivity or shift its product quality. That is why the campaign keeps a research / master / working cell bank hierarchy (RCB → MCB → WCB) — a structured taxonomy of disjoint tiers, characterized (tested for identity, sterility, and freedom from contaminants) and passage-limited (capped at a maximum number of culture generations, since "passage" counts how many times the cells have been split and re-grown) — rather than one vial nobody can vouch for. The biological side of the seam (is this still a CHO cell? what is its species?) and the manufacturing side (is this WCB within its validated passage limit and fully characterized?) are both true of the same WCB-CHO-001, and only an authored crosswalk holds them together. In the running example both halves resolve to numbers: the biological side names the species as obo:NCBITaxon_10029 (Chinese hamster), while the manufacturing side reads WCB-CHO-001 at passage 8 against a validated passage limit of 40 and four characterization verdicts (identity, sterility, viral, genetic) all PASS — and the seed train it inoculates accumulates passages (12, then 16) back toward that limit. The borrowed CLO/NCBITaxon IRIs answer "what species?"; only the local, dated passage and characterization data answer "is this vial still safe to inoculate from?" — which is exactly why the seam is author-maintained rather than imported (the cell-bank gate that reads this is CQ-17 in the release-gate chapter). The target chapter names this seam at the day-one discovery boundary, where the antibody is still a product concept — an information artifact that exists before any cell makes the molecule; the reuse phase is where the seam becomes a maintained artifact, because every borrowed edge that crosses it is one the author justified term by term.

Evaluation: does the reuse hold up?

Reuse is validated the same way every phase is — by loading the graph, reasoning over it, and running the competency question that depends on it. CQ-20 is that test, and it is a pure reuse check: it asks the reasoned graph to return the host organism's stable external taxon IRI, which exists only because bp:HostOrganism was aligned to obo:NCBITaxon_10029. The test is written in SPARQL — the query language for this kind of graph, the way SQL queries a relational database:

# queries/CQ-20.rq — CQ-20: the host is named by a stable public identifier, not the string "CHO".
PREFIX bp: <https://example.org/bioproc#>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
PREFIX obo: <http://purl.obolibrary.org/obo/>
SELECT ?host ?taxon WHERE {
bp:WCB-CHO-001 bp:hasHostOrganism ?host .
?host a ?cls .
?cls rdfs:subClassOf ?taxon .
FILTER(STRSTARTS(STR(?taxon), STR(obo:)))
}

The check in cq-catalog.json requires the single row host = CHO-host, taxon = NCBITaxon_10029 — and validate.py returns exactly that. Note the query starts at bp:WCB-CHO-001, the working cell bank at the root of every genealogy chain, and traverses an edge that lives in align.ttl (bp:HostOrganism rdfs:subClassOf obo:NCBITaxon_10029): the test passes only because the reuse was done, and would fail the instant the host were typed by the bare string "CHO" instead. A second reuse-dependent test runs in the other direction: because bp:derivedFrom was aligned to the OBO Relation Ontology's obo:RO_0001000 derives from and declared owl:TransitiveProperty (in bioproc.ttl), CQ-01 walks DS-001's lineage with a single property path (bp:derivedFrom)+ — back through the polishing intermediate, the viral-filtered and viral-inactivated pools, the capture pool, the clarified harvest, the bioreactor batch, the seed-bioreactor culture, the shake-flask culture, and the three cell-bank tiers (WCB, MCB, RCB) — to 11 ancestors, with no join table and no hand assertion (the drug substance DS-001 is the subject of the walk, not an ancestor). Reuse of a relation (RO) is doing the work here, not reuse of a class: the same alignment that gives partners a shared derivedFrom IRI is what lets one + operator replace a recursive query. The 11-ancestor count is read off the raw asserted graph by the property path and needs no reasoner; the same long-range edges are also available as inferred derivedFrom triples once the OWL-RL closure runs, and that transitive inference is its own structural test (CQ-22). Across the full dataset the alignment scales to a graph of 2120 triples (a triple is the atomic unit of this data — a single subject-predicate-object statement, e.g. one rdfs:subClassOf edge) closing to 7137 under OWL-RL (the reasoner deduces every implied triple, growing 2120 stated facts to 7137 total), with the distinct IOF classes consumed grown to 27 (7 Core + 20 biopharma) — the measurable yield of shopping before building. The full per-CQ acceptance harness, where all 23 of these checks run as PASS/FAIL tests, is the subject of the validation chapter.

And because the purification steps are typed by shared IOF process classes rather than private codes, a reviewer can read the campaign's quality trajectory straight off the chain: high-molecular-weight aggregate sits at 4.1 % in the Protein A capture pool, falls to 1.4 % after the iof:PolishingProcess step, and reads 1.287 % in the released drug substance — comfortably under the 2.0 % specification (this is CQ-05). The aggregate is removed at polishing, not wished away; borrowing the step type is what lets a partner's tool locate "where did quality change?" without parsing prose (the HMW attribute and its 2.0 % criterion are authored in the taxonomy chapter).

What the reuse buys a learning model: a validated ground truth

The same reused alignment that lets a partner's tool line up your batches is what lets a model learn over them without learning a lie — and that is the load-bearing tie between this phase and the companion volume on machine learning. A learning system needs a ground truth, and an ontology-grounded graph is a stricter ground truth than a flat training table in three concrete ways that the reuse phase, specifically, is what supplies.

First, the training data can be validated before a model ever sees it. The release gate's SHACL shapes (the constraint language that checks a graph conforms to required shapes before it is trusted) refuse a batch node missing its bp:derivedFrom parent or a CQA with no unit-bearing value — and a subgraph that passes those shapes is exactly a training set with no silent holes for a model to confabulate over. This is the same machinery the ontologies-and-AI chapter points at a retrieval: a graph that fails its shapes is the hollow or mislabelled input that a fluent model completes from memory. The reuse decision is upstream of that check — the value passes its shape only because it arrived through QUDT already wearing its unit, which is the alignment this phase made.

Second, the unit of learning is the batch, not the row, and the graph already knows it. A Raman spectrum yields hundreds of correlated columns but a single batch is one genuinely independent observation, so a naive random train/test split leaks information across the boundary and reports an honest-looking accuracy that collapses in production — the cardinal sin the models-and-validation chapter names. Grouped cross-validation — leave-one-batch-out, holding every row of one batch together on one side of the split — is the only honest evaluation, and the (bp:derivedFrom)+ lineage walk above is the grouping key: every row tracing to the same BATCH-2026-001 belongs together because the genealogy says so, not because a spreadsheet column happened to match. The reused, transitive derivedFrom relation is doing double duty — a traceability edge for a regulator and a leakage-proof fold boundary for a validator.

Third, the graph is a check on the model, not just its fuel. Type a model's output back as a triple and the reasoned graph can contradict it: a predicted release for a lot whose host fails CQ-20, or whose iof:PolishingProcess HMW never crossed the 2.0 % line, is a claim the OWL-RL closure and the SHACL gate can refuse before it reaches a human — the validation paradox that a model is only ever as trustworthy as the structured knowledge it is anchored to, restated as a reasoner that outranks the predictor. For a graph-native model the borrowed edges become features directly: "host is a Chinese-hamster line" is obo:NCBITaxon_10029, "this resin is at cycle 38 of 200" hangs off iof:ChromatographyMedium, and a graph-neural or retrieval-augmented model reads typed, dereferenceable facts rather than re-derived strings. None of this is free of the field's hard limit — pooling two sites' "day-7 feed rate" still fails unless every contributing process is semantically harmonised first, which the ML frontier chapter names as the expensive, unfinished prerequisite; the reuse-and-alignment discipline of this chapter is precisely that prerequisite, done once, in the open.

The unsolved part: a borrow can drift, and a near-miss can hide

Reuse trades one risk for another. By aligning to external IRIs the model inherits their meaning — but also their changes. NCBI Taxonomy revises, IOF ships new releases, a CLO term can be obsoleted; an edge verified against Release_202602 is true at that release and is a maintenance obligation thereafter, which is why the governance chapter treats the alignment as versioned content under change control. The subtler hazard is the near-miss that looks faithful: NCIt's Indication and Mechanism of Action are bound as lexical bridges to a shared concept name, not as BFO category claims, because NCIt is a thesaurus, not a BFO-partitioned ontology — assert them as if they were entailment-grade and you have smuggled an unfaithful claim into the graph. There is a deeper limit the alignment cannot touch at all: the cell line it so carefully names is alive. An IRI implies a crisp, stable identity, but no owl:sameAs and no borrowed CLO term answers whether the culture at passage 60 is the same entity as the culture at passage 5, and no alignment can certify that the vial labelled WCB-CHO-001 actually contains the line the label claims. Cell lines have been confused and cross-contaminated across the life sciences for decades — which is exactly why stable identifiers and authentication exist [10] — and for a manufacturing root node a misidentification is the worst possible error: asserted with full confidence, it propagates transitively through derivedFrom to every descendant batch, and no amount of downstream data integrity catches it, because every downstream fact is correctly derived from a wrongly identified root. Reuse gives the root a stable, public name; it cannot make the thing at the root be what the name says. The discipline that keeps reuse honest is the same one that makes it valuable: borrow only what genuinely matches, name the source, and flag — never fake — the gap.

Why it matters

The reuse phase is where a model decides whether it will be interoperable or merely internally consistent. Borrow the host as obo:NCBITaxon_10029 and a partner can ask "which lots came from a Chinese-hamster line?" across every program's graph at once; borrow the Protein A step as iof:CaptureStep and a contract manufacturer's IOF-aware tool lines up your batches without translation; borrow the IgG as obo:GO_0071735 and the antibody's molecular function joins the public web of biomedicine; borrow the excipients up to ChEBI and the formulation is traceable structure rather than a buried recipe. CQ-20 passes, CQ-01 returns its 11 ancestors, and a regulator's structured-data expectations are already half-met — all because the survey was done first. Skip it and mint private codes for the species, the unit operations, the antibody, and the excipients, and every one of those connections becomes a retrofit project paid for at the worst possible time, when a lot fails and the recall must be scoped now. The cell bank is the cheapest place in the lifecycle to get identity right and the most expensive to get wrong, and the reuse phase is where that bill comes due — or does not.

In the real world

Reuse-first is established practice in research informatics and still maturing in manufacturing. GO, PRO, OBI, CLO, and the disease ontologies are mature, funded, and reused by default across public databases [6][7][10]; a real antibody program's target already has a PRO IRI and a GO annotation waiting, and CHO cells — which dominate commercial antibody production — have a sequenced public genome and stable taxonomy and cell-line identifiers ready to borrow. The RCB/MCB/WCB hierarchy and its characterization are not optional good practice either; they are an expectation of the regulatory guidance on cell substrates, so every real program already maintains exactly the banking lineage this reuse anchors. The IOF biopharma modules were released recently and are substantial — a Released vocabulary of unit operations, equipment, materials, and QbD parameters — but they are not yet adopted at the scale the OBO side enjoys, and they have real holes (no fill-finish class). The frontier this book walks is precisely the join: carrying a target's OBO identity all the way into the IOF-described cell line, bioreactor batch, and drug product, authoring the OBO–IOF crosswalk by hand at the living cell where the two worlds meet, and publishing the alignment as the maintained, versioned reuse artifact it is rather than a one-time mapping spreadsheet.

Key terms

  • IRI — Internationalized Resource Identifier; a global, dereferenceable web address (like a URL) that names exactly one thing, so independent systems can confirm they mean the same entity — the unit of identity the whole reuse phase borrows.
  • Triple / closure — a triple is the atomic unit of this data (one subject-predicate-object statement, e.g. a single rdfs:subClassOf edge); the closure is the larger set of triples a reasoner derives from the stated ones (here 2120 asserted → 7137 after OWL-RL).
  • Reuse (NeOn scenario) — the lifecycle phase of surveying, finding, selecting, and aligning to existing ontologies before authoring new terms; a non-functional requirement that here underwrites CQ-20 (stable host identity) and, transitively, CQ-01 (the 11-ancestor lineage walk).
  • Layered survey — looking for reusable terms at the foundational (BFO), mid-level (IOF Core), domain (IOF biopharma + OBO members), and cross-cutting (QUDT, PROV-O, SOSA, SKOS, Dublin Core, AFO) levels — each justified by a different one of the entities inside a single batch.
  • Continuant / occurrent split — BFO's distinction between things that persist (the batch material, the vessel, the purity, the role) and things that happen (the fermentation, the capture step); what keeps equipment, material, and activity separable so a vessel can host many batches without contradiction.
  • Registry — a catalog used to find candidates: OLS4 and OBO Foundry and BioPortal for OBO, the published IOF release for IOF, LOV and publishers' sites for the cross-cutting layer; what it cannot supply is the domain knowledge to pick the right hit.
  • Selection criteria — coverage, maturity, adoption, logical compatibility, license, and import cost; the tests a found term passes before it earns an alignment edge, where a near-miss that asserts a falsehood is worse than a flagged local mint.
  • Lexical alignment vs. owl:imports — asserting rdfs:subClassOf to an external IRI buys shared vocabulary (so a partner's IOF tool aligns your batches) without loading the external file; owl:imports pulls the external chain in and buys cross-ontology entailment.
  • align.ttl — the single reuse artifact: a file of rdfs:subClassOf / rdfs:subPropertyOf edges from local bp: terms up to verified external IRIs, with the verification source recorded per term.
  • OBO–IOF seam — the unbridged boundary between the biomedical (OBO) and manufacturing (IOF) ontologies; both BFO-grounded, so a single bp:CellLine can wear both, but the crosswalk through the living cell is author-curated and maintained.
  • Grounded ground truth (ontology for ML) — an ontology-grounded, SHACL-validated graph used as the verified input and the cross-check for a learning model: training data with no silent holes, the lineage walk as a leakage-proof grouping key, and a reasoned graph that can contradict a model's prediction.
  • Leave-one-batch-out cross-validation — grouped evaluation that keeps every row of one batch on one side of the train/test split, because a batch (not a row) is the independent observation in bioprocess; the reused transitive bp:derivedFrom lineage supplies the grouping key.

Where this leads

The parts are bought and the parts list is written. With the borrowed vocabulary fixed, the conceptualization phase can finally draw the local terms — the small set of classes the survey proved no one else publishes: the fill-finish process and container-closure system (with bp:fillsInto), the charge and aggregate variants, the GS1 packaging tiers, and the governance records, each one a documented gap in the borrowed vocabularies rather than a near-miss faked into an edge. The next chapter, Conceptualization: Classes and the Taxonomy, turns from borrowing to authoring: naming each local bp: class, fixing its BFO category, and arranging the taxonomy so that every new term either specializes a reused one or earns its place as a justified, flagged local mint.