Skip to main content

The Biologic and Its Data Shadow

📍 Where we are: Part I, Chapter 1 — having met the two products of biomanufacturing in the Preface (the molecule and its data), we now look closely at that second product: the data shadow that every batch casts.

In the Preface — Making the Same Medicine Twice — we made a claim that sounds almost philosophical: a biologic is manufactured twice. Once as a molecule, grown inside living cells, and once as a body of data that proves the molecule is what we say it is. This chapter takes that claim and makes it concrete. We are going to follow a single batch of medicine and count its shadow.

The word "shadow" is deliberate. A shadow is not the object, but it is cast by the object, follows it everywhere, and tells you a great deal about its shape. The data shadow of a batch is every number, record, signature, and sensor trace generated while that batch was made and tested. By the end of this chapter you will see that the shadow is not optional paperwork bolted on at the end. It is, in a real regulatory and scientific sense, part of the product.

The simple version

Think of an expensive bottle of wine. The wine in the glass is the product — but a serious collector also wants the provenance: where the grapes grew, the weather that year, who bottled it, how it was stored, and proof the bottle was never tampered with. For wine, that paper trail is a nice-to-have. For a biologic medicine going into a human body, the paper trail is the law, and it is enormous. This book is about that paper trail — except almost none of it is paper anymore.

What this chapter covers

  • Why a biologic's complexity forces data to the center of the story
  • The many kinds of data one single batch produces, and roughly how much
  • The legal rule that "if it isn't documented, it didn't happen"
  • The four families of data we will return to again and again
  • Why the verb in this book's title — manage — is the hard part

The molecule that is defined by how it is made

A quick reminder, because everything else follows from it. A biologic is a large, fragile protein medicine — a molecular machine made of thousands of atoms — that can only be built by living cells, not by ordinary chemistry. Our recurring example is the monoclonal antibody (mAb), a Y-shaped protein engineered to lock onto one specific target in the body, such as a cancer signal. (The companion guide Biologic Drug Manufacturing covers the biology in depth; here we care about its consequences for data.)

A small-molecule drug like aspirin can be written as an exact chemical formula. A biologic cannot. It is too big, too flexible, and decorated with tiny sugar chains that the cells attach as they grow it. Two factories with the same gene but different culture conditions — pH, temperature, nutrient levels — can produce subtly different molecules, because living cells perform post-translational modifications (PTMs), chemical edits made to the protein after it is built, such as glycosylation (the attaching of sugar chains) that shift with the cell's state. This is why the industry repeats the phrase "the product is the process" — the manufacturing process does not merely make the medicine, it helps define what the medicine is [1].

That single fact is the reason data sits at the center of biomanufacturing rather than at the edge. If the process defines the product, then the record of the process is your only proof of what the product actually is. Modern regulatory thinking formalized this. Under Quality by Design (QbD) — a development philosophy in which quality is built in deliberately rather than tested in afterward — you must identify which process settings and which product properties truly matter, and understand how they relate [1]. The international guideline ICH Q8(R2) — ICH being the International Council for Harmonisation, the body that aligns drug regulation across the US, EU, and Japan — turned this into expectations: define what the product needs to do, identify its critical quality attributes (the handful of product properties that have to stay within limits, formalized later as CQAs), and map the design space (the proven range of process conditions within which the product reliably meets those attributes) [2]. Every one of those words is, in practice, a request for data.

The data shadow of a single batch

Now picture one batch — one run through the factory, producing perhaps a few kilograms of purified antibody (the drug substance — the bulk antibody before it is filled into vials) that, after fill/finish, become tens of thousands of vials of finished drug product. Watch the data it throws off.

What gets captured: the six data sources

A single run does not throw off one kind of data; it throws off six, each born at a different physical step. The Preface already sorted the same shadow into six sources, grouped by how each record behaves (continuous sensor readings, at-line and off-line lab results, phase-transition events, the batch master record, process events, and the audit trail with signatures). Here we cut the shadow a second, complementary way — by the physical station that creates each record — which folds the Preface's phase-transition and process events into the control system's alarms and events log, keeps its batch master record as the EBR, and surfaces two stations the Preface had bundled into context: environmental monitoring and material genealogy. Same shadow, two slices — not a different set of records: sensor time-series, alarms and events, the EBR, analytical results, environmental monitoring, and material genealogy. The first three are born on the manufacturing floor — in the production bioreactor and at harvest and clarification — and the last three are born in the QC lab, the cleanroom, and the warehouse. Meet them in turn.

Time-series from sensors. Inside the bioreactor (the large vessel where cells grow — for example a 2,000-liter stirred-tank vessel such as a Cytiva Xcellerex XDR or a Thermo Fisher HyPerforma single-use bioreactor), probes measure temperature, pH (acidity), dissolved oxygen, stirring speed, and more, often every few seconds, for one to three weeks. These continuous streams are the heartbeat of the batch. The framework that pushed the industry to measure such critical quality and performance attributes during processing (in real time or on rapid at-line measurement) is the FDA's Process Analytical Technology (PAT) initiative [3]. The numbers are large: ten to twenty probes sampled around 0.5 Hz over one to three weeks easily exceed ten million points per batch. That is the raw acquisition rate; a historian rarely stores every sample. To save space it records a reading only when the value changes by more than a small set amount — a band of change too small to bother recording is called a deadband — so a flat, steady signal is kept as a handful of points rather than thousands. (This technique of storing only meaningful changes is called exception reporting, or compression; the AVEVA PI System's swinging-door algorithm is a well-known example.) The stored point count is therefore far smaller than the acquisition rate, which is itself a data-integrity design choice we revisit later. Each point arrives as a timestamped tag — BR101.Temp.PV, BR101.pH.PV (the .PV suffix means process value, the live measured reading) — landing in a process Historian database such as the AVEVA PI System (formerly OSIsoft PI) or Rockwell FactoryTalk. A few rows look like this:

timestamp,tag,value,unit,quality
2026-06-13T14:03:00Z,BR101.Temp.PV,37.0,degC,Good
2026-06-13T14:03:00Z,BR101.pH.PV,7.05,pH,Good
2026-06-13T14:03:00Z,BR101.DO.PV,42.3,%sat,Good
2026-06-13T14:03:02Z,BR101.Temp.PV,37.0,degC,Good
2026-06-13T14:03:02Z,BR101.pH.PV,7.04,pH,Good

Alarms and events. Every time a value drifts out of range, an operator opens a valve, or a pump starts, the control system logs it — what happened, when, and often who triggered it.

The electronic batch record (EBR). This is the master narrative of the batch: the step-by-step recipe that was followed, the materials added, the parameters confirmed, and the human sign-offs at each stage. It lives in a Manufacturing Execution System (MES) — commercial examples include Körber PAS-X, Siemens Opcenter Execution Pharma, and Rockwell PharmaSuite. We will meet it properly in later chapters; for now, know it is the spine the rest of the data hangs from.

Analytical results. Once material is harvested and purified — through a whole downstream train of unit operations: capture chromatography on a Protein A column, low-pH viral inactivation (an acid hold that kills enveloped viruses), polishing chromatography on cation- and anion-exchange columns that strip residual aggregate and host-cell protein, a viral-filtration step with its post-use membrane integrity test, and a final UF/DF (ultrafiltration/diafiltration) that concentrates and buffer-exchanges the antibody into drug substance — the Quality Control (QC) laboratory tests it for identity, purity, potency, and safety; these results are what eventually drive QC and release. Each of those steps casts its own data: a column's elution chromatogram and pool fraction, a hold step's pH-and-temperature trace and validated log-reduction value (LRV) (the factor by which a viral-clearance step cuts the virus count), a filter integrity-test pass/fail, a diafiltration's diavolume count and final concentration. The results land in a Laboratory Information Management System (LIMS) — the QC lab's system of record. Each test produces results, and behind each result sits a raw instrument file: the original output of a chromatograph or mass spectrometer, often a large and proprietary digital file — a single high-performance liquid chromatography (HPLC) run, for instance, is commonly stored as a vendor-specific .ch file or a .d directory of 5–100 MB.

Environmental monitoring (EM). Cleanrooms are watched constantly — airborne particles, microbial samples, temperature, humidity — to prove the surroundings stayed clean while the batch was open.

Material genealogy. Every raw material, growth medium, filter, and single-use plastic bag carries a lot number, a supplier, and a certificate. The thread linking a finished vial back to every ingredient and consumable that touched it is its genealogy, or lineage.

The six sources that make up one batch's data shadow — very different kinds of data, all describing the same run:

Data sourceWhat it capturesTypical system or format
Sensor time-seriesTemperature, pH, dissolved O₂ tracesHistorian; ~10 M readings per batch
Alarms & eventsOut-of-range drifts, valve and pump actions, who actedControl-system event log
Electronic batch record (EBR)The recipe followed, materials added, parameters, sign-offsMES (PAS-X, Opcenter, PharmaSuite)
Analytical resultsIdentity, purity, potency, safety — plus raw instrument filesLIMS; .ch / .d files, 5–100 MB
Environmental monitoringAirborne particles, microbial counts, temperature, humidityEM system
Material genealogyLot numbers, suppliers, and certificates for every inputGenealogy / lineage records

A single batch as a hub: six data sources — sensor time-series, alarms and events, the EBR, analytical results, environmental monitoring, and material genealogy — all connecting into one central batch anchor A single batch's data shadow: six sources, all anchored to one shared s88.batch identity. Original diagram by the authors, created with AI assistance.

Notice the variety, not just the volume. Tidy numeric streams sit beside large binary instrument files, beside human-signed forms, beside supplier certificates. They live in different systems, in different formats, and the hard work of this book is connecting them. Shared standards exist to make that connection possible — ISA-88 (also issued as ANSI/ISA-88, the batch-control model that defines recipes, procedures, and equipment hierarchy) gives the batch a common structure, and OPC UA (Open Platform Communications Unified Architecture) is the modern protocol that lets instruments and historians exchange those tagged readings between vendors — though, as later chapters show, a standardized information model (an OPC UA Companion Specification) for a mammalian-cell bioreactor does not yet exist, so the semantics still vary plant to plant — but applying them across an entire facility is the real labor.

Anatomy of a data shadow: one batch's six-source record

It helps to see the six sources not as a list but as one record. Dissect a single batch — call it BATCH-2026-001 — and every source resolves to the same identity card. Each source has a role (what it captures), a format, and a home system, but they all share one anchor: the batch's identity. That anchor is an ISA-88 batch entity — in a relational store, a single s88.batch row carrying batch_id, recipe_id, unit_id, product_id, status, and the released lot. Every sensor tag, every QC result, every consumed material lot points back to that one row. Strip the anchor away and you have six disconnected islands; keep it, and you have one batch that can tell its whole story.

Anatomy card of one batch's data shadow: six sources, each with its role, format, and home system, all bound to a central s88.batch entity by batch_id, with the relationship edges that link them One batch dissected: six heterogeneous sources, each with its role and home system, all anchored to a single s88.batch entity and joined by genealogy edges. Original diagram by the authors, created with AI assistance.

This s88.batch row is not an abstraction we invented for the figure — it is exactly the artifact the companion implementation guide builds. Open-Source Bioprocess Data Systems shows the concrete schema in its batch and equipment model chapter: the s88.batch table and the s88.genealogy edges that stitch every source to its batch are real SQL, and the first sensor stream that feeds it is wired up in the upstream bioreactor chapter. The card above is the bridge: the physical step (a production bioreactor run) casts a data point, that point becomes a row in s88.batch, and the row is what an auditor — or the next batch's analytics — actually queries.

The same anchor, written as a graph

A relational s88.batch row is one way to express the anchor; a knowledge graph is another, and naming it now previews how the companion Ontologies for Biopharmaceutical Manufacturing book treats the very same identity. In a graph, each fact is a triple — a subject, a predicate, and an object — and the join key becomes an edge you can walk rather than a foreign key you must join. The batch's lineage and a sensor reading hung on it read like this (in Turtle, the text syntax for RDF; a means "is a", and bp: is just a short alias for a vocabulary namespace):

bp:BATCH-2026-001 a bp:Batch ;
bp:derivedFrom bp:SEED-001 ; # the seed culture it grew from
bp:occursIn bp:BR-101 . # the vessel, kept a SEPARATE node
bp:BR101-Temp-001 a bp:SensorReading ;
bp:ofBatch bp:BATCH-2026-001 ; # the same anchor the SQL row carries
bp:value "37.0"^^xsd:float ;
bp:unit unit:DEG_C .

Two design choices in those few lines are the whole bridge to Book 4. First, bp:derivedFrom is declared a transitive property, so lineage to any depth — the vial back to its cell bank — is inferred from only the immediate parent edges, never hand-asserted; that single relation is the spine the ontology's genealogy chapter is built on. Second, the batch (a continuant — a thing that persists and bears qualities through time) is kept a different node from the run that made it (an occurrent — a process that happens and is over) and from BR-101 the vessel: collapse "the batch" into "the bioreactor" and a vessel that processes a hundred batches a year inherits one batch's release facts and breaks every lineage trace through it.

That graph is also the natural home of the rule that required data must be present. A relational NOT NULL constraint can demand a value exists; the graph's equivalent is a SHACL (Shapes Constraint Language) shape that gates a release — "every released lot must carry exactly one in-range monomer-purity result and an attributable signature, or it fails." And the questions this chapter keeps asking — which lots share a failed batch's cell-bank ancestor? — become one-line SPARQL (the query language for RDF) traversals up the derivedFrom spine, a competency question the model is built to answer. The release gate, the SHACL shapes, and those queries are worked end-to-end in the ontology's release-gate-and-SHACL chapter; the point here is only that the s88.batch anchor and the graph's batch node are the same identity wearing two clothes — a foreign key in SQL, a walkable IRI in RDF.

The management pipeline: capture, contextualize, protect, connect, retain

Having the six sources is not the same as managing them. A managed data shadow flows through a pipeline of five verbs: capture at the moment of creation, contextualize by joining metadata, protect so nothing can be silently altered, connect across the systems that hold the fragments, and retain for the years the law demands. The shared s88.batch anchor is where the six converge; the pipeline is what turns that pile of records into something an auditor can trust and an engineer can learn from.

Flow diagram: six data sources converge on the shared s88.batch anchor entity, then pass through a five-stage pipeline of capture, contextualize, protect, connect, and retain

We return to each verb in depth — capture and contextualization in The Lifecycle of a Data Point and Where Data Is Born, protection in Data Integrity and ALCOA+, and connection in Semantic Interoperability and The Digital Thread. The point here is that "manage" is a pipeline, not a single act.

If it isn't documented, it didn't happen

In ordinary work, you do the job and the paperwork is a chore afterward. In medicine manufacturing, the rule is inverted. The records are the evidence that the job was done correctly, and without them the work is treated as if it never occurred. This is the heart of cGMPcurrent Good Manufacturing Practice, the body of regulation governing how medicines are made.

In the United States, 21 CFR Part 211 — Title 21 of the Code of Federal Regulations, the codified body of US federal rules — spells this out. It requires a master production record (the approved recipe), a batch production and control record for every batch (the as-executed account), the recording of test results, and the protection of backup data — all kept on file for years after the batch is released [4]. The data shadow is not a courtesy. It is a legal obligation, batch by batch.

The records must also be trustworthy, which is a separate problem from merely existing. The FDA's guidance on data integrity describes the qualities reliable records must have, often summarized by the acronym ALCOA — Attributable, Legible, Contemporaneous, Original, and Accurate — backed by audit trails that capture who changed what and when [5]. Regulators now extend these five into ALCOA+, appending four further qualities (Complete, Consistent, Enduring, Available); we return to the full set in Data Integrity and ALCOA+. The word Contemporaneous matters most here: you must record an action as it happens, not reconstruct it from memory later. A reading written down an hour after the fact is, in the eyes of a regulator, a different and weaker kind of evidence.

caution

"Contemporaneous" is why so much of biomanufacturing now happens through validated computer systems rather than notebooks. A sensor that timestamps its own reading the instant it takes it is the strongest possible witness. A human transcribing that number into a logbook later is the weakest. Much of modern data management exists to keep records close to the moment they are born.

Because most of these records are now electronic, a second rule applies: 21 CFR Part 11, which sets the conditions under which electronic records and electronic signatures are accepted as the legal equals of paper and ink [6]. When an operator clicks "approve" in a batch record system, Part 11 is what makes that click binding. The same expectations exist outside the United States: in the European Union, EU GMP Annex 11 governs computerised systems and is the close counterpart to Part 11 (and currently being modernized — a draft revised Annex 11 was issued in 2025 to address networked, multi-system data integrity), so a product sold on both sides of the Atlantic must satisfy both.

But a click is only binding if the system under it has been proven to work, which is a separate obligation from the record. Before a Historian, LIMS, or MES may hold a GMP record at all, it must be validated — proven with documented evidence to do what it is intended to do and nothing it should not. The traditional discipline is Computerized System Validation (CSV), classically structured as the V-model's three qualification rungs — IQ/OQ/PQ (Installation, Operational, and Performance Qualification: proof that the system was installed right, operates right, and performs right on its real workload). Over two decades CSV calcified into screenshot-everything paperwork, so the FDA's Computer Software Assurance (CSA) reframing now pushes effort toward critical thinking and risk: validate a system that auto-releases a lot far harder than a label printer. The full V-model, the GAMP 5 software categories, and the CSV-to-CSA shift get their own treatment in Validating Computerized Systems; the point here is that the data shadow is only trustworthy if the systems casting it are themselves qualified.

Four families of data

The shadow is large, but it is not formless. Throughout this book we will sort data into four families. Meet them briefly now; each gets its own treatment later.

  • Process data — the CPPs. A Critical Process Parameter (CPP) is a setting that, if it varies too much, will change the product — bioreactor temperature, pH, feed rate. These are the "how we made it" numbers, mostly the sensor time-series above. Identifying which parameters are truly critical is exactly the QbD exercise that ICH Q8 demands [2].

  • Quality data — the CQAs. A Critical Quality Attribute (CQA) is a property of the product itself that must stay within limits to keep the medicine safe and effective — its purity, its potency, the pattern of sugar decorations on the antibody. These are the "what we made" numbers from the QC lab — and the analytical procedures that produce them now have their own QbD guideline, ICH Q14 (Analytical Procedure Development) [11], a companion to the method-validation guideline ICH Q2(R2), which ICH adopted alongside Q14 in 2023 [12]. The A-Mab case study — a freely published, hypothetical-antibody example written by an industry consortium to show QbD in action — is a worked illustration of teams systematically ranking an antibody's CQAs and CPPs [9].

  • Metadata. Data about data: the units, the timestamp, the instrument, the operator, the calibration status. A bare number — "37" — is meaningless. 37 degrees Celsius, measured by probe TT-101, calibrated last Tuesday, at 14:03 is information. (Probe tags follow an equipment-specific schema — BR-01-TT-01 might mean bioreactor unit 01, temperature transmitter 01 — and the exact convention must be fixed in the facility's data dictionary so every system reads the same name the same way.) Metadata, and the audit trails that protect it, are central to the integrity rules above and to the risk-based system controls described in GAMP 5, the standard guide for validating computerized systems in regulated industry [10].

  • Master data. The stable reference information that does not change batch to batch: the approved recipe, the product specification, the list of materials, the equipment register. If process and quality data are the story of one batch, master data is the unchanging cast of characters every batch shares.

Hold these four loosely for now. The point is simply that the shadow has structure, and naming its parts is the first step to managing it.

From many islands to one story: batch genealogy as the key

Structure is necessary but not sufficient. The deeper problem is that the four families and the six sources live in different systems — a Historian, an MES, a LIMS, an EM platform, a genealogy ledger — each with its own database, its own identifiers, and its own idea of what a "batch" is. A purity result in the LIMS and a temperature trace in the Historian describe the same run, but nothing in either system says so unless something deliberately joins them.

The join key is batch genealogy. If every record — sensor tag, QC result, consumed lot, cleanroom sample — carries the same batch identity (the s88.batch anchor above), then the scattered islands become one connected story: this vial came from this lot of drug substance, made in this bioreactor run, fed by these media lots, tested by these methods, in these cleanroom conditions. That single thread is what lets a regulator trace a finished vial back to every ingredient and condition that shaped it, and it is what the companion guide implements as s88.genealogy edges over the s88.batch table in its batch and equipment model. Building and maintaining that thread — the work of contextualization — is the central engineering act of this book.

Why system integration remains unsolved

It would be comforting to say that standards have already solved the connection problem. They have not. This is the genuinely hard, still-open part of the data flow, and it is worth stating plainly.

The standards exist — ISA-88 for batch structure, ISA-95 for the plant-information hierarchy, OPC UA for vendor-neutral exchange — but the systems that actually hold the data were mostly built before those standards were universal, or by vendors with little incentive to interoperate. A vendor Historian, a LIMS, an EBR/MES, and a chromatography data system (CDS) frequently do not speak ISA-88 or OPC UA natively. They expose proprietary APIs, export idiosyncratic file formats, and disagree on something as basic as how a batch is named. Stitching them into one genealogy is therefore still, in most facilities, a manual or custom-coded effort — a brittle web of point-to-point integrations and reconciliation spreadsheets that breaks whenever a system is upgraded. ISPE's GAMP 5 devotes its discussion of interfaces and system integration to exactly this risk: data crossing a boundary between two computerized systems is where lineage is most easily lost, and where validation effort must concentrate [10]. The aspiration of this book — one connected, queryable batch story — is real, but reaching it across a multi-vendor plant remains an unsolved, labor-intensive problem rather than a finished one.

Why "manage" is the operative word

We could have called this book Recording Data in Biomanufacturing. We did not, because recording is the easy part. The hard part is everything the verb manage implies.

The data must be captured at the moment of creation, faithfully and contemporaneously. It must be contextualized — joined to the metadata that turns a number into a fact. It must be protected so it cannot be silently altered or lost, satisfying the integrity and electronic-records rules above [5][6]. It must be connected across the many systems that hold its fragments, so the genealogy of a vial can actually be traced. And it must be retained — the regulatory minimum under 21 CFR Part 211 is one year after the batch's expiration date [4], but firms in practice keep records for a decade or more, because a safety question can surface long after the medicine has shipped.

That is a tall order, and the industry increasingly treats process data not as a regulatory burden but as a genuine asset — fuel for analytics, process understanding, and the predictive models that are reshaping the field [7][8]. A well-managed data shadow does not just defend a batch in an audit; it teaches you how to make the next batch better — the analytics and predictive models we develop in From Data to Knowledge: SPC, Multivariate Analysis, and Continued Process Verification and Machine Learning, Soft Sensors, and Hybrid Models.

Why the batch anchor is also a modeling discipline

The same s88.batch identity that lets an auditor trace a vial is what keeps a model honest, and it is worth previewing why, because the two uses of the data are deeply linked. A naive analyst pools every sensor row across all batches, shuffles them, and splits the pile into a training and a test set — and is delighted by the result. The delight is data leakage: rows from one batch land in both halves, so the model is effectively tested on data it has already seen, and the score is fantasy. The fix is to split by the batch anchor — a grouped, leave-one-batch-out cross-validation (the validation scheme that holds out whole batches, never individual rows, so the test batch is genuinely novel), which is only possible if every row already carries its s88.batch key. Good data management is therefore a precondition for trustworthy ML, not a separate concern; the companion Machine Learning & AI for Biomanufacturing book makes this the foundation of how it splits its data and validates its models.

Two further ideas the data shadow seeds, both developed in that book. An applicability domain is the input region a model was actually trained on; a prediction made outside it — a feed rate or temperature the model never saw — is extrapolation, and the metadata that records each batch's operating window is exactly what lets a model flag "I am being asked about a batch unlike any I learned from." And model drift — a model going quietly stale as the living process and its hardware move (a probe fouls, a media lot changes, the cell line adapts over passages) — must be told apart from genuine process drift; distinguishing the two needs the contextualized, timestamped shadow as its reference, and is the heart of the MLOps lifecycle chapter. Crucially, a model that helps make a medicine is itself a record under the same rules as the rest of the shadow: under GMP it is locked (frozen and version-pinned, never silently relearning) and only ever changed under formal change control, exactly the contemporaneous, attributable discipline this chapter demands of every other record.

Why it matters

If you take one idea from this chapter, take this: the data shadow is as essential to the product as the molecule. Lose the molecule and you lose one batch. Lose the data — or fail to make it trustworthy — and you can no longer prove that any of your batches are what you claim, which can halt a product entirely. Regulators do not ask to taste the medicine; they ask to see its records [4][5]. The shadow is how a biologic proves it is itself.

In the real world

Filled glass vials of a drug product on a production line

Filled vials of drug product. Behind every vial stands a far larger volume of data proving how it was made and tested.

Vials. Image by CSIRO, https://commons.wikimedia.org/wiki/File:CSIRO_ScienceImage_11474_Vials.jpg, licensed under CC BY 3.0 (https://creativecommons.org/licenses/by/3.0/), via Wikimedia Commons.

The data shadow grows even larger as the industry modernizes. The field is advancing toward continuous and intensified processing, where cells produce nonstop rather than in a single fed-batch tank (a vessel that is filled, run once to completion, then emptied). Continuous processing means continuous data — there is no neat "end of batch" moment, so the sensor streams never stop and the need for real-time PAT measurement and live integrity becomes acute [3]. On a real plant floor this data follows a well-worn path: probes on the bioreactor feed a control system — on a large plant a DCS (Distributed Control System, such as Emerson DeltaV or Siemens SIMATIC PCS 7), or, on smaller mobile equipment units (skids — the pre-assembled, self-contained units rolled onto the floor), a PLC (Programmable Logic Controller) paired with SCADA (Supervisory Control and Data Acquisition) software such as Siemens WinCC, the screens operators watch and steer the run from. The control-system layer is the subject of the Automation and Process Control Data chapter. That control system streams tagged readings into a Historian (the AVEVA PI System, formerly OSIsoft PI) and surfaces them in the MES batch record (Körber PAS-X, Siemens Opcenter), while the QC lab's chromatographs and mass spectrometers deposit their raw files into a chromatography data system (CDS) such as Waters Empower or Thermo Chromeleon. At the same time, frameworks like GAMP 5 push the field toward a risk-based, lifecycle view of the very computer systems that capture all this data [10]. The shadow is not shrinking. Managing it well is becoming the defining engineering challenge of modern biomanufacturing. The companion implementation guide assembles exactly this path — sensors to Historian to MES to LIMS — into one open-source reference architecture, so you can see the whole data shadow wired together in working code.

Key terms

  • Data shadow — the full body of records a batch generates: sensor traces, batch records, test results, and signatures; as essential to the product as the molecule.
  • Biologic — a large, complex protein medicine made by living cells rather than by chemistry alone.
  • Monoclonal antibody (mAb) — a Y-shaped protein, every copy identical, engineered to lock onto one specific target.
  • Drug substance — the purified bulk antibody, before it is filled into vials.
  • Drug product — the finished, filled vials of medicine, produced from drug substance in the fill/finish step.
  • The product is the process — the principle that how a biologic is made helps define what it is.
  • Quality by Design (QbD) — building quality into a process deliberately by understanding which parameters and attributes matter.
  • Bioreactor — the vessel in which living cells are grown to make the product.
  • Fed-batch — the conventional batch mode in which a bioreactor is filled, run once to completion, then emptied (contrasted with continuous processing).
  • Electronic batch record (EBR) — the as-executed digital account of how one batch was made, held in a Manufacturing Execution System (MES).
  • Historian — a database built to store streams of timestamped process readings (tags) from instruments and control systems.
  • LIMS (Laboratory Information Management System) — the database that holds QC sample registrations, test results, and specifications.
  • Environmental monitoring (EM) — the continuous watch on cleanroom conditions (airborne particles, microbial counts, temperature, humidity) that proves the surroundings stayed clean while the batch was open.
  • Post-translational modification (PTM) — a chemical edit a living cell makes to a protein after building it; glycosylation, the attaching of sugar chains, is the key example for antibodies.
  • PAT (Process Analytical Technology) — an FDA framework for measuring critical quality and process attributes during manufacture rather than only after.
  • cGMP (current Good Manufacturing Practice) — the regulations governing how medicines must be manufactured.
  • ALCOA — the five data-integrity qualities records must have: Attributable, Legible, Contemporaneous, Original, Accurate. Regulators now extend this to ALCOA+, adding four further qualities (Complete, Consistent, Enduring, Available); the full treatment appears in Data Integrity and ALCOA+.
  • Audit trail — a secure log of who changed what data, and when.
  • Critical Process Parameter (CPP) — a process setting that, if it varies too much, changes the product.
  • Critical Quality Attribute (CQA) — a property of the product that must stay within limits for it to be safe and effective.
  • Metadata — data about data: units, timestamps, instruments, operators; what turns a bare number into a fact.
  • Master data — stable reference information shared across batches: recipes, specifications, equipment registers.
  • Material genealogy — the traceable lineage linking a finished vial back to every ingredient and consumable that touched it.
  • The management pipeline (capture, contextualize, protect, connect, retain) — the five verbs that turn a pile of records into a trustworthy, connected data shadow; each is developed in a later chapter.
  • ISA-88 batch entity (s88.batch) — the single record that anchors a batch's identity (batch_id, recipe_id, unit_id, product_id, status, lot) and that every data source carries as its join key; in the companion guide it is a real database table.
  • System integration — the work of joining records across separate computerized systems (Historian, MES, LIMS, CDS); the boundary between two systems is where data lineage is most easily lost, and the part of the data flow that remains stubbornly manual or custom-coded.
  • Log-reduction value (LRV) — the factor by which a downstream viral-clearance step (a low-pH hold, a viral filter) cuts the virus count; a validated number cast as data by each clearance step.
  • Knowledge graph / triple — an alternative to relational rows in which each fact is a subject-predicate-object triple and the s88.batch join key becomes a walkable edge; the form Book 4 uses for the same batch identity.
  • SHACL (Shapes Constraint Language) — a way to gate graph data with required-structure rules ("every released lot must carry one in-range monomer result and a signature, or it fails"), the graph analogue of a database constraint.
  • SPARQL — the query language for RDF graphs; the way a lineage question ("which lots share a failed batch's cell-bank ancestor?") becomes a one-line traversal up the genealogy spine.
  • Grouped (leave-one-batch-out) cross-validation — splitting model training and testing by the batch anchor rather than by individual rows, so a model is never tested on data from a batch it trained on; the guard against data leakage, and a reason the batch key matters for analytics as well as audit.
  • Applicability domain — the input region a model was actually trained on; a prediction outside it is extrapolation, flaggable only if each batch's operating window was recorded as metadata.
  • Model drift — a predictive model going quietly stale as the living process and its hardware move; distinct from genuine process drift, and detectable only against the contextualized, timestamped data shadow.
  • Locked model — a deployed model frozen and version-pinned, never silently relearning, changed only under formal change control; the same contemporaneous, attributable discipline applied to a model as to any other GMP record.
  • Computerized System Validation (CSV) / Computer Software Assurance (CSA) — the discipline of proving, with documented evidence, that a data system does what it should (the V-model's IQ/OQ/PQ rungs), and the FDA-led shift toward risk-based critical thinking over exhaustive paperwork.

Where this leads

We have surveyed the shadow from above and seen its scale, variety, and legal weight. But a shadow is made of individual points, and the only way to truly understand data management is to follow one. In the next chapter, The Lifecycle of a Data Point, we zoom all the way in: we will track a single measurement from the instant a sensor or analyst creates it — through capture, processing, contextualization, review and use, reporting, retention, and disposal. Along the way we will draw the line between raw data and metadata, and confront the idea at the core of everything that follows: data without context is just noise.