Skip to main content

Polishing Chromatography: Trajectory Models and Resin Lifetime

📍 Where we are: Part IV · Downstream, Learned — Chapter 15. The last chapter cleared viruses by a validated log-reduction margin and handed us a pool that is safe but not yet pure enough to be drug substance. Now the product meets its final columns. Polishing is where the separation gets hard — where the impurities left to remove are versions of the antibody itself, differing by a single charge or a stuck-together dimer — and where the pooling decision stops trading product against junk and starts trading product against product.

The capture pool PApool-001 was concentrated and far cleaner than the harvest, and viral safety proved a clearance margin on top. But it still carries impurities that capture cannot touch: charge variants of mAb-A — acidic species (chiefly from deamidation of asparagine and from sialylation) and basic species (chiefly from un-clipped C-terminal lysine), each shifting the molecule's net charge in opposite directions from the main form — aggregates (high-molecular-weight, HMW, species — two or more antibodies stuck together), and fragments (low-molecular-weight, LMW). These are not foreign contaminants; they are the product, slightly wrong. Removing them is the job of one or two polishing columns — typically cation exchange (CEX), anion exchange (AEX), or a mixed-mode resin (the porous beads packed in the column whose surface carries the charged groups the protein binds to) — run in bind-and-elute or flow-through mode. The polishing pool feeds UF/DF and becomes the drug substance DS-001.

Here the pooling decision becomes its sharpest. On a CEX gradient the acidic variants elute first, the main (target) species in the middle, and the basic variants last — so where you cut the peak directly sets the charge-variant composition of the product, which is a release CQA (Critical Quality Attribute — a measured property the drug must meet to be releasable, defined in QC and release). Cut too wide and you keep main species but drag in acidic and basic tails; cut too narrow and you protect purity but throw away yield. This is the same locked-guard-band pattern as capture, but the attribute on the line is now CEX_main_pct — and in our running example that number must land between 60 and 80 percent, with the golden batch (our representative in-spec reference batch, BATCH-2026-001) at 70.686.

The simple version

Imagine sorting a bag of nearly identical coins where a few are very slightly heavier (acidic) or lighter (basic) than the rest, a few have fused together into doubles (aggregates), and a few have chipped into half-coins (fragments). You pour them down a tilted chute and they arrive in order — light, normal, heavy — with the doubles lagging behind. Your job is to hold a bucket under the chute and catch only the stream of normal coins: start the bucket a moment too early and you scoop up light ones, stop too late and you catch heavy ones and doubles. A polishing column is that chute; the "pooling decision" is when to start and stop the bucket. And because the chute itself wears out — the slope flattens after thousands of uses until the coins stop separating cleanly — you also need to know when to replace the chute. This chapter learns both: when to catch, and when the chute is too worn to trust.

What this chapter covers

We start where the maturity actually is. The production-grade computational model of a polishing column is mechanistic, not machine-learned — the same general-rate-model lineage as capture, now applied to mixed-mode and ion-exchange polishing (Cytiva GoSilico, the open-source CADET) — and we attribute it as such. Then we place learning where it earns its keep on this step: trajectory / charge-variant pooling, where a learned-but-locked rule on the in-line gradient sets the cut points that govern CEX_main_pct and clip the HMW tail; separating HMW and LMW, where the size-exclusion (SEC) attributes SEC_HMW_pct and SEC_LMW_pct bound how aggressive the cut must be; and resin lifetime / column-integrity prediction, the slow problem of an asset wearing out, where ML and classical moment analysis decide the cycle-count question — when to repack. The runnable artifact, examples/platform/ml/resin_lifetime.py, models a locked charge-variant pooling window against the running example's real CEX and SEC release values and projects a governed repack cycle from a resin-health trend.

The polishing step, as a set of learnable decisions

A bind-and-elute CEX polishing cycle has the same skeleton as capture — equilibrate, load, wash, elute, strip, clean — but the elution is the whole story. Instead of one sharp product peak releasing on a pH drop, polishing typically runs a shallow salt or pH gradient that pulls the charge variants off the resin in charge order, smearing them across many column volumes (one column volume = the volume of the packed bed itself, the natural yardstick for elution width). The acidic species, holding less positive charge, let go first; the main species in the middle; the basic species, clinging hardest to the negatively charged CEX resin, last. The chromatogram is therefore not a spike to catch but a trajectory to slice, and three decisions sit on that trajectory:

  • Where to cut the front (the lead cut). Start collecting too early and the acidic shoulder is in the pool, pushing CEX_acidic_pct up and CEX_main_pct down. Start later and the pool is purer in main species but you have discarded recoverable product.
  • Where to cut the back (the tail cut). Stop too late and the basic shoulder — and any aggregate riding the tail — enters the pool. Stop earlier and you protect CEX_basic_pct and SEC_HMW_pct at the cost of yield.
  • When the column is too old to separate cleanly. Every cut above assumes the resin still resolves the variants. As the bed ages over its validated cycle life, plate count falls, peaks broaden and tail, and the charge variants stop separating — until the same cut points that gave 70 percent main species start giving 64. The cycle-count decision — when to repack or replace — is a prediction problem in its own right.

Each is a place where a model reads a cheap in-line signal and produces a decision or a number, and each has a bounded, auditable consequence: you can grade the pool a rule chose against the release assay of what it actually collected. That auditability is why downstream ML — here as in capture — has moved into plants as monitoring and locked rules, never as autonomous control of the cut.

A note on what a "cut" physically is. On a real system the lead and tail cuts are not abstract fractions; they are setpoints on a monitored channel — a UV280 absorbance threshold on the rising and falling edge (e.g. start collecting when UV crosses 100 mAU — milli-absorbance units — ascending, stop at 200 mAU descending). It can also be a conductivity value on the salt ramp, a fixed elution volume, or a peak-slope trigger. The model's job is to choose those setpoints, once, at design time. The consequence of the choice — the charge-variant and size composition of what landed in the pool tank — is read out hours later by the slow lab assays. That lag between decision and label is the structural fact that shapes every modeling choice on this step.

The mature tool is mechanistic — and it now reaches polishing

As with capture, the strongest computational model of a polishing column is a set of partial differential equations, not a fitted network. The general rate model plus a steric mass action (or, for mixed-mode, a multi-component) isotherm describes how each charge variant competes for binding sites as the gradient develops, and a calibrated model predicts the elution order, the peak overlap, and therefore the pool composition any pair of cut points would yield — across conditions it was never run at, with the extrapolation that comes from encoding real physics [1]. Concretely, the general rate model couples a column mass balance — a bookkeeping equation that tracks where the protein goes as it is carried down the column by flow (axial convection), smeared out along the way (axial dispersion), and ferried across the liquid film at each bead surface (film mass transfer) — with diffusion into the pores inside each bead (intraparticle pore diffusion), and the steric mass action (SMA) isotherm — Brooks and Cramer's 1992 formulation (an isotherm is just the equation relating how much protein is bound to the resin to how much is free in solution at equilibrium) — sets the binding equilibrium for each species i through its characteristic charge νi, equilibrium constant Ki, and steric factor σi, with the salt counter-ion competing for the same charged ligands. Because acidic, main, and basic variants differ only in net charge, they differ in ν, and that single parameter difference is what spreads them along the salt gradient. Calibrating the model means fitting those per-species parameters from a small set of gradient experiments (Hahn's mixed-mode polishing case fit a steric-mass-action / colloidal-particle-adsorption isotherm from six experiments and then optimized and validated in silico) and then sweeping cut points in software rather than at the bench.

This is production technology in CMC process development: Cytiva's GoSilico (ChromX/DSPX) and the open-source CADET are the exemplars (CADET being the open general-rate-model + SMA solver of record), and a peer-reviewed industrial case modeled a mixed-mode polishing step for an antibody end to end — simulation, optimization, and design-space definition, describing fragments, aggregates, and HCP — entirely mechanistically [2]. It must be attributed correctly: these are mechanistic models, not machine learning. Calling a calibrated general-rate-model twin "AI polishing" is the same category error this book refused for capture. The distinction is not pedantic: a mechanistic twin extrapolates (predicts beyond the conditions it was calibrated on) because it encodes conservation laws and a binding mechanism, so it can predict a load or gradient it has never seen; a data-fit surrogate only interpolates (predicts between conditions it was trained on) within its training envelope and silently fails outside it. Knowing which you have is the difference between a model you can use to define a design space and one you can only use to monitor.

Evidence

Mechanistic chromatography modeling of polishing steps (general rate model + steric mass action / multi-component isotherms) is mechanistic, not ML, and is the most mature deployed computational tool here — Cytiva GoSilico (production, CMC process development) and open-source CADET, with a peer-reviewed mixed-mode antibody-polishing case [1][2] (peer-reviewed-independent and vendor documentation). Any vendor headline (yield uplift, lots saved) is vendor-self-reported and must carry that label; the modeling capability is established, the specific savings are not independently audited. The learned layer this chapter builds sits beside the mechanistic twin — it reads the trace, sets a locked cut, and watches the resin age — it does not replace the physics.

Trajectory modeling and charge-variant pooling

The pooling decision on a polishing gradient is, at its core, the same locked-guard-band rule as capture — collect while a monitored signal stays inside a band, divert otherwise — but the band now sits on a charge-ordered trajectory and trades two product-quality attributes against yield. Three things make it harder than the capture pool.

First, the signal that matters is not always the raw UV. UV280 tells you protein is eluting but not which charge variant. The richest in-line signal for charge-variant resolution is Raman (or, increasingly, multi-angle light scattering for aggregate content), and the most ambitious research in this space used a convolutional neural network on Raman spectra to make charge-variant pooling decisions on a CEX polishing step, reporting R² (a goodness-of-fit score where 1.0 is a perfect fit) between 0.94 and 0.99 for the predicted charge-variant fractions [3]. In detail, the CNN was calibrated on Raman spectra acquired during process-scale cation-exchange chromatography and quantified acidic, main, and basic species plus total protein simultaneously — R² of 0.94 (acidic), 0.99 (main), 0.96 (basic), and 0.99 (total protein) — so the network turns a raw spectrum into the four numbers a pooling rule needs, in line. The convolutional architecture matters because a Raman spectrum is a 1-D signal with locally correlated peaks; convolution learns the band shapes that shift with charge state, where a plain regression on raw intensities would overfit. The classical baseline here is not raw-intensity regression but a chemometric PLS model on preprocessed spectra — scatter-corrected (SNV) and Savitzky-Golay-smoothed, the same stack the soft-sensor chapter runs — and part of what the CNN buys is learning that band-shape preprocessing end-to-end instead of hand-crafting it. That is a genuine deep-learning result on exactly this problem — and it is research, on one separation, not a routine deployment. (Note the contrast with the capture-step Raman work that predicted many quality attributes during Protein A: that one used K-nearest-neighbours, not deep learning, and must not be cited as a deep-learning case [4].)

Second, the rule is learned at design time, then locked. A model trained on historical polishing cycles — each with its CEX charge-variant and SEC aggregate release assays — can learn the lead and tail cut points that best trade CEX_main_pct against recovery while keeping SEC_HMW_pct under its ceiling. But that cut is then frozen and runs as a fixed rule. Letting a model re-draw the cut online, batch to batch, to chase a quality target is precisely the adaptive control of a CQA that draft EU/PIC/S GMP Annex 22 draws its sharpest line against — it requires locked models and a predetermined change-control plan, not live self-modification. The training/validation discipline that earns the lock is specific: split historical cycles so that future cycles and new resin lots sit in the test fold, not just a random shuffle (a random split lets the model memorize a campaign's idiosyncrasies and report an optimistic score); choose the cut by constrained search over the candidate setpoints; then freeze the chosen setpoints, the preprocessing, and the model version under configuration control, and validate the frozen object against held-out batches before it ever runs in GMP. After lock, the only path to a new cut is a documented change-control retrain, not an online update.

Third, two specs are on the line at once. The cut must satisfy the charge-variant spec (CEX main in 60-80 percent) and the size spec (HMW under 3 percent, LMW under 2 percent). Aggregates often ride the trailing edge of the peak, so the tail cut does double duty: it sets CEX_basic_pct and clips HMW. A pooling model that optimizes charge variants while ignoring the aggregate tail can pass CEX and fail SEC.

Grounding it in the running example's real numbers

The polishing pool's job is to take the feed charge-variant composition and enrich it toward the main species. In our running example the released values for the golden batch BATCH-2026-001 are real, read straight from examples/datasets/hplc_results.csv: CEX_main_pct = 70.686, CEX_acidic_pct = 21.551, CEX_basic_pct = 10.452, with SEC_HMW_pct = 1.287 and SEC_LMW_pct = 0.439. Across the six campaign batches, CEX main runs 66.699 to 70.686 and HMW runs 1.086 to 1.719 — comfortably inside both specs, but the out-of-specification (OOS) sibling BATCH-2026-004 failed on a different attribute (host-cell protein — a process impurity carried over from the production cells — at 128.0 ng/mg against a 100.0 ceiling), not on charge variants or aggregates. That is a useful caution the module makes explicit: a polishing pool that is perfectly in spec on the attributes this step controls tells you nothing about an impurity an earlier step let through.

The module's pooling model treats the elution as a charge-ordered gradient (acidic → main → basic) and a locked centre window — keep fractions 0.18 to 0.88 of the peak — that enriches the main species by partly discarding the acidic lead and basic tail. The keep-fractions are the locked rule's effect on each variant; the batch's acidic/main/basic split is the real measured release value, so the pooled purity the window implies is grounded, not invented. The arithmetic is deliberately transparent: each variant's pooled mass is its real feed percentage times the fraction the centre window keeps of it (acidic 0.62, main 0.97, basic 0.55), and the pooled main percentage is that variant's kept mass renormalized over the kept total — pooled_main = 100 · (main·0.97) / (acidic·0.62 + main·0.97 + basic·0.55). This is a gentle polishing enrichment (main rises a few points, not fifteen), which is the realistic regime for a single CEX cut, and it is why the keep-fractions are close to one another rather than a hard binary include/exclude.

Separating HMW and LMW: the size dimension

Charge is one axis; size is the other, and the two are measured by different assays. Size-exclusion chromatography (SEC) is the release assay that reports SEC_HMW_pct (aggregates) and SEC_LMW_pct (fragments), and the polishing step is where aggregate content is most actively controlled — capture concentrates everything, including aggregates, and polishing is the chance to leave them behind. Where the charge variants spread along the gradient, aggregates tend to partition to the peak edges: HMW species, being larger and often stickier, frequently trail the main peak, so the tail cut that protects CEX_basic_pct also clips HMW. Whether HMW actually rides the tail is molecule-specific, though: aggregates usually elute late on CEX, but for some antibodies they co-elute with or precede the main species, which is exactly why a second orthogonal column (AEX flow-through or mixed-mode) is often what removes the aggregate the CEX tail cut cannot. This coupling is why a one-dimensional "maximize main species" objective is wrong — the cut has to satisfy a vector of attributes.

A learned pooling model that respects this is a small multi-objective problem: choose lead and tail cuts to maximize step yield subject to CEX main at least 60, HMW under 3, LMW under 2 percent. Stated as an optimization, it is maximize yield(lead, tail) s.t. the three inequality constraints, which over two scalar decision variables is small enough to solve by grid search or any constrained optimizer — the difficulty is never the optimizer, it is that each candidate's true objective and constraints are only knowable through a slow assay. In the small-data regime of a real campaign — a handful to a few dozen cycles with full release assays — this is not a job for a data-hungry network; it is a constrained optimization over a few cut parameters, ideally informed by the mechanistic twin's prediction of where each species elutes. The honest division of labour mirrors the rest of downstream: physics predicts the elution profile, a learned/locked rule places the cuts, and the release assays grade the result. In practice the cleanest pipeline is hybrid in exactly this sense — the calibrated general-rate-model twin produces a predicted elution profile for each species (including the aggregate tail it was parameterized to describe), the constrained search runs over that simulated profile to place lead and tail, and the real release assays of the first few cycles confirm or correct the placement before the cut is locked.

Resin lifetime and the cycle-count decision

The slow, expensive, genuinely hard problem in polishing is time — the resin is not the same on cycle 200 as on cycle 1. A polishing resin is validated for a finite cycle life, and across that life it degrades in measurable ways: the plate count (column efficiency, N) falls, peak asymmetry (As, tailing) rises, back-pressure creeps up, dynamic binding capacity (how much protein the bed can still hold while flowing) is lost as the ligands — the charged groups on the resin that bind the protein — slowly foul (get coated and blocked), and — most directly for a charge CQA — selectivity and retention drift, so the variants land at slightly different absolute positions than the frozen cut was tuned for. That selectivity shift (not just the broadening modes) is what moves where each variant elutes, and — the consequence that matters — charge-variant resolution degrades until the locked cut points that gave 70 percent main species start giving less, drifting the pool toward its spec edge. A pooling rule that is perfectly tuned today will slowly mistime itself as the bed ages. The decision this forces is the cycle-count decision: when to repack or replace the column, balancing the cost of an expensive resin against the risk of a drifting CQA.

The vocabulary here is concrete and classical. Plate count N = 5.54 · (t_R / w_½)² measures how many theoretical equilibration stages the bed behaves like — higher is sharper — from the retention time t_R and the peak width at half height w_½; equivalently the height equivalent to a theoretical plate, HETP = L/N, where L is the bed length, so a rising HETP means a degrading bed. Asymmetry As = b/a is the ratio of the trailing to leading half-widths measured at (typically) 10 percent peak height; As = 1 is symmetric, and a creeping rise above ~1.2–1.5 signals tailing from channeling, fines, or fouling. These are the numbers an aging resin moves, and they are exactly what the production monitors compute — so the resin-health story is, at bottom, a story about watching N fall and As rise and deciding when their combination crosses a line.

Two families of method address this, and the honest reading is that they cooperate rather than compete.

Classical, production-grade: moment analysis and transition analysis. The most mature deployed approach is not ML at all. At commercial CDMO scale, Samsung Biologics uses moment analysis and direct transition analysis — computing plate height (HETP), transition width, and asymmetry from a pulse or step injection — for near-real-time column-integrity monitoring of large-scale GMP columns across all chromatography steps [5]. Moment analysis is the statistical-moments reading of a tracer pulse: the zeroth moment is the area (mass recovered), the first moment is the mean retention (centre of gravity of the peak), and the second central moment is the variance — peak spread — from which HETP follows directly without assuming a Gaussian peak. Direct transition analysis does the same from the breakthrough/step transition the process already runs, so it needs no extra tracer injection — a real operational advantage at scale. These are deterministic algorithms on the chromatographic profile, and they are production. They answer "is this column still packed and performing within its qualified envelope?" directly from physics, and they are the baseline any ML approach must beat or augment, not replace.

Data-driven monitoring and projection. On top of (or alongside) moment analysis, ML tracks features of the chromatographic profile across cycles to detect aging before it shows up in yield or quality. A pilot study used on-line Process Analytical Technology (PAT) with Principal Component Analysis (PCA) and batch-level modeling to detect Protein A resin aging some 20 to 25 cycles before observable yield decline, with a proposed strategy to extend resin life — a modeled benefit, not a validated GMP outcome [6]. Model-based strategies for resin lifetime optimization and supervision — including PCA/PLS statistical monitoring coupled with a hybrid deterministic lumped-kinetic aging model carrying two explicit aging parameters — and feature-mining of chromatographic profiles to monitor column performance fill out the literature; the feature-mining work (AstraZeneca) mined PCA/PLS, similarity scores, and SOP-MPT features from absorbance profiles and flagged degradation in cases where traditional HETP/asymmetry could not detect the change [7][8]. The pattern is consistent: physics (moment analysis) measures current health; a fitted trend projects when health will cross an action limit; the repack itself happens under governed change control, never as a model silently extending its own validated cycle count.

Hero diagram of learned polishing chromatography: a horizontal CEX elution trajectory running left to right with a shallow salt-or-pH gradient line rising beneath it, the antibody charge variants drawn as three overlapping bell curves in charge order — an acidic shoulder first, a tall main-species peak in the centre, a basic shoulder last — with an HMW aggregate bump riding the trailing edge; a learned-but-locked pooling band over the centre with a lead-cut marker on the rising edge and a tail-cut marker on the falling edge, the collected slice shaded and labelled pooled CEX main 78 percent against the 60 to 80 percent spec; below the trajectory a small multi-objective callout listing the three constraints CEX main at least 60, HMW under 3, LMW under 2; on the right a separate resin-aging panel showing a resin-health index falling roughly linearly over cycle number, a horizontal action-limit line at 0.80, the fitted trend extrapolating to a projected repack cycle, and a note that moment analysis HETP and asymmetry measure current health while the trend projects when to repack; a banner across the top noting the mechanistic mixed-mode twin sits beside this learned layer — physics for the elution, ML for the locked cut and the aging trend. Polishing, learned: a charge-ordered elution trajectory with overlapping acidic, main, and basic variants and an aggregate tail; a learned-but-locked guard band picks the lead and tail cuts that govern CEX main, basic, and HMW at once; and a resin-health trend — anchored by production-grade moment analysis — projects the governed repack cycle. The machine-learning layer sits beside the mechanistic mixed-mode twin, not instead of it. Original diagram by the authors, created with AI assistance.

Building it: locked pooling plus a resin-lifetime projection

The runnable module frames two tasks on the running example's real release data. First, charge-variant pooling: model the elution as a charge-ordered gradient and apply a locked centre window, reporting the pooled CEX_main_pct the window implies against the batch's real feed composition. Second, resin lifetime: fit a resin-health index against cycle number and project the cycle at which it crosses an action limit — the governed repack recommendation — validated as an extrapolation (fit the first 60 percent of cycles, test the tail) rather than an interpolation, because a lifetime projection that has only ever been graded on cycles it was fit on is worthless.

The pooling function is intentionally arithmetic, not a black box: it reads the batch's real charge-variant row, applies the locked per-variant keep fractions, and renormalizes. The lifetime function is a single-feature LinearRegression of a resin-health index on cycle number, fit on the first 60 percent of an illustrative degradation history and scored on the held-out tail with r2_score, then inverted to solve for the cycle where the fitted line crosses the action limit. The health index itself collapses plate count and asymmetry into one number — health = 0.5·(N/N₀) + 0.5·(As₀/As) — so it is 1.0 like-new and falls as N drops and As rises.

# examples/platform/ml/resin_lifetime.py — locked charge-variant pooling + resin lifetime.
import numpy as np
import pandas as pd
from sklearn.linear_model import LinearRegression
from sklearn.metrics import r2_score

CEX_MAIN_LOW, CEX_MAIN_HIGH = 60.0, 80.0 # real CEX_main release spec, %
HMW_HIGH = 3.0 # SEC HMW spec ceiling, %
HEALTH_ACTION_LIMIT = 0.80 # resin-health action limit (illustrative)

def pooling_window(batch_id="BATCH-2026-001"):
"""Locked guard band on a charge-ordered gradient (acidic -> main -> basic).
The cut FRACTIONS are the locked rule; the batch's acidic/main/basic split is
the REAL measured release value, so the pooled purity is grounded."""
res = load_polishing_results().loc[batch_id] # real release row, wide
acidic, main, basic = res.CEX_acidic_pct, res.CEX_main_pct, res.CEX_basic_pct
lead_cut, tail_cut = 0.18, 0.88 # collect 0.18..0.88 (locked)
keep_acidic, keep_main, keep_basic = 0.62, 0.97, 0.55 # window's effect per variant
m_ac, m_mn, m_bs = acidic * keep_acidic, main * keep_main, basic * keep_basic
tot = m_ac + m_mn + m_bs
pooled_main = 100.0 * m_mn / tot
return {"feed_main": round(main, 2), "pooled_main": round(pooled_main, 2),
"yield_frac": round(tot / (acidic + main + basic), 3),
"in_spec": bool(CEX_MAIN_LOW <= pooled_main <= CEX_MAIN_HIGH)}

def _synthetic_cycle_history(n_cycles=200, seed=2026):
"""Illustrative per-cycle health: plate count N decays ~linearly, asymmetry rises.
SHAPE is the real phenomenology; the numbers are illustrative."""
rng = np.random.default_rng(seed)
cyc = np.arange(1, n_cycles + 1)
N = 3200.0 - 4.2 * cyc + rng.normal(0, 35, n_cycles) # plate count falls
As = 1.05 + 0.0016 * cyc + rng.normal(0, 0.01, n_cycles) # asymmetry rises
health = 0.5 * (N / N[0]) + 0.5 * (As[0] / As) # 1=like-new, in [0,1]
return pd.DataFrame({"cycle": cyc, "health_index": health})

def resin_lifetime(action_limit=HEALTH_ACTION_LIMIT):
"""Fit health_index ~ cycle and project the action-limit crossing, validated
by leave-future-out: fit first 60% of cycles, test the tail (extrapolation)."""
hist = _synthetic_cycle_history()
x, y = hist[["cycle"]].to_numpy(float), hist["health_index"].to_numpy(float)
cut = int(0.6 * len(hist))
reg = LinearRegression().fit(x[:cut], y[:cut])
r2_tail = r2_score(y[cut:], reg.predict(x[cut:]))
slope, intercept = float(reg.coef_[0]), float(reg.intercept_)
cross = (action_limit - intercept) / slope # cycle where health hits the limit
return {"slope_per_cycle": round(slope, 6), "extrapolation_r2": round(r2_tail, 3),
"projected_repack_cycle": int(round(cross)), "health_now": round(y[-1], 3)}

Running it on the running example's real release values gives the verified output:

### resin_lifetime.py ###
Real CEX charge-variant + HMW release values (hplc_results.csv):
test CEX_main_pct CEX_acidic_pct CEX_basic_pct SEC_HMW_pct
batch_id
BATCH-2026-001 70.686 21.551 10.452 1.287
BATCH-2026-002 69.085 22.744 9.220 1.247
BATCH-2026-003 70.404 19.508 10.172 1.086
BATCH-2026-004 67.879 21.447 8.289 1.280
BATCH-2026-005 66.699 21.756 9.413 1.169
BATCH-2026-006 69.171 20.927 10.408 1.719

Charge-variant pooling [BATCH-2026-001]: feed CEX_main 70.69% -> pooled 78.2% (acidic 15.24%, basic 6.56%), yield 0.854, in_spec=True
Charge-variant pooling [BATCH-2026-004]: feed CEX_main 67.88% -> pooled 78.67% (acidic 15.89%, basic 5.45%), yield 0.857, in_spec=True

Resin lifetime: slope -0.001317/cycle, extrapolation R2=0.814 (fit first 60% of cycles), health now 0.76 -> projected repack at cycle 156 (action limit 0.8)
ASSERT ok: locked charge-variant window keeps pooled CEX_main inside the 60-80% spec.

Read it as a process engineer would. The locked centre window takes the golden batch's feed of 70.69 percent main species and enriches it to a pooled 78.2 percent — comfortably inside the 60-to-80 spec — at the cost of about 15 percent of the mass (step yield 0.854), with the acidic share of the pool falling from 21.55 to a renormalized 15.24 percent and the basic from 10.45 to 6.56 (these are the pool's composition after renormalizing, not the fraction of each variant that survived the cut). The same locked rule applied to BATCH-2026-004 enriches its lower feed of 67.88 to a pooled 78.67 — a quiet but important point: the locked window is robust enough that the polishing step itself is in spec for the OOS batch, because that batch failed on host-cell protein, an attribute polishing-as-modeled-here does not control. In spec on this step does not clear an upstream impurity.

The resin-lifetime model projects a repack at cycle 156 with a held-out extrapolation R² of 0.814 — honest about being an extrapolation, which is the only kind of projection a lifetime model can be — from a fitted health slope of about -0.0013 per cycle and a current health index of 0.76. Note that 0.76 is already below the 0.80 action limit at the last observed cycle: the trend would have triggered the repack recommendation before the column reached 200 cycles, which is exactly the early-warning behaviour the projection exists to provide. The 156-cycle projection and the 0.80 health action limit are illustrative; the shape (slow plate-count decay plus rising asymmetry, validated leave-future-out) is the real phenomenology, and the 0.814 tail R² is the honest admission that a straight line fit on early cycles only roughly tracks the later ones — good enough to schedule a governed repack with margin, not good enough to ride to the last safe cycle. And what actually governs the repack is not the point crossing (cycle 156) but a conservative lower bound on it: the same logic the viral-safety chapter used to report the lower 95 percent confidence bound of a clearance factor applies here — bootstrap or interval-bound the crossing and repack against the early edge of that interval, never the central estimate.

Anatomy of one pooling-and-lifetime record

A polishing step produces two records that must be read together: the pool it collected this cycle, and the resin health of the column that collected it. A pool that is in spec on a healthy column means one thing; the same pool on a column three cycles from its action limit means something else entirely. The record below ties them.

Identity card unpacking one polishing pooling-and-lifetime record for the CEX step on BATCH-2026-001: a header naming the cycle, the column, and its cycle count; a live-signals block listing UV280, conductivity (the salt gradient), pH, and an optional Raman charge-variant signal as the streamed channels; a learned-but-locked rule block showing the lead cut at 0.18 and tail cut at 0.88 of the peak with a note that the cut was learned at design time and is now frozen; a multi-objective constraint block listing CEX main at least 60, HMW under 3, LMW under 2; a pooled-result block showing feed CEX main 70.69 percent enriched to pooled 78.2 percent with acidic 15.24, basic 6.56, step yield 0.854, all flagged in spec; a size block showing the SEC HMW 1.287 and LMW 0.439 the tail cut protected; a resin-health block showing plate count N and asymmetry As from moment analysis, a health index of 0.76, the 0.80 action limit, and the projected repack at cycle 156 with extrapolation R2 0.814; a lineage row showing the pool derivedFrom PApool-001 and feeding DS-001; and a governance footnote that the cut is locked design-time and the lifetime model is monitoring not autonomous, per draft Annex 22. One polishing record, fully unpacked: the live gradient signals, the learned-but-locked lead and tail cuts, the multi-objective constraints (CEX main, HMW, LMW) the cut must satisfy at once, the pooled charge-variant and size results it produced, and — the field that makes it honest — the resin-health index and projected repack cycle of the column that produced them, with lineage from capture pool to drug substance. Original diagram by the authors, created with AI assistance.

Read field by field, the record is the chapter in miniature.

  • Header — cycle, column, cycle count. Identifies which physical column and how many cycles into its validated life this pool came off. This is the field that turns an in-spec pool from "fine" into "fine, but on a column at cycle N of a finite life," and it is the join key between the pool record and the resin-health record.
  • Live signals — UV280, conductivity, pH, optional Raman. The streamed channels available the instant the cut is made. UV280 (mAU) sees total eluting protein; conductivity is the salt gradient itself, the x-axis the variants spread along; pH matters on a pH-gradient or mixed-mode step; the optional Raman channel is the only one that resolves which charge variant is eluting (the CNN-Raman research result), and where present it feeds the four charge-variant fractions directly.
  • Learned-but-locked rule — lead cut 0.18, tail cut 0.88. The two setpoints that define the centre window, learned at design time on historical cycles and now frozen under configuration control. The "locked" flag is not decoration: it is the field a GMP reviewer checks to confirm the cut is not self-modifying.
  • Multi-objective constraints — CEX main at least 60, HMW under 3, LMW under 2. The vector of release specs the single cut must satisfy simultaneously (the multi-objective reason a single CEX number is never the whole grade). A pool that is excellent on CEX main but over on HMW is still a failed cut.
  • Pooled result — feed CEX main 70.69 → pooled 78.2 percent; acidic 15.24, basic 6.56; step yield 0.854; in spec. What the window actually collected, in charge-variant terms, with the yield cost of the enrichment. The in-spec flag here is the pooling_window() assertion made visible.
  • Size result — SEC HMW 1.287, LMW 0.439. The aggregate and fragment content the tail cut protected, from the orthogonal SEC assay. It appears beside the charge-variant result because the same tail cut sets both; reading them together is the point.
  • Resin-health block — plate count N and asymmetry As (from moment analysis), health index 0.76, action limit 0.80, projected repack cycle 156, extrapolation R² 0.814. The field that makes the record honest. N and As are the production-grade moment-analysis measurements of current health; the health index and projection are the fitted trend's read on future health and the governed-repack recommendation that follows.
  • Lineage — derivedFrom PApool-001, feeds DS-001; pool POLpool-001. The genealogy edge: this polishing pool descends from the capture pool and becomes the drug substance, so a CQA question on DS-001 can be traced straight back through this record — the same derivedFrom genealogy Book 4's instance graph models for this exact chain, asserted backward as DS-001 derivedFrom POLpool-001 derivedFrom PApool-001.
  • Governance footnote — cut locked design-time; lifetime model is monitoring, not autonomous; per draft Annex 22. The compliance statement that the whole record is built to support: the cut does not move online and the lifetime model recommends rather than acts. The record is auditable after the fact against the SEC and CEX release assays — which is exactly why this locked-rule-plus-monitoring pattern passes GMP review where an autonomous cut-redrawing controller would not.

The ontology underneath: why a feature pulled by IRI is a feature you can trust

Every number this chapter consumes — the feed CEX_main_pct, the SEC_HMW_pct ceiling, the cycle count, the derivedFrom parent — is a feature the pooling and lifetime models read, and the quiet question no ML chapter usually asks is what gives that feature its meaning. The honest answer is a vocabulary, not a column header. In a brittle pipeline the model joins on the string "CEX_main_pct"; rename the column in one LIMS export and the join silently returns nulls, and a pooling rule fed nulls makes a confident, wrong cut. The robust alternative is to pull each feature by its IRI (Internationalized Resource Identifier — a globally unique web name for a concept, not a local key): the charge-variant feature is bp:cexMainPct, a datatype property pinned to a drug-substance material in Book 4's relations chapter, so the same number means the same thing to the LIMS that wrote it, the historian that trended it, and the model that read it — the identifier discipline that makes data FAIR. That is the difference between a feature that happens to be named a thing and one that is that thing.

The deeper payoff is that the release gate validates the training data, not just the lot. Book 4's bp:ReleaseShape is a SHACL (Shapes Constraint Language — a closed-world validator that fails when a required result is missing, where a reasoner would shrug it off as unknown) shape that targets every released material and enforces exactly the panel this chapter trades against: bp:cexMainPct present, singular, and inside 60.0–80.0; bp:hmwPct at or below 2.0. Run that same shape over the rows that feed pooling_window() and resin_lifetime() and you get, for free, the guarantee a training set most often lacks — that every label is present, typed xsd:float, and in its physically possible range — before the model ever fits on it. A pooling model trained on a row where the CEX result was dropped by a LIMS integration, or filed twice under two IRIs, is a model trained on a lie; the release shape catches that at the seam between systems, which is precisely where required results go missing. The gate that decides whether DS-001 may reach a patient is the same gate that decides whether the data was clean enough to learn from.

Two more semantic distinctions earn their keep here. First, the BFO continuant/occurrent cut keeps a measurement (the CEX_main_pct quality that inheres in the pool) distinct from the run (the polishing cycle, an occurrent that happened once and is gone) and from the column (equipment that persists across hundreds of cycles) — three different aisles, so the resin-health record can attach a cycle count to the column, a duration to the run, and a charge-variant value to the pool without ever collapsing them into one fuzzy node. Second, the bp:derivedFrom lineage spine — asserted DS-001 derivedFrom POLpool-001 derivedFrom PApool-001 — is not just a traceability nicety; it is the natural grouping key for honest cross-validation. Cycles that share a campaign, a resin lot, or a parent capture pool are not independent draws, so a leave-one-batch-out (or leave-one-lot-out) split that groups on the lineage edge is what stops the optimistic score a random shuffle would report — the same non-independence the training-data chapter warns about, now expressed as a graph traversal rather than a hand-maintained batch list. And when a GraphRAG assistant is asked what did DS-001 derive from, and was its charge profile in spec — the knowledge-graph grounding the ontology book closes on — it answers by walking those typed bp:derivedFrom edges and citing the bp:cexMainPct value the SHACL gate already validated, so the model narrates a true lineage instead of inventing a plausible one.

None of this is new modeling machinery on top of the ML — it is the ML resting on a vocabulary that already exists. The features have IRIs, the labels pass a shape, the lineage is a typed edge, and the measurement is not the run. That is what makes the locked cut and the projected repack trustworthy rather than merely computed.

The data discipline underneath: an auditable training set

The same record that grades a cut is, under GMP, a regulated artifact, and that constrains the data this chapter learns from more than the algorithm does. Each CEX and SEC release value carried into hplc_results.csv has a data shadow — the governed trail every batch casts — that must be ALCOA+ (Attributable, Legible, Contemporaneous, Original, Accurate, plus Complete, Consistent, Enduring, Available — the data-integrity attributes a regulated record must satisfy) and 21 CFR Part 11 compliant: the result is attributable to the analyst and instrument that produced it, time-stamped contemporaneously, and signed. A pooling model is only as trustworthy as that provenance, because a label whose origin cannot be reconstructed cannot anchor a validated decision — the decision-versus-label lag this chapter keeps flagging is exactly the window where an un-attributed or back-dated value would poison the training set. And the row itself is not free-floating: in an ISA-95 / B2MML-structured plant the result ties to a batch, a material lot, and a unit operation with their units and timestamps explicit, so the model's input carries its manufacturing context rather than a bare float. That grounding is what lets a reviewer trust that the 70.686 the model read is the same 70.686 Quality released BATCH-2026-001 on.

The unsolved part: resolution decay, small data, and the locked-model paradox

The hard, unsolved problem here is the same shape as capture's, sharpened. The column ages, its resolution decays, and a frozen cut slowly drifts the pool toward its spec edge — but the very fix a data scientist wants, let the cut adapt to the aging resin, is the thing draft Annex 22 most explicitly forbids for a CQA-affecting function. This is the locked-model paradox in its purest form: the conditions that make adaptation valuable (a slowly non-stationary process) are exactly the conditions a static, validated model is least suited to, yet a static, validated model is the only kind permitted on the critical path. The field is left in an awkward middle ground: detect the decay well (moment analysis and aging monitors are good at this), but respond through governed retraining and validated repack, not through a model quietly rewriting its own cut. Knowing when a model has decayed enough to warrant a controlled retrain — and proving the new cut is at least as safe — is the open MLOps problem of downstream, and it is genuinely unsolved at scale. The deepest difficulty is that the trigger for retraining (drift in resolution) and the constraint on retraining (revalidate everything before the new cut can run) operate on different clocks: drift is continuous, revalidation is discrete and slow, so there is always a window in which the plant knows the cut is mistiming and is not permitted to fix it live. Managing that window — with margin in the original cut, conservative action limits, and a pre-authorized change-control path — is the real engineering, and it is closer to operations discipline than to machine learning.

Two further difficulties are specific to polishing. First, small data bites hardest exactly here. The labels that train a charge-variant pooling model are full SEC + CEX release assays, which arrive once or twice per batch from a slow lab — the cold-start, sparse-reference reality the whole book keeps meeting. A campaign of a few dozen cycles, each with one charge-variant readout, is not enough to fit a data-hungry model that generalizes to a new resin lot, a new scale, or a feed shifted by an upstream change. Worse, the labels are not independent draws from a fixed distribution: cycle N+1's resin is slightly older than cycle N's, so the very data you would use to fit a pooling rule is itself drifting under you, which means a model that fits the campaign's average will mistime both its early healthy cycles and its late degraded ones. This is the canonical case for a hybrid approach — the mechanistic twin predicts the elution profile, the data only places the cut — and the honest verdict is that pure-ML charge-variant pooling remains research, not routine [3].

Second, the resin-health index itself is a modeling choice, not a measurement. Our module collapses plate count and asymmetry into a single 0-to-1 number with a 50/50 weighting; a real action limit and weighting must be justified against the attribute that actually drifts (here, charge-variant resolution), validated, and tied to the moment-analysis envelope the production monitors already use. And the baseline is itself a choice: normalizing against a single like-new cycle (as the module does, for transparency, with N₀ carrying its own σ=35 measurement noise and As₀ its σ=0.01) passes that one reading's noise into every later health value — a real validation would baseline against an averaged set of qualified early cycles. The 50/50 weighting is a placeholder for a real question — does this column degrade first in plate count or first in asymmetry, and which of those better predicts the day CEX_main_pct will leave spec? — that can only be answered by regressing the health features against the actual charge-variant resolution over a column's life, which loops straight back into the small-data problem. A health index that is not anchored to a real quality consequence is a dashboard number, not a decision criterion — and the gap between "the index crossed 0.80" and "the pool will drift out of spec at cycle 160" is the same validation work that separates a soft sensor from a release decision everywhere in this book.

What this chapter adds to the model suite

This chapter contributes examples/platform/ml/resin_lifetime.py to Book 5's example suite as a standalone runnable module — run it directly with python resin_lifetime.py (it is one of the suite's runnable modules rather than one of the twenty-one gated by the run_all.py harness, but it guards its own claim with an assert all the same) — the polishing pooling-and-lifetime model, anchored to the running example's real CEX and SEC release values. It provides:

  • load_polishing_results() — pivots the real charge-variant (CEX_main/acidic/basic_pct) and aggregate (SEC_HMW_pct) release values from examples/datasets/hplc_results.csv into one row per batch across all six campaign batches, including the OOS BATCH-2026-004.
  • pooling_window() — the locked charge-variant pooling rule: a centre window on a charge-ordered gradient that enriches CEX_main_pct toward spec while clipping the acidic lead and basic tail, reporting the pooled composition, step yield, and an in-spec flag against the real 60-to-80 percent CEX main spec. The pooled-purity assertion guards the claim so it cannot silently rot.
  • resin_lifetime() — the cycle-count model: fits a resin-health index against cycle number and projects the governed repack cycle, validated leave-future-out (fit the first 60 percent of cycles, test the tail) so the projection is honestly an extrapolation, with the slope, extrapolation R², and projected repack cycle reported.

It deliberately complements rather than duplicates the capture module: where chromatography.py classifies phases and times a single sharp peak with a recovery soft sensor, resin_lifetime.py works the polishing problem — a charge-ordered trajectory, a multi-objective cut, and the slow asset-aging decision the capture chapter named as unsolved.

Like the rest of the suite it is open and reproducible by construction: it depends only on the permissively-licensed open-source stack — NumPy, pandas, and scikit-learn's LinearRegression/r2_score — pinned in the suite's environment, with no proprietary solver in the loop (the mechanistic twin it defers to has an open exemplar of its own in CADET). The pooling rule reads the committed examples/datasets/hplc_results.csv so anyone can run it against the real release values, and the lifetime model's only synthetic input is seeded (seed=2026), so the projected repack at cycle 156 and the 0.814 tail R² reproduce byte-for-byte on every machine — the same determinism the open-source analytics chapter relies on when its SPC example trends this very CEX_main_pct attribute. That is the bar the whole companion stack holds: a claim you cannot re-run is a claim you cannot check.

Why it matters

Polishing is the last place to fix the product before it becomes drug substance, and its pooling decision is one of the few in manufacturing where a cut directly sets a release CQA — CEX_main_pct — live, on a moving signal. Getting the learning layer right means three concrete things: a pooling rule that is learned and validated but locked, placing the lead and tail cuts on evidence rather than a fixed historical window while satisfying the charge-variant and size specs at once; a clear deference to the mechanistic mixed-mode twin for the elution physics, so neither method pretends to be the other; and a resin-lifetime model that turns the expensive, risky repack decision from a fixed calendar rule into a monitored, projected, governed one — anchored by production-grade moment analysis, not floating on a black-box health score. Get the boundary right and polishing becomes a model citizen of downstream ML: real, deployed-adjacent, and honest about its small-data ceiling and its locked-model constraint. Blur it — let a model redraw its own cut to chase a target, or trust a health index nobody validated against a quality consequence — and you have stepped over the line the regulators have drawn brightest, on the very step that decides whether the drug substance is in spec.

In the real world

Mechanistic modeling of polishing is production technology in CMC: Cytiva GoSilico and open-source CADET model mixed-mode and ion-exchange polishing for real molecules, with a peer-reviewed mixed-mode antibody-polishing case in the literature [1][2] — mechanistic, not ML, with vendor savings figures vendor-self-reported. On the column-integrity side, the deployed reality is algorithmic, not deep learning: Samsung Biologics runs moment analysis and direct transition analysis (HETP, transition width, asymmetry) for near-real-time integrity monitoring of production-scale columns (production) [5]. On the learning side, the deployed-adjacent reality is monitoring and prediction rather than autonomous control: on-line PAT plus PCA has flagged Protein A resin aging 20 to 25 cycles ahead of yield loss (pilot) [6], model-based resin-lifetime supervision and chromatographic-profile feature mining are active research [7][8], and CNN-guided Raman for charge-variant pooling is a striking but still research result (R² 0.94 to 0.99) [3]. The pattern matches the ISPE Pharma 4.0 picture the whole book reports: downstream ML clusters in monitoring, prediction, and human-in-the-loop review — not in autonomous control of a critical quality attribute. The open-source analytics chapter shows the same shape of model running in code (its SPC example trends this very CEX_main_pct attribute), and Book 1's purification chapters and Book 4's downstream ontology describe the same physical step through their own lenses.

Key terms

  • CQA (Critical Quality Attribute) — a measured property the product must meet its specification on to be released; here the charge-variant composition CEX_main_pct, which the pooling cut directly sets (defined fully in QC and release).
  • Polishing chromatography — the final purification column(s) (CEX, AEX, or mixed-mode) that remove product-related impurities capture cannot: charge variants, aggregates, and fragments; yields the pool that feeds UF/DF and becomes drug substance.
  • Charge variants — versions of the antibody differing by charge (acidic, main, basic), measured by CEX as CEX_acidic/main/basic_pct; on a CEX gradient they elute in charge order, so the pooling cut sets their composition.
  • Trajectory / charge-variant pooling — choosing the lead and tail cuts on a charge-ordered elution to set the pool's charge-variant composition; the cut is learned and validated at design time but locked in production.
  • HMW / LMW (aggregates / fragments) — high- and low-molecular-weight species measured by SEC as SEC_HMW_pct and SEC_LMW_pct; aggregates often ride the peak tail, so the tail cut clips them, coupling the size and charge specs.
  • Multi-objective cut — the polishing cut must satisfy the CEX main, HMW, and LMW specs simultaneously, so a one-dimensional "maximize main species" objective is wrong.
  • Steric mass action (SMA) — the ion-exchange binding isotherm (Brooks and Cramer, 1992) that gives each species a characteristic charge, equilibrium constant, and steric factor; the per-species charge difference is what spreads the variants along the salt gradient in a mechanistic twin.
  • Mechanistic polishing model — a first-principles column simulator (general rate model + steric mass action / multi-component isotherm; GoSilico, CADET); the mature production tool here, and not machine learning.
  • Resin lifetime / cycle-count decision — when to repack or replace an aging column; the resin's plate count falls and asymmetry rises over its validated cycle life until charge-variant resolution decays.
  • Plate count (N) / HETP / asymmetry (As)N = 5.54·(t_R/w_½)², HETP = L/N, and As = b/a; the classical efficiency, plate-height, and peak-tailing measures that an aging bed moves and that moment analysis computes.
  • Moment analysis / transition analysis — deterministic computation of plate height (HETP), transition width, and asymmetry from a pulse or step (the statistical moments of the profile); the production-grade, non-ML column-integrity monitor that ML aging models augment, not replace.
  • Resin-health index — a modeling choice that collapses plate count and asymmetry into a single number against an action limit; only a decision criterion once validated against the quality attribute that actually drifts.
  • IRI / feature by IRI — an Internationalized Resource Identifier is a globally unique web name for a concept (e.g. bp:cexMainPct); pulling a feature by its IRI rather than a column-name string is what makes the model's input mean the same thing across LIMS, historian, and code (the identifier discipline).
  • SHACL release shape — the same closed-world bp:ReleaseShape gate that decides whether DS-001 may be released also validates the model's training rows are present, typed, and in-range; the gate that clears a lot is the gate that clears the data (the release gate and SHACL).
  • derivedFrom lineage as grouping key — the typed bp:derivedFrom edge (DS-001 derivedFrom POLpool-001 derivedFrom PApool-001) marks which cycles share a lot or parent pool, so it is the natural leave-one-batch-out split that keeps cross-validation honest (relations and genealogy).
  • ALCOA+ / 21 CFR Part 11 — the data-integrity attributes (Attributable, Legible, Contemporaneous, Original, Accurate, plus Complete, Consistent, Enduring, Available) and the regulation that make a release value an auditable record; a label without that provenance cannot anchor a validated decision (the data shadow).

Where this leads

The polishing pool is now in spec on charge variants and aggregates — the product is as pure as chromatography can make it, but it is still dilute and in the wrong buffer to be a drug. The next chapter, UF/DF and Drug Substance: Soft-Sensing Concentration and Excipients, takes up the final downstream step — concentrating the protein and exchanging it into its formulation buffer — and the soft-sensing problem of knowing the protein concentration and excipient levels in real time, as the pool DS-001 finally earns the name drug substance.