(Lecture 11.) Last week asked how a control strategy is held once it’s running — and a running in-line model is a prediction from a model over a spectrum. This week is where that promise gets cashed in: first the rules for when a computational result can be trusted at all, then the mathematics that actually builds the model. It is also the last full lecture of the term — final paper presentations run alongside it.
The one idea
A predictive model is a hypothesis about data in exactly the way a method is a hypothesis about a molecule — fit on the evidence you have, valid only within its range, tested by every new sample. Chemometrics is machine learning; it just had a 40-year head start in regulated analytical science, so it carries things generic ML underplays: interpretability, small-n honesty, and a validation tradition. Almost everything below is a safeguard against building a model that fits your data instead of one that predicts new data.
On teaching this honestly. Nobody in the department specialises in machine learning, and this session does not pretend otherwise. The goal is not to make you model builders; it is to make you competent clients and reviewers — able to say what a model must do, what data it needs, whether its output can be trusted, and where the regulations draw the line.
Part 1 — Miscellaneous methods: the rules for a trustworthy computational result
Three things called “AI”
| Kind | What it is | Where it shows up | How it’s governed |
|---|---|---|---|
| Predictive ML / chemometrics | A model mapping inputs to a number or class, fit on labelled data | NIR/Raman calibration, peak detection, predictive stability, (Q)SAR | Like an analytical method — Q2, Q14, lifecycle |
| Generative / LLMs | A model that produces fluent text (or code, images) from a prompt; non-deterministic | Drafting narratives, extracting data from legacy PDFs, literature triage, code assistance | Emerging AI guidance plus the existing data-integrity, Part 11, and CSV floor; a human verifies every factual output |
| Agentic systems | An LLM given tools and allowed to act in a loop | Early pilots — automated investigation triage, lab-system orchestration | Least mature; treated as a computerised system with a human decision-maker in the loop |
The regulatory landscape (early 2026)
There is not yet a binding, AI-specific regulation for pharmaceutical analysis — there is a fast-forming framework sitting on a floor that already applies: the EU AI Act (horizontal, risk-tiered, phasing in through 2026–2027); FDA draft guidance (2025) on AI to support regulatory decisions, with a risk-based credibility framework (evidence scales with model influence × decision consequence); FDA discussion papers on AI in manufacturing; an ICH reflection paper signalling AI will be addressed within the existing quality framework, not a separate track; and ISPE GAMP / PDA guidance. The floor underneath all of it is already binding: GMP data integrity (ALCOA+), 21 CFR Part 11 / Annex 11, computerised-system validation, and Q9 quality risk management. The through-line: credibility proportionate to consequence.
Where a language model can and cannot sit
| Can (with human verification) | Cannot |
|---|---|
| Draft an OOS-investigation narrative from analyst notes | Decide the OOS outcome, or state a root cause as fact |
| Extract structured data from legacy CoAs, reports, PDFs | Be the sole record of that data — the extraction is verified against the source |
| Triage literature; summarise a method-transfer report | Contribute an uncited claim to a regulatory document |
| Assist with chemometrics / analysis code | Run unreviewed code that produces a reportable result |
The hard line is data integrity: an LLM output is not deterministic, not inherently traceable to a source, and can be confidently wrong. Anything that becomes a GMP record or informs a GMP decision must be verified and attributable to a person.
Worked case — where a prediction already replaces an experiment
ICH M7 (assessment of DNA-reactive impurities) is the clearest example of a computational prediction being formally accepted in lieu of data. For a new impurity, M7 allows a mutagenicity conclusion drawn from two complementary (Q)SAR systems — one expert-rule-based, one statistical — to substitute for an Ames test. If both predict non-mutagenic and there is no conflicting knowledge, no bacterial assay is run: two orthogonal models (the same instinct as orthogonal analytical methods), a defined scope, documented and versioned, with expert review on top. Predictive stability (ASAP) is on the same trajectory, given a formal home in the modernized Q1 annex — as is the dissolution IVIVC biowaiver: an in-vitro model standing in for a clinical study.
Part 2 — The method: chemometrics and machine learning
One continuum
| Linear latent-variable (classic chemometrics) | Nonlinear / modern ML | |
|---|---|---|
| Methods | PCA, PLS, PLS-DA, MCR, PCR | Random forests, gradient boosting, SVM, neural nets, deep learning on raw spectra |
| Best when | Relationships are roughly linear; n is small; you must explain the model | Genuinely nonlinear response; complex matrices; image or high-dimensional data; n is large |
| Regulated setting | The default — transparent, established | Used where it clearly wins, with extra credibility evidence |
The regulated setting pushes toward the transparent end, but the discipline — training/test separation, applicability domain, drift monitoring, reproducibility — is identical across the continuum.
Preprocessing — part of the method, not a tidy-up
| Method | Removes | Note |
|---|---|---|
| Mean-centering | The common offset — always done | |
| Scaling (autoscale, Pareto) | Differences in variable magnitude | Autoscale gives every variable equal weight — powerful and dangerous |
| SNV / MSC | Multiplicative scatter, path-length variation | The default for diffuse-reflectance NIR |
| Savitzky–Golay derivatives | Baseline slope (1st) and offset+slope (2nd) | Amplifies noise — needs smoothing |
Preprocessing choices are locked with the model and revalidated if changed — as much part of the method as the column and mobile phase in an HPLC method.
PCA, PLS, and MCR
- PCA re-expresses many correlated variables as a few uncorrelated components: scores show where each sample sits (clustering by batch, site, season); loadings show how wavelengths combine into each component; Hotelling’s T² and Q-residual are the outlier and out-of-domain detectors a running PAT model depends on.
- PLS regresses spectra against a reference value through a few latent variables built to be relevant to y. The one place people cheat: too many latent variables fits noise, chosen by cross-validation that must reflect how the model will actually be used — RMSECV and especially RMSEP (an independent test set) are what count, not RMSEC.
- MCR-ALS resolves a matrix of mixed spectra into pure-component spectra and concentration profiles with minimal assumptions — the catch is rotational ambiguity, resolved only by chemical judgment about which constraints (non-negativity, unimodality) to apply.
Design of experiments
Factorial, fractional-factorial, response-surface, and D-optimal designs map how several factors jointly affect a response with far fewer runs than one-factor-at-a-time, revealing interactions. This is the machinery behind the analytical robustness study and MODR and the manufacturing design space (Q8).
Worked case — the model that passed cross-validation and failed in production
A PLS model for tablet assay by NIR: 6 latent variables, RMSECV 0.9%, R² 0.99. In production it was biased 2–3% and trended with batch. The cause: the calibration set had 3 replicate spectra per tablet, and cross-validation left out random spectra, not whole tablets or batches — information leakage between training and test, generic to ML, not just chemometrics. Re-run with batch-blocked cross-validation, the honest RMSECV was 2.1% and the model needed rebuilding with a wider set. The model fit beautifully and predicted badly, and only a validation design that mirrored real use revealed the gap.
The notebook thread
The shared Python/Jupyter exercise: load real spectral data — the same in-line NIR dataset from last week’s continuous-manufacturing case — preprocess it, fit PLS and a tree ensemble, and build the validation that tells them apart, then show what an out-of-domain sample does to the prediction and to the T²/Q diagnostics that are supposed to catch it.
Where the analyst sits
For predictive models the analyst states what the model must do in measurable terms, owns the quality of the training data, and decides per sample whether it is in the applicability domain. For generative tools the analyst is the verifier. Neither is machine-learning expertise; both are analytical judgment, the STEAM “M” and “A” at once — and it is why Week 1 said every scientist is now a data scientist. The refrain: science → evidence → reduced uncertainty → control → regulatory confidence → patient trust.
On the job
- Ask, for any AI tool your future employer uses: is this predictive (governed like a method) or generative (governed by data-integrity and human verification)? The answer changes what you’re allowed to trust it with.
- You are far more likely to be handed a pre-built PLS model and asked to defend a single day’s predictions than to build a model from scratch — know how to read RMSECV/RMSEP and a T²/Q diagnostic before you know how to fit one.
- “The model still fits the reference method” is not the same claim as “the model is still valid” — the difference is exactly the leakage failure in the worked case above.
For discussion
- M7 accepts two (Q)SAR predictions in place of an Ames test, but not one. Why two, and why does that mirror how you use analytical methods?
- Rank these by the credibility evidence they need: an LLM that summarises papers; an NIR model that releases tablets; a (Q)SAR call on an impurity; an LLM that drafts an OOS narrative. What drives the ranking?
- Why does RMSEC always improve as you add latent variables, while RMSEP eventually gets worse? What is happening to the model?
- You have 40 tablets, 3 NIR spectra each. Describe a cross-validation scheme that will not lie to you, and one that will.
The capstone
An applied problem, worked in teams over the following weeks: a comparability exercise after a manufacturing change; an out-of-specification investigation; a method transfer to a second site; or a specification-setting exercise for a new attribute. Deliverable: a short analytical control strategy and a defence of it — the argument a regulator would have to follow.
Final paper presentations
The paper is a critical analysis of an analytical method, technique, or problem of the student’s choosing — argued the way the course has argued all term: what question does this measurement answer, what decision does it support, how is the result defended, and where does it fail? Strong papers interrogate a real method, incident, or guideline rather than survey a topic; name the failure modes and their detectability (Week 2); connect the technique to a regulatory expectation and a patient consequence; and say what would change the conclusion.
Format
(Instructor: fill in — capstone team size and deliverable format; presentation length and Q&A; whether the written paper is due before or after the talk; grading split across paper, presentation, and capstone; peer-review expectations. This session absorbs what were two full lecture weeks plus presentations — confirm the pacing works in a single 3-hour slot, and consider whether presentations should spill into the make-up slot if numbers require.)
Source note. AI/regulatory material anchored in ICH M7(R2), the FDA draft guidance Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products (2025), FDA discussion papers on AI in manufacturing, the ICH reflection paper on AI, the EU AI Act, and ISPE GAMP guidance — all on the base of GMP data integrity (ALCOA+) and 21 CFR Part 11 / EU Annex 11. Chemometrics follows the standard literature (Brereton, Chemometrics; Martens & Næs, Multivariate Calibration; Esbensen) and ASTM E1655 / ISO 12099. The overfitting/leakage case is a composite of commonly reported failures. (Instructor: confirm current AI-guidance versions and the software the class will use for the notebook thread.)