FMEA in Detail — Scoring, Scaling, and Where It Breaks

Failure Mode and Effects Analysis worked end to end on an HPLC assay method: the RPN formula, a full failure-mode table with a before/after action, why detection runs backward, and the known weaknesses that make RPN easy to over-trust.
A banner titled 'FMEA in Detail — Scoring, Scaling, and Where It Breaks' with the tagline 'Find the important failures. Take the right action. Don't over-trust the number.' and the note that this is a worked example from an HPLC assay method, with practical guidance for real decisions. Panels: (1) The One Idea — RPN is a prioritization tool, not a measurement; it tells you which failure mode to look at first, not how much worse one is than another, and treating it like it does is the single most common way an FMEA goes wrong, with the note to focus effort where it matters for the patient and avoid gold-plating the rest; (2) How FMEA Works — a five-step chevron: list the process steps (e.g., sample prep, separation, detection, data analysis, reporting), identify failure modes (what could fail at each step), analyze effects and causes (what would it do, why would it happen, how would we catch it), score S, O, D (Severity, Occurrence, Detection), take action and re-score (implement risk controls and re-score to show the improvement) — captioned as a living tool, updated when methods, instruments, materials, or knowledge change; (3) The RPN Formula — RPN = S × O × D, with Severity (how bad the effect is on the patient or the decision, 1 = negligible to 10 = catastrophic e.g. potential patient harm), Occurrence (how often it is expected to happen, 1 = remote to 10 = very high frequency), and Detection (how likely current controls are to catch it before it matters, with the callout that a high score means poorly detected because the scale runs backward — 1 = almost certain to detect, 10 = very unlikely to detect); (4) Typical 1–10 Scales (Examples) — a table mapping score bands (10, 5, 1) to example Severity, Occurrence, and Detection descriptions, with a note to use defined, documented criteria tailored to the method, product, and patient risk; (5) Worked Example — HPLC Assay Method FMEA, a five-row table (mis-integrated peak, wrong diluent used, column-to-column carryover, drifting calibration curve, co-eluting unknown degradant) each with process step/failure mode, effect, cause, current control, S, O, D, RPN, a risk control action, and a re-scored RPN, with the note that carryover (RPN 200) outranks drifting calibration (RPN 108) even though a wrong release decision from drifting calibration may seem worse — because detection was poor (D = 8) — RPN doing its job of surfacing the blind spot, not just the scariest-sounding failure; (6) Why the Same RPN Can Mean Very Different Things — two failure modes (A: rare but severe, S=9 O=2 D=5; B: more frequent, less severe, S=5 O=3 D=6) both scoring RPN 90, with the point that a severity-first reviewer would act on A first regardless of the tied RPN, which is exactly why you shouldn't rank by RPN alone and why any mode with severity ≥ 9 should be automatically flagged for action; (7) Known Limitations — RPN is an ordinal product, not a true measurement (100 is not twice as bad as 50); detection and occurrence are often guesses, and Q9(R1) highlights this subjectivity, asking for defined scales, cross-functional input, and documented assumptions; different (S,O,D) combinations can give the same RPN with very different meaning; it can miss low-probability, high-severity events (use severity-first rules); it is not a substitute for scientific judgment — it's a tool to structure it; (8) FMEA vs. FMECA — a side-by-side comparison: FMEA uses S×O×D (RPN), prioritizes failure modes, is simple and widely used, good for method development; FMECA adds criticality analysis (e.g., a severity/probability matrix), often includes failure-mode ratios, is better for high-risk or regulated products, and most analytical FMEAs are effectively FMECAs in practice; (9) When to Use Other Risk Tools — a table of five tools (fault tree analysis for working backward from a failure, useful for OOS root-cause investigation; HACCP for identifying and controlling critical points, useful in manufacturing or sample handling; HAZOP for deviations from design intent, useful in process/engineering systems; risk ranking and filtering for comparing many unrelated risks, useful for site or portfolio decisions; Ishikawa/fishbone/PHA for first-pass hazard identification, useful for early method or process review); (10) From Risk to Control Strategy — a five-step numbered flow: identify CQAs (attribute risk assessment), assess method parameters (method FMEA), implement controls (e.g., robustness, system suitability), set specifications (Q6) and stability program (Q1), monitor and manage change (Q14), with the note that a control strategy is the output of risk management, not a separate exercise; (11) Worked Case — Nitrosamine Risk Assessment, a bulleted walkthrough: identify hazard (potent mutagenic carcinogens, e.g. NDMA, NDEA, drug-specific), analyze risk (synthetic route, nitrite sources, secondary amines, recovered solvents), control risk (route changes, nitrite scavengers, tighter limits at ppb levels), communicate (to the agency, on a deadline), analytical challenge (need methods sensitive enough to detect at the acceptable intake). Footer: 'Science + Risk-Based Thinking = Better Medicines for Patients,' Temple University branding, and the tagline 'All science ultimately serves people.'

The one idea

RPN is a prioritization tool, not a measurement. It tells you which failure mode to look at first — it does not tell you how much worse one failure mode is than another, and treating it like it does is the single most common way an FMEA goes wrong.

Mechanics

Failure Mode and Effects Analysis decomposes a method or process into steps, and for each step asks: what could fail (failure mode), what would that do (effect), why would it happen (cause), and how would we catch it (controls)? Each mode is scored on three independent 1–10 scales and multiplied:

Risk Priority Number = Severity × Occurrence × Detection

  • Severity — how bad the effect is for the patient or the decision. A wrong release decision (a failing batch shipped, or a good batch scrapped) scores high; a re-run that costs a day scores low.
  • Occurrence — how often the cause is expected to produce the failure, from historical data or, absent that, engineering judgment.
  • Detection — how likely the existing controls are to catch the failure before it matters. High detection score = poorly detected — this scale runs backward from the other two, and it is where most FMEAs go wrong: a “10” means “we would almost certainly miss this,” not “we’d definitely catch it.”

Modes with a high RPN, or a high severity regardless of RPN, get a corrective action; the mode is then re-scored to show the action actually moved the number, not just noted “action taken.”

Worked example — an HPLC assay method

Failure modeEffectCauseCurrent controlSODRPNActionRe-scored RPN
Mis-integrated peakWrong reported assay valueManual integration override without documented rationalePeer review of chromatograms846192Require documented integration parameters; lock auto-integration settings8 × 4 × 2 = 64
Wrong diluent usedLow or erratic recoverySimilar-looking bottles stored adjacent on the benchAnalyst training63590Segregate diluent storage; barcode-scan verification at weigh-in6 × 3 × 2 = 36
Column-to-column carryoverGhost peak misread as an impurityInsufficient wash gradient between injectionsNone — relies on visual inspection558200Add a blank injection after each sample series; extend wash time5 × 5 × 3 = 75
Drifting calibration curveSystematic bias in reported resultStandard degraded between preparation and useSystem suitability at run start only926108Add a mid-run suitability check; shorten standard hold time9 × 2 × 3 = 54
Co-eluting unknown degradantImpurity result reported lowInsufficient resolution between API and degradantResolution check in system suitability934108Switch to an orthogonal column for confirmatory testing9 × 3 × 2 = 54

Two things worth noticing in this table: the carryover mode (RPN 200) outranks the drifting-calibration mode (RPN 108) even though a wrong release decision from a drifting curve is arguably worse — because carryover’s detection score was so bad (8: nobody was actually looking for it). That is RPN doing its job: surfacing the blind spot, not just the scariest-sounding failure.

Why the same RPN can mean very different things

Failure modeSODRPN
A92590
B53690

Both score 90. Mode A is a rare but severe failure that’s moderately well detected; mode B is a more frequent, less severe failure that’s poorly detected. A severity-first reviewer would act on A first regardless of the tied RPN — which is exactly the argument for not ranking a whole FMEA by RPN alone, and for flagging any mode with severity ≥ 9 for action independent of its RPN.

FMEA vs. FMECA

FMECA adds a formal criticality analysis on top of FMEA — instead of (or alongside) the RPN product, each failure mode’s criticality is assessed against a defined severity/probability matrix, often with failure-mode ratios when one cause can produce several distinct failure modes. In practice, most analytical-development FMEAs are really FMECAs in miniature: teams already flag “any severity ≥ 9 regardless of RPN” as an action trigger, which is a criticality rule, not a pure RPN rule.

Known weaknesses — worth teaching so students don’t over-trust the number

  • RPN is an ordinal product treated as if it were interval data; an RPN of 100 is not “twice as bad” as 50, and — as shown above — different (S, O, D) triples give the same RPN with very different meaning.
  • Detection and occurrence are often guessed. Q9(R1) explicitly flags this subjectivity and asks for it to be managed (defined scales, cross-functional scoring, documented assumptions).
  • Many programs now supplement or replace RPN with a severity-first criticality matrix, or with risk ranking and filtering when comparing failure modes across unrelated processes.

When to reach for something else

FMEA decomposes one process step by step and scores every mode on the same three scales — it’s the right tool when the process is defined and you’re building or revising its control strategy. Reach for fault tree analysis instead when you’re working backward from a failure that has already happened and need to trace its root cause; reach for risk ranking and filtering when you’re comparing risks that don’t share a process or a scale at all.