
Failure Mode and Effects Analysis (FMEA) helps a team identify how a process, product, or service could fail, understand the effects, and prioritize preventive action before customers feel the impact.
The 2017 British Airways catastrophe shows why that discipline matters. An early estimate put the potential cost as high as £100 million, while IAG later reported an estimated gross cost of about £80 million. The IT outage disrupted tens of thousands of travelers and exposed how a single failure can travel through an interconnected operation.
This guide explains the method, walks through a practical scoring model, and uses the British Airways incident to show where preventive controls could reduce operational risk. You can also run the free FMEA template in Process Street.
In this article:
- Your free FMEA template
- What is Failure Mode and Effects Analysis?
- When do we use FMEA?
- How do you conduct and document FMEA?
- What happened to British Airways?
- Why did the outage happen?
- How can FMEA help prevent a repeat?
Your free FMEA: Failure Modes and Effects Analysis template

If you already understand the method, start with the Failure Modes and Effects Analysis checklist. It gives your team a repeatable way to define scope, identify potential failures, score risk, assign actions, and document the result.
Run the checklist as a standalone risk assessment or connect it to a broader quality management system. A consistent workflow makes ownership and evidence visible, which matters when the same analysis must be reviewed after a change, incident, or audit.
Open the FMEA template in Process Street to copy, customize, and run it with your team.
FMEA? Or, what is Failure Mode and Effects Analysis?

In its most simple form, FMEA is a method for identifying potential problems and prioritizing them so that you can begin to tackle or mitigate them.
It is how you approach your process management from a worst-case-scenario mindset. The perfect job for a pessimist.
Failure Mode and Effects Analysis can also be seen referred to as:
- Potential Failure Modes and Effects Analysis
- Failure Modes
- Effects and Criticality Analysis (FMECA)
Let’s start with some basic terminology to put things into context.
Failure modes: In any process, there are multiple ways that things can go wrong. Each of these ways you can think of are known as modes in the context of FMEA. It could be as simple as Jenny from accounting is off ill and this creates a problem for the accounts receivable department. Or it could be some hugely complicated problem in a massive manufacturing plant operating with automated technology. The core concept is still the same: the way something can go wrong in the process is seen as a mode of failure; a Failure Mode.
Effects analysis: This one is even simpler. What are the effects of an element of the process failing? Is it that Jenny’s work gets delayed a day? Or does Darren have twice the workload that day to keep things on track? But what is the effect of either of those outcomes on the broader company performance? What if the extra workload on Darren causes him to make some errors in the rush? Darren’s a nice guy, but he’s only human. Analyzing the effects of the initial problem means following the path of causality and investigating each potential problem from there.
If you want to know more about criticality analysis specifically, check out this article: Criticality analysis: What is it and how is it done?
So, within the framework of FMEA, the point of the exercise is to be able to identify all the different failure modes and then evaluate the potential damaging effects of each. It is vital to assess a number of different ways the effects can be damaging.
It could be how severe the consequences are, like with the British Airways error. It could be how frequently they occur; how many defects are present in every manufacturing batch, for instance. Or, it could be how elusive the failure is: difficult to stop, to see, or even to identify at the very beginning.
These potential effects need to be prioritized in order for us to tackle them systematically. Within the FMEA framework, we prioritize them in the order mentioned above: severity, frequency, ease of detection.
With priority defined, a company can begin to work through their long list of potential process failures either one by one or in related groups. By tackling the most dangerous problems first and the smaller harder ones last, a company is minimizing its exposure to risk as effectively as possible.
When do we use FMEA?

Short answer: all the time.
Maybe I’m being a bit hyperbolic. You don’t actually have to use FMEA all of the time, but at Process Street we operate with a strong focus on continuous iteration, improvement, and process optimization. Our guiding principles of business process management leave us with the intention of forever assessing and reviewing our internal processes, and FMEA is one method of process improvement; albeit, a highly risk-focused one.
The key use cases are really at moments of change. Continual iteration and assessment is encouraged, but not all businesses have the resources to have a dedicated process team and the time amongst their other workers to contribute to the investigation. Continually performing FMEA could result in greater inefficiencies than benefits for a company not equipped to do so.
The primary use case would be when you’re designing a new process. A process could take the form of a new way for a team to operate: new policies and procedures. Or, a new process could apply to providing a whole new service, where you’re entering into uncharted territory without previous similar processes to build from. It could even apply to a new product, in the manufacturing, distribution, or even customer use of that product as part of a wider risk assessment.
The framework is broad and general so that it can apply to a multitude of different scenarios.
When building a new process of any form, the process should be well documented and mapped. Once this design has been done, you can see the FMEA as a risk assessment of your concept. This technique should allow you to live out the process and explore what it will be like in action. As you discover the different failure modes, you will be able to adapt and improve the process before deployment.
FMEA is also used often in process reapplication or redesign. Perhaps you’re taking an existing process and looking to improve it, or to use it again in an adapted way on a slightly different use case. Whenever a process is to undergo significant change, FMEA is used to detect those potential new problems which could have arisen from the changes.
The real moment where you have to use FMEA is when you’ve failed to use it effectively in the past: when disaster strikes. This is the position British Airways is in now. After the problem has occurred, you need to launch a full investigation into what happened, why it happened, and how it could be stopped from happening again in future: don’t let the normalization of deviance set in!
Of course, you can’t stop every problem ever from happening. We’ve all seen Minority Report. But with effective use of FMEA, you can re-evaluate your processes to hopefully minimize the damage of these errors or incidents.
Conducting and documenting Failure Mode and Effect Analysis: A user’s guide

FMEA works best as a facilitated team exercise, not a spreadsheet completed in isolation. Bring together people who understand the process, its technology, its customers, and its controls. Their different perspectives make the list of failure modes more complete and the scoring more credible.
- Define the scope. Name the process, product, or service being assessed, its boundaries, the expected result, and what is out of scope.
- Break the scope into functions or steps. For an ATM, that might include reading a card, authenticating the account, authorizing the withdrawal, dispensing cash, and recording the transaction.
- Identify failure modes, effects, and causes. Ask how each function could fail, what the customer or operation would experience, and what could cause that failure.
- Record current controls and score the risk. Score severity (S), occurrence (O), and detection (D) on the chosen scale. Document the rationale so another reviewer can understand the rating.
- Prioritize and assign actions. Choose controls that prevent the cause, reduce the effect, or improve detection. Give every action an owner and due date, then rescore after implementation.
In the classic model, the Risk Priority Number is calculated as S × O × D. A criticality score may also be calculated as severity × occurrence. These results help order the work, but teams should not rely on RPN alone: different combinations can produce the same number, and a catastrophic effect can remain unacceptable even with a low occurrence score.
The scoring example in Nancy R. Tague’s The Quality Toolbox uses an ATM to show how functions, effects, causes, controls, and ratings fit into a grid. That structure is still practical because it links each recommended action to a specific failure path.
Three disciplines keep the analysis useful: involve all key stakeholders, understand the scope, and document the process. Documentation creates the audit trail needed to review assumptions, confirm that actions were completed, and revisit the analysis when the process changes.
What happened to British Airways?

On May 27, 2017, British Airways experienced a power loss at a UK data center. The disruption affected check-in, baggage, flight operations, and customer communications across a holiday weekend. More than 700 flights were canceled over three days, with approximately 75,000 travelers affected.
Contemporary reporting initially discussed a possible £100 million impact. IAG later said the outage had an estimated gross cost of about £80 million. The difference matters: £100 million was an early forecast, not the final reported cost. The headline remains a useful reminder of the scale of exposure, but the later figure is the stronger historical benchmark.
The measurable effects included cancellations, compensation, hotel and rebooking costs, extra customer-service work, and operational recovery. The harder-to-measure effects included customer frustration and damage to trust. As then-IAG chief executive Willie Walsh acknowledged, the incident was damaging to the British Airways brand even though he believed the brand remained resilient.
An FMEA team would describe those as effects attached to different functions: unavailable systems prevent check-in and dispatch, disrupted dispatch cancels flights, canceled flights strand travelers, and delayed communications increase both cost and frustration. Following that chain keeps the analysis tied to operational outcomes.
Why did it happen?

British Airways said there was a loss of power to the UK data center, followed by an uncontrolled return of power that damaged IT systems. Reporting at the time focused on the uninterruptible power supply (UPS), the equipment intended to maintain a stable supply and support recovery when mains power fails.
The precise sequence and individual responsibility were disputed in public reporting, so the safest conclusion is narrower than some early accounts suggested: the outage involved a power interruption and a restoration process that caused physical damage and delayed recovery. It is not necessary to assign blame to recognize the control problem.
From a process perspective, the critical questions are clear. Who was authorized to interrupt and restore power? What written procedure governed the work? Which technical interlocks or approvals prevented an unsafe restart? Was trained supervision required? Could the system detect the wrong state before power reached vulnerable hardware? Were failover and recovery plans tested under realistic conditions?
Those questions turn a dramatic incident into analyzable failure modes. They also avoid treating “human error” as a root cause. People act inside systems. A robust process anticipates slips, unclear handoffs, time pressure, and incomplete information, then designs controls that make a dangerous action harder to perform and easier to detect.
How can FMEA help us prevent this happening again?

FMEA could not guarantee that an outage would never occur, but it could make the failure paths visible before a real incident tested them. For a data-center power procedure, the team would map each step from shutdown authorization through power isolation, maintenance, restoration, validation, and operational handback.
Three analyses are especially relevant. First, review the task-level power restoration procedure and its technical controls. Second, review contractor onboarding, competence, supervision, and access using a documented contractor onboarding process. Third, review business continuity, redundancy, failover, and recovery so a localized failure cannot disable a wider operation.
For example, a team might score uncontrolled power restoration with severity 9, occurrence 2, and detection 8. The classic RPN would be 144. The exact number depends on agreed scales and evidence, but the high severity and weak detection should trigger action: interlocks, independent approval, a controlled restart sequence, monitoring, and a tested rollback plan.
The most effective controls prevent the cause or limit the effect instead of relying only on reminders. That could mean physical or software interlocks, role-based access, two-person verification, validated runbooks, live telemetry, staged energization, failover capacity, and rehearsed incident response. Every control should have an owner, evidence, and a review trigger.
Risk management tools help teams connect that analysis to ongoing work. In Process Street, one platform combines Docs for controlled process knowledge, Ops for running accountable workflows, and built-in AI to help teams work with process information. Teams can route approvals, assign corrective actions, capture evidence, and keep a record of what changed without separating the procedure from its execution.
The lesson is broader than aviation. At moments of change, test how the process can fail, prioritize the most consequential paths, and verify that the controls work in practice. FMEA turns pessimism into a disciplined form of prevention.
For more quality management resources: