Run by TEEPTRAK SAS, a production-monitoring software company. No purchase required; plants using any system, or none, are assessed by the same rules.
Factory Excellence Indexby Factory Excellence Awards
Factory Excellence Index

Loss and downtime problem-solving: fixing losses so they stay fixed

This pillar scores what a plant does once a loss is visible, from choosing which losses to work on to proving that a fix actually reduced them. It also covers the two loss areas that most often escape formal problem-solving, equipment care and changeovers with short stops.

What this pillar covers and why it is where losses actually shrink

Visibility tells you what happened and daily routines make sure someone reacts. Neither makes the loss smaller next month. That happens only when a plant works on a loss worth the effort, reaches a cause it can change, and then checks that the change held.

It is also where effort is easiest to waste. A 5 Whys done on a minor loss, or a fix that is never checked, costs the same hours as one that works. Plants with weak problem-solving are rarely idle: they are busy solving the same problems again, often with the same people, and the Pareto looks the same every quarter.

The five questions follow the life of a loss. The first is about choosing, the second about analysis and the third about verification. The last two look at equipment losses shared between maintenance and production, and at changeovers and short stops, which are often left out of problem-solving because each one looks too small to matter.

Named methods (5 Whys, fishbone, A3, 8D, SMED, TPM) appear in the answers as examples. The assessment does not reward the method's name. It rewards triggers, reviews and data checks around whichever method you use.

The five levels as you would see them on the floor

Place your plant in the table, then check it against your pillar index from the assessment. The full level definitions are on the levels page.

LevelWhat you see on the floorWhat the numbers look like in meetings
1 ReactiveWork goes to the loudest problem or the last breakdown. Fixes are applied with no analysis and not checked. Maintenance reacts and the two departments blame each other. Changeovers are not timed.The same breakdown is discussed as if it were new. 'We fixed that last month' is a common sentence, and nobody can show what the fix changed.
2 AwareA monthly review of downtime totals; choices are made on experience. Causes are discussed informally. Planned maintenance exists but is postponed whenever production is under pressure. Changeover duration is roughly known.Monthly totals by department. The causes recorded are symptoms such as 'sensor fault' or 'operator error'. Actions sit in minutes and are checked when someone remembers.
3 StructuredA weekly Pareto of losses by reason with the top items assigned. 5 Whys or a fishbone is used for major events. Planned maintenance is respected and breakdown data is shared. Every changeover is timed; short stops are estimated.A weekly Pareto on the wall and an action list tracked to completion. Actions are closed when the work is done, not when the loss drops, so the same reasons reappear.
4 ProactiveA Pareto per line in hours and money, with named owners for the top losses. A standard method with written triggers, and a manager reviews each analysis. Joint equipment reviews; operators handle basic care. Changeover reduction done on key lines; short stops measured on the bottleneck.Before-and-after charts for every closed action. Recurrence is discussed openly. Maintenance and production present one set of equipment loss figures.
5 ExcellentPlant targets broken into loss budgets per line. Problem-solving skills at every level, with analysis quality coached. Verified fixes update the standard and are copied to similar equipment. Operator care and planned maintenance form one system. Changeovers run to a standard with target times.Losses tracked weekly against each line's budget with trends. The share of problems that come back is tracked. Every changeover over target and every short-stop cause is reviewed across lines.

The five questions and how plants misjudge themselves

Problem-solving is easy to over-rate because the artefacts exist: templates, a Pareto, an action tracker. The questions ask what those artefacts lead to. Each check below uses records you already have.

1. How do you decide which losses to attack? We are asking whether the size of the loss drives the choice. The usual over-rating: a weekly Pareto exists, but the team works on the fifth item because it is easier, or the top bar is 'other' and stays there. A Pareto ranked by number of stops rather than stop time hides the long ones; one ranked only by time hides the short ones.

Self-check: Pareto against workload

List what each team actually worked on last month. Put that list next to last month's Pareto for the same lines. Count how many of the top three losses had someone working on them. If the answer is one or none, the Pareto is decoration.

2. How are root causes analysed? We are asking whether analysis reaches a cause the plant can change, and whether someone checks the quality of the reasoning. Plants over-rate by counting completed forms. A 5 Whys that stops at 'operator error' or 'bearing worn' is a form, not an analysis.

Self-check: read the last five analyses

Take the last five analyses on file. For each, read the final cause. Does it name something the plant can change in a standard, a design, a schedule or a training plan? If it names a person, or restates the symptom in other words, it does not count. Then check whether your triggers for starting an analysis are written down anywhere.

3. When a countermeasure is put in place, how do you check it worked? We are asking whether actions are closed on evidence that the loss changed. An action tracker with every line marked 'done' is Level 3, not Level 4: it proves the work happened, not that it worked.

Self-check: five closed actions

Pick five actions closed last quarter. For each, find the loss figure for the two to four weeks before and after the action. Then look for the same stop reason in this month's Pareto for the same machine. Count how many actions you can prove, and how many came back.

4. How do maintenance and production work together on equipment losses? We are asking whether equipment losses are treated as a shared problem with shared data. Over-rating is common when a joint meeting exists but maintenance presents and production listens, or when operator care means cleaning without any inspection standard.

Self-check: slots and standards

Count the planned maintenance slots scheduled on the bottleneck last quarter and how many were done on the planned date. Then ask an operator on that line to show you the cleaning and inspection standard and to explain what they look for during inspection. A standard nobody can explain is not in use.

5. How are changeovers and short stops managed? We are asking whether both are measured and attacked, not just known. Under-rating is common here: plants assume short stops need automatic capture, when a tally sheet on the bottleneck for a week is a legitimate start. Over-rating is common too: a changeover workshop was run once, and the times have drifted back since.

Self-check: spread and tally

Take last month's changeover times for your most frequent changeover and look at the spread between the fastest and the slowest, not the average. A wide spread means there is no standard in use. For short stops, have the operator on the bottleneck keep a pen tally of every stop too short to log, for one shift. Compare the total with what your records show.

What Level-4 plants do differently

Plants at Level 4 on this pillar do not analyse more problems than others. They analyse fewer, chosen by value, and they refuse to close any of them without data. These are the practices that tend to set them apart.

  • A weekly loss review per line, same day, same time. Thirty minutes at most. The Pareto shows stop time in hours and in money, using one cost per hour agreed once with finance. Each of the top three losses has a named owner and a one-line status.
  • Written triggers for formal analysis. For example, any stop over two hours, or the same reason three times in a week. Below the trigger, the team fixes and records. Above it, a structured analysis is mandatory.
  • A 15-minute manager review of every analysis. A short checklist: is the problem stated with data, was the cause confirmed rather than assumed, does the chain end on something the plant controls, and does the countermeasure address that cause.
  • A closure rule. An action is closed only when the loss data shows a change over at least two to four weeks. For rare failures, they agree a longer window or a leading measure such as an inspection finding.
  • Maintenance and production share one set of equipment figures. Failure frequency and repair time per critical machine are reviewed jointly every week, and operators on the bottleneck carry out basic cleaning, lubrication and inspection against a written standard.
  • Changeovers with a standard and a target. Tasks that can be done while the line runs are separated from those that cannot, each step has a time, and every changeover that runs over target gets a reason written on the board.
  • Short stops counted by cause on the bottleneck. Even by hand. Once the top causes are known, they go into the weekly Pareto like any other loss.
A threshold worth writing down

Decide what counts as a recurrence for your plant, for example the same failure on the same machine within 90 days of a closed action. Write it down and count recurrences every month. It is a direct measure of whether your problem-solving is working, and it costs nothing to keep.

A 90-day plan to move up one level

Your report gives a specific next move for each question. The plan below combines them for a plant that reviews downtime monthly, analyses informally and tracks actions loosely. Keep it to the bottleneck line; a narrow plan that holds is worth more than a plant-wide one that fades.

  1. Weeks 1–2Every Monday, list last week's longest stops on the bottleneck and pick one to work on. Agree a cost per hour of lost output with finance and write it down. Add one column to the action list: 'how will we know it worked?' Hold the first 30-minute weekly meeting between the maintenance and production leads on that week's breakdowns.
  2. Weeks 3–6Build a weekly Pareto of stop time by reason, in hours and money, and assign the top three to named people. Pick one analysis method, train supervisors on it with a real case from the line, and write the triggers for using it. Start recording the start and end time of every changeover on the bottleneck, and run a one-shift pen tally of short stops.
  3. Weeks 7–12Apply the closure rule: no action closed without two to four weeks of before-and-after data. Start manager reviews of each analysis. Film the most frequent changeover, separate the tasks that can be done while the line runs and time each step. Start basic operator care on the bottleneck. At week 12, repeat the self-checks and retake the assessment.

Who owns what during the 90 days:

  • Production manager: the weekly loss review, the Pareto and the closure rule.
  • Maintenance manager: the joint weekly review, planned maintenance compliance and the operator care standard.
  • Finance or controlling: the agreed cost per hour, set once and not reopened every meeting.
  • Supervisors and team leaders: analyses on triggered events and the changeover records.
  • Continuous improvement lead: trains the method, sits in on the first manager reviews, and runs the changeover analysis with the team.
  • Operators on the bottleneck: basic care against the standard and the short-stop tally.

Evidence a jury looks for in verification

For this pillar, jurors in verification look for a chain: a loss chosen from data, an analysis, a countermeasure and a before-and-after check. One complete chain, with its records, is more convincing than a folder of templates.

  • Weekly Paretos for several consecutive weeks, with named owners against the top losses.
  • Completed analyses (5 Whys, A3 or 8D) showing the manager's review and the written triggers that started them.
  • The action list with before-and-after loss data attached to closed actions.
  • Planned maintenance compliance records and completed operator care checklists for the bottleneck.
  • Changeover time records against the standard, and the list of over-target changeovers with reasons.
  • Short-stop counts by cause, whether from a system or a tally sheet.
  • At Level 5, the list of similar machines or lines where a verified fix was copied.

What does not count: blank templates, training certificates on their own, a single showcase A3, an action list marked 'done' with no data, or a changeover workshop report with no changeover records after it. In the interview, expect to walk through one loss from the Pareto to the verified result.

Pitfalls that keep plants stuck on this pillar

  • Analysing everything. When every stop needs a form, forms get filled and nothing is analysed. Triggers exist to protect analysis time for the losses that matter.
  • Root causes that end on a person. 'Operator did not follow the instruction' leads to retraining and nothing else. Ask why the instruction was easy not to follow.
  • Closing on completion. The work is done, the action is closed, and the loss is back three weeks later with a new action number.
  • 'Other' as the top reason. Your largest loss has no owner and no analysis. Split it before you do anything else.
  • Reopening the money debate. If the cost per hour is argued about in every review, the Pareto in money is ignored. Agree it once and date it.
  • Planned maintenance as a buffer. Postponing it for production feels free this week and shows up as breakdowns later. Track postponements so the trade-off is visible.
  • One changeover workshop. Times fall, then drift back because no standard or target time was written and nobody reviews the overruns.
  • Copying a fix without checking fit. A countermeasure spread to similar machines still needs its own before-and-after check on each one.

Where to go next

Take the FEI assessment to see where your plant stands on this pillar and on the five others, with a 90-day move for each of your three biggest gaps. Problem-solving fixes individual losses; the next pillar, the improvement engine, covers how a plant plans, resources and sustains improvement across many of them. Verified plants can be considered for Factory Excellence Recognition.

Questions

Which method should we use: 5 Whys, A3 or 8D?

Any of them, as long as it is one standard method with written triggers and a manager review. 5 Whys is enough to reach Level 3. A common split is 5 Whys for line-level events and A3 or 8D for larger or customer-facing problems; what matters for the score is that the choice is written down and followed.

How do we put a money value on downtime without a long argument with finance?

Agree one cost per hour of lost output on the bottleneck with finance, write down how it was calculated and date it. It does not need to be precise; it needs to be stable, so that losses can be ranked against each other. Review it once a year, not in every meeting.

How long should we wait before closing an action?

The model uses two to four weeks of loss data after the change. For failures that happen only a few times a year, that window is too short: agree a longer one, or use a leading measure such as an inspection result, and write the rule down so the choice is not made case by case.

Our changeovers are set by customer demand. Does the changeover question still apply?

Yes. The question is not about how many changeovers you run but about how each one is done: whether it is timed, whether there is a standard and a target, and whether overruns are reviewed. The more changeovers you run, the more that discipline is worth.

Do we need a formal TPM programme to score well on maintenance and production?

No. The answers describe practices: protected planned maintenance, shared breakdown data, joint reviews, basic operator care and failure-mode analysis. Plants that run TPM will recognise them, but the assessment scores the practices, whatever the programme is called.