Andon: The Lean Signal System Explained

Andon signal light and workstation show how an abnormality becomes visible

Turn this article into takeaways for your work.

Each assistant summarizes the article only for you and suggests best practices for your work.

A machine makes a noise it shouldn't. One person hears it. What happens in the next thirty seconds decides whether that noise costs you a minute or a shift, and andon is the system that decides it.

Andon is the signalling layer of a lean production system: the boards, lights, cords, buttons and dashboards that turn one person noticing something wrong into a named responder arriving while the problem is still small. It doesn't diagnose anything and it doesn't fix anything. Its only job is to make an abnormality impossible to miss and impossible to ignore, quickly enough that the problem can still be solved where it started.

Most writing about andon argues that stopping work is a good idea. This page assumes you agree and looks at the machinery instead.

Key Facts: Andon

  • The Lean Enterprise Institute defines andon as "a visual management tool that highlights the status of operations in an area at a single glance and that signals whenever an abnormality occurs" (LEI Lean Lexicon).
  • Toyota states that when equipment stops, "the andon (problem display board) lights up to notify workers of the abnormality", and that it also lights "when the stop cord is pulled so that workers can call the person in charge" (Toyota Production System).
  • A pull is not automatically a shutdown. Under a fixed-position stop system the line halts "at the end of the work cycle", and only if the problem cannot be solved during that cycle (LEI Lean Lexicon).
  • The lineage runs back to Sakichi Toyoda's automatic loom, which Toyota says built "the capability to make judgments into the machine itself" (Toyota).

What Andon Actually Is

Andon is the Japanese word for a paper lantern, and the industrial meaning kept the sense of the original: a light that tells you something. The Lean Enterprise Institute's lexicon describes the typical form as "an overhead signboard with rows of numbers corresponding to workstations or machines", where a number lights when a sensor detects a problem or "an operator pulls a cord or pushes a button" (LEI).

A worker and sensor activate an Andon signal for a responder

Two things follow, and both get lost when people picture andon as purely an alarm.

The first is that andon is a status display before it's an alert. A board showing every station running normally tells the area leader that no intervention is needed anywhere, which is as useful as a red light telling them where to walk. An andon that only exists when something is wrong is a klaxon.

The second is that it has two trigger paths, machine and human, and they aren't interchangeable. A sensor trips on conditions you already knew to look for. A person pulls on conditions nobody wrote a rule for: a smell, a part that fits but feels wrong, a customer sentence off the script. Build only the sensor path and you detect exactly the failure modes your designers anticipated.

Andon is also the best-known instance of visual management, which LEI defines as placing tools, parts and performance indicators in plain view "so the status of the system can be understood at a glance by everyone involved" (LEI).

Where this page stops and jidoka begins

Jidoka is the parent concept and it owns the argument. LEI states it as "providing machines and operators the ability to detect when an abnormal condition has occurred and immediately stop work" (LEI). If you're weighing up whether building quality in beats inspecting it at the end, read that page.

Andon is the wire between the detection and the stop, carrying "something is wrong here" from whoever knows it to whoever can act on it, with a clock attached. Jidoka says stopping is right. Andon is how anyone finds out there's something to stop for: the signal, the ladder it climbs, the clock it runs, and the record it leaves.

Where the Word and the System Came From

Sakichi Toyoda's automatic loom, built in the early twentieth century, stopped itself when a thread broke. Toyota describes its significance as building "the capability to make judgments into the machine itself" rather than simply automating hand work (Toyota). A loom that halts on its own is self-contained: detection and response sit in the same box. An assembly line isn't a loom. Detection and response are held by different people in different places, and that gap is what andon was invented to close. Toyota's account describes both halves of the modern arrangement: equipment that stops itself and lights the board, and an andon "set to light up when the stop cord is pulled so that workers can call the person in charge".

Jidoka is one of two pillars of the Toyota Production System. The other is just-in-time, which Toyota summarises as "making only what is needed, when it is needed, and in the amount needed". Just-in-time strips out the inventory buffers that used to hide a quality problem for a shift, so an unsignalled defect now arrives at the next station exactly when that station needs a good part. That's why a serious lean implementation needs a working andon, not just a policy about stopping.

The Anatomy of an Andon System

Every andon system, factory floor or service desk, is the same six stages. Most failures trace to exactly one of them, which makes this a useful diagnostic.

Six stages of Andon from trigger and signal through escalation, response, resolution and record

Stage What it does It fails when
Trigger A sensor trips or a person pulls, pushes, clicks or calls The threshold is so tight normal variation trips it, or so loose real abnormality passes
Signal The status becomes visible to everyone who needs it, instantly and unasked It reaches a screen nobody watches, or an inbox rather than a place
Escalation If nobody answers inside the window, the signal climbs a tier automatically Escalation depends on the person who raised it chasing someone
Response A named role arrives, assesses, and decides whether work continues or stops "The team" is named instead of a role, so everyone assumes someone else is going
Resolution The immediate condition is cleared and normal work resumes The fix is a workaround recorded as a fix
Record The event, its cause and its countermeasure are logged The signal clears and leaves no trace, so the same problem is new every time

Two stages deserve a note. The trigger stage is where poka-yoke belongs: mistake-proofing is the cheapest trigger because it detects without relying on anyone paying attention. The response stage is often already written down, because the reaction plan column of a control plan specifies what happens when a characteristic goes out of limits. If you have one, most of your response protocol is drafted and needs only a signal attached.

The record stage gets cut for time and shouldn't be. Without it, pull number 400 for the same cause looks identical to pull number one, and the andon becomes an efficient way of solving the same problem forever.

Andon Board, Cord, Light or Digital

"Andon" covers four fairly different pieces of equipment. Mixing them is normal, but each is better at something the others are worse at.

Form What it is Strongest at Weakest at
Andon board Overhead signboard, one light or number per station Area-wide status at a glance, showing normal as clearly as abnormal Detail. It tells you where, not what
Andon cord or button A pull cord or button at the workstation Speed and low friction. Nothing to log in to, nothing to classify Carrying information beyond "here, now"
Andon light Coloured tower on a machine, red for problems and green for normal operations (LEI) Machine-level status readable from anywhere on the floor Human-raised issues, which have no machine to sit on
Digital andon Dashboard, mobile alert or ticket queue with automatic escalation Routing, timestamping and the record stage Being seen. A screen is easy to look away from, a light overhead is not

The trade-off is between visibility and detail. A physical light is unmissable and nearly contentless. A digital alert is rich, routable and self-logging, and it competes for attention with every other notification the responder gets that hour. Teams that go digital and then wonder why response times got worse have usually swapped a place for a feed. The practical answer is both: physical channel for attention, digital channel for escalation timing and the record.

What Each Signal State Obliges Someone to Do

A colour that means "attention" and nothing more will be ignored within a month. Every state needs a specific owner and a specific required action, posted where the signal is. The scheme below is a common convention, not a standard: LEI confirms only the base pair, red for problems and green for normal operations, so define the rest locally and write it on the board.

State Common meaning Who owns the response Required action
Green Running to plan, no help needed Nobody None. Green is information, not silence
Yellow or amber Help requested, work continues for now Team leader Attend within the window and decide: fix inside the cycle, or escalate
Red Work has stopped, or will stop at the fixed position Supervisor or area manager Attend in person, authorise the restart, own the record
Blue or white Material or tooling shortage rather than a defect Materials or logistics role Replenish, and log the shortage as a separate cause class
Repeated red Same cause signalled again inside a defined window Engineering or the improvement owner Treat the repeat as the real event, open a root-cause investigation

The last row is the one most systems lack and most need. A repeat pull is a different event from a first pull, because it proves the previous resolution was a workaround.

The Escalation Ladder and Response-Time Design

An andon without a clock is a suggestion. The escalation ladder converts a signal into an obligation, and it works by making the next tier automatic rather than requested.

A clock beside ascending steps illustrates timed Andon escalation

Tier Trigger Responder Target response time
0 Operator spots it and can correct it inside the cycle The operator Immediate, inside the work cycle
1 Operator signals for help, work continues Team leader for that zone Within the remaining work cycle. On a Toyota line at 60 second takt, that's the 60 second interval
2 Tier 1 cannot resolve inside the cycle, or work stops Area supervisor or shift manager Under 5 minutes from the original signal
3 Stop exceeds a defined duration, or the cause repeats Engineering, maintenance or quality Under 15 minutes, countermeasure owner assigned
4 Stop threatens customer commitment, safety or delivery Plant or operations leadership Under 30 minutes, decision authority attached

Three design rules separate a working ladder from one that exists on a laminated card.

Response time runs from the signal, not from the handover. Tier 2's five minutes starts at the original pull, not when tier 1 gave up. Otherwise every tier restarts the clock and a four-tier ladder legitimises a fifty-minute response.

Every tier names a role reachable in that time. A fifteen-minute tier assigned to an engineer covering three buildings is fifteen minutes on paper only. Set the tier to what the named person can achieve, or change the person.

The window comes from the work, not from a round number. Takt time gives you tier 1 for free on a paced line: the responder gets the remainder of the cycle, because that's the time available before the product moves. Off a paced line, derive it from how long the abnormality can persist before it gets expensive, and record that reasoning in your standard work.

A pull is not a shutdown

The commonest misconception about andon is that pulling the cord stops the line. Under a fixed-position stop system it doesn't, at least not straight away. LEI describes the method as stopping the line "at the end of the work cycle", and only "if a problem is detected that cannot be solved during the work cycle". The operator signals the supervisor, the supervisor assesses whether the issue can be corrected before the cycle ends, and if it can, resets the signal and the line never stops (LEI).

That is what makes a high pull rate affordable. If every pull were an immediate shutdown, the cost of raising a hand would be enormous and the system would suppress itself within weeks. The fixed-position stop makes the signal cheap and the stop rare, and the gap between those two numbers is where the value sits. The automated equivalent is automatic line stop, which LEI defines as ensuring "that a production process stops whenever a problem or defect occurs".

The Cultural Precondition Nobody Budgets For

Hardware is the cheap part. The expensive part is a workforce that believes pulling the cord is a contribution rather than a confession, and no amount of equipment produces that belief.

The evidence is worth taking seriously because it surprised the researcher who found it. Amy Edmondson went into two Boston hospitals expecting better-performing teams to report fewer medication errors. She found the opposite: teams scoring higher on teamwork measures showed higher detected error rates. Her conclusion, in her own words, was that "better teams probably don't make more mistakes, but they are more able to discuss mistakes" (Behavioral Scientist). That launched the research programme that became psychological safety.

Read it against an andon board and the implication is uncomfortable. Reported abnormality rate measures willingness to report at least as much as actual abnormality. A quiet board in a blame culture and a quiet board in a stable process look identical from the supervisor's office.

The difference is made in small, repeated moments. Whether the responder arrives and asks what happened or asks who did it. Whether a pull that turns out to be nothing is thanked or sighed at. Whether the andon log is ever quoted in a performance review, which is the fastest way to kill a programme permanently. Whether leadership pulls the cord themselves during a gemba walk. This is where andon stops being a tool and becomes part of the quality culture that total quality management covers.

Metrics That Tell You Whether the Andon Works

Andon generates its own telemetry, which makes it easy to measure honestly. The catch is that the most obvious metric reads backwards.

Paired gauges and a magnifying glass illustrate reading Andon metrics together

Metric Definition Healthy direction The trap
Pull rate Signals raised per shift or per thousand units Rising early, then stable A fall reads as improvement and usually isn't
Response time Signal to responder physically present Falling, tightly distributed An improving average hides a worsening tail. Track the 90th percentile
Resolution time Responder present to normal work resumed Falling for repeat causes Pressure here produces workarounds recorded as fixes
Stop rate Share of pulls that became an actual stoppage Falling while pull rate holds Falling because people stopped pulling is the same number for the opposite reason
Repeat rate Share of pulls whose cause was signalled in the last 30 days Falling A high repeat rate is a countermeasure problem, not an andon problem
Countermeasure closure Share of logged causes with a permanent fix Rising to a stable ceiling Closing items by relabelling them as accepted risk

Be blunt about pull rate: a falling pull rate is usually bad news. There are two explanations for fewer signals. Either the process genuinely became more stable, or people stopped telling you when it isn't. The second is far more common, and it's the failure mode Edmondson's hospital data describes.

Tell them apart by checking pull rate against an independent measure of quality. If pulls fall while escaped defects, complaints, scrap and rework fall too, the process improved. If pulls fall while any of those hold steady or rise, the andon is being suppressed. That cross-check belongs in whatever routine you use to monitor the process, because it's invisible in the andon data alone.

Repeat rate is what connects andon to improvement. Each repeated cause is an invitation to run five whys or a fuller root cause analysis and close the loop through your kaizen cycle. A healthy pull rate alongside a stubborn repeat rate means the andon is doing its job and the problem sits downstream of the signal.

Andon Outside the Factory

The pattern (unmissable signal, named responder, bounded window, automatic escalation, logged record) transfers to any operation where someone can spot a problem earlier than the system notices it. The equipment changes completely. The six stages don't.

A laptop alert and headset illustrate Andon signals outside manufacturing

Setting Trigger Responder Equivalent of "stopping the line"
Software incident response Failed health check, error-budget burn, or an engineer declaring an incident On-call engineer, then incident commander Freeze deploys, roll back, block the release pipeline
Customer support desk Third contact about the same defect in one shift Support lead, then product owner Pull the article, pause the campaign, hold the order flow
Field service Technician finds an unsafe or out-of-spec condition on site Dispatcher, then service engineering Suspend that job type until assessed
Clinical care Barcode mismatch at dispensing, or a nurse's concern about an order Pharmacist or rapid response team Hold the dose until reconciled
Back-office processing Exception rate on a queue crosses a threshold Team leader, then process owner Stop releasing new work into the queue

Software incident response is the closest structural match, and Google's SRE practice makes the parallel explicit. Its guidance insists that "everybody involved in the incident knows their role and doesn't stray onto someone else's turf", which is exactly the tier discipline of an escalation ladder. It also argues for declaring early, because "it is better to declare an incident early and then find a simple fix and close out the incident than to have to spin up the incident management framework hours into a burgeoning problem" (Google SRE Book). That's the fixed-position stop argument in different vocabulary. The service version has a longer history than most people assume: Amazon has publicly listed "our customer service Andon Cord" among its own internally driven innovations, alongside Kindle FreeTime and AutoRip (Werner Vogels, Amazon CTO).

One adaptation is mandatory. A factory has a shared physical space, so the signal can simply be visible. A distributed team has none, so the response window has to be enforced by tooling rather than by line of sight. Write those windows into a customer-facing commitment and the structure is the one used in service level agreements: a defined trigger, a named owner, a clock, and a consequence when the clock runs out.

How to Build an Andon System That People Actually Pull

Start from the abnormalities you already have. Spend a week logging what goes wrong, where, and how long it took someone to notice. That list is your trigger specification. Buying hardware first produces a system that signals whatever the vendor anticipated.

Define response before you define signals. For each abnormality class, write down the responding role, the window, and what they're authorised to decide on arrival. A signal with no defined responder trains everyone to ignore signals.

Pick the cheapest trigger that works, and make the signal a place rather than a feed. Cord or button where seconds matter and hands are busy, sensor where the condition is measurable, digital flag where the work is distributed. Put it where the responder already looks. Resist classifying at the moment of pull: every field you add is friction, and friction is the enemy of pull rate.

Pilot on one cell, one shift, one queue for two to four weeks. Check whether pull rate rises (it should, sharply, because you're surfacing abnormalities that were previously absorbed) and whether response time holds as volume climbs. If response degrades under pilot load the ladder is wrong, and rolling out makes it worse everywhere at once.

Instrument the record from day one: cause, timestamp, tier reached, resolution, countermeasure owner. Retrofitting a log after six months means six months of pulls that taught you nothing.

Review the log weekly with the people who pull. Sort by repeat rate, pick the top cause, assign a countermeasure owner, close it before the next review. This is where andon becomes improvement rather than firefighting, and it's the clearest proof to the floor that pulls matter.

Limitations and Honest Failure Modes

Andon can surface a problem quickly, but response capacity, root-cause discipline and trust determine whether the signal leads to a lasting fix.

An alarm above a blocked paper queue illustrates Andon response capacity limits

Limitation Why it matters
Andon signals, it doesn't solve Weak root-cause discipline gets you faster firefighting, not fewer fires
Alarm fatigue is a real ceiling Thresholds tuned too tight produce constant signalling, and a floor that has learned to ignore a light cannot be un-taught cheaply
Suppression is invisible in the data A programme dying of blame culture and a genuinely stable process produce the same quiet board
The response tier is the real constraint Signalling capacity is cheap and response capacity isn't. An andon whose tier 1 role is already fully loaded creates overburden rather than flow
Digital versions decay quietly A light that burns out is noticed in minutes. A dashboard nobody has opened in three weeks looks like a healthy one

Frequently Asked Questions about Andon

What does andon mean?

Andon is the Japanese word for a paper lantern. In lean production it means a visual signal system that shows an area's status and alerts a responder when something goes wrong. The Lean Enterprise Institute defines it as a visual management tool showing operations at a single glance, which signals whenever an abnormality occurs.

What is the difference between andon and jidoka?

Jidoka is the principle that work should stop the moment an abnormality appears, so defects never travel downstream. Andon is the mechanism that makes jidoka possible, carrying the abnormality from whoever detected it to whoever can respond, within a defined window. Jidoka is the philosophy, andon is the wire.

Does pulling the andon cord stop the production line?

Not immediately, under a fixed-position stop system. The pull signals the team leader, who assesses whether the problem can be fixed within the remaining work cycle. If it can, the leader resets the signal and the line never stops. If it cannot, the line halts at the end of the cycle, at a fixed position.

What do the andon colours mean?

The base convention is green for normal operations and red for a problem. Many plants add yellow for help requested without a stop, and blue or white for a material shortage. Beyond red and green there's no universal standard, so define each state's meaning and required action locally and post them at the signal.

Is a falling andon pull rate a good sign?

Usually not. Fewer signals can mean the process stabilised, or that people stopped reporting, and the second is more common. Check pull rate against independent quality measures such as escaped defects, scrap, rework and complaints. If those fall too, the process improved. If they don't, the andon is being suppressed.

Can andon work outside manufacturing?

Yes, and the six stages transfer intact: trigger, signal, escalation, response, resolution, record. Software teams run it as incident alerting with an on-call ladder, support desks as a repeat-issue flag routed to a named lead. The one adaptation: distributed teams have no shared sightline, so escalation timing must be enforced by tooling rather than by a visible light.

What's the difference between andon and poka-yoke?

Poka-yoke is mistake-proofing, a device or design that makes an error impossible or self-detecting, and it sits at the trigger stage. Andon is what happens next: carrying that detection to a responder with a clock attached. Poka-yoke reduces how often you need the signal, andon decides what the signal is worth.

An andon system is easy to buy and hard to earn. The lights, cords and dashboards are a few weeks of work. What takes longer is a floor that believes a pull is welcome, a ladder whose response times are real rather than aspirational, and a weekly habit of reading the log with the people who raised the signals. Get those right and the hardware barely matters. Get them wrong and you've bought an expensive set of lights everyone has learned to walk past.

About the author

Tara Minh

Tara Minh

Senior Operations & Growth Strategist

Tara Minh is Senior Operations & Growth Strategist at Rework, helping B2B SaaS leaders scale without breaking their teams. With 8+ years in revenue operations and process optimization, Tara turns messy workflows into systems people actually follow. Readers get practical frameworks they can use to cut waste, align teams, and grow on purpose.