The router reboots. The ERP goes into an unplanned maintenance window. The plant's internet link fails for an hour. On a serialized packaging line, should production stop? If not, how do you make sure nothing goes wrong while the line is disconnected — and that everything is put right when it reconnects?
This guide explains how a well-designed line handles outages, what limits apply, and which conditions should stop production regardless.
First, which link went down?
"The network is down" can mean several different things, with different consequences.
| What failed | Typical effect on a well-designed line |
|---|---|
| Internet or cloud connection | Production unaffected if the ERP is on-site and serials are buffered locally; cloud event delivery is delayed. With a cloud-hosted ERP, treat this as an ERP outage (next row) |
| ERP unavailable (link or system) | Production continues from local serials; registrations queue |
| Factory network between line software and devices | Line cannot print or verify — production stops |
| Line software PC or local storage | Line cannot track serials — production stops until recovered |
| Coder, camera or reject device | Line cannot serialize or verify safely — production stops |
The rest of this guide is mostly about the first two rows: losing the upstream systems while the line itself is healthy.
The three ingredients of riding through an outage
1. A local buffer of serials
If the line had to fetch each serial from the ERP, an ERP outage would stop it instantly. Instead, line software keeps a local pool of serials, downloaded in advance.
Think of a water tank on a roof. When the mains supply is cut, the house still has water — for as long as the tank lasts. The pool size, divided by the line rate, tells you how long that is. At 120 packs a minute, a pool of 7,200 serials lasts about an hour at full speed (less, once rejects are counted).
2. Local storage that survives restarts
Every serial status and carton record is written to local storage as it happens. If the outage is accompanied by a power blip or PC restart, the line software knows exactly which serials were sent, verified, rejected or packed.
3. A persistent outbox
Results bound for the ERP — "carton 0457 contains these twelve serials" — and for any cloud or traceability system are written into an outbox on local disk before sending. During an outage the outbox simply grows. Like a post tray on a desk, letters pile up but none is lost; when the post room opens, they go out.
When the connection comes back: reconciliation
Reconnection is where the real care is needed.
Deliver the backlog in order. Registrations go to the ERP oldest first, so that, for example, a carton is registered before the pallet that contains it.
Resolve the uncertain ones. Some messages may have been sent just as the connection failed, with no reply received. Line software should ask the ERP whether it already has them before resending, so nothing is registered twice. This is one of the core techniques of exactly-once serialization.
Separate errors from outages. If the ERP now rejects a registration for a business reason — perhaps the production order was closed during the outage — that is not a connectivity problem. It should become an exception for a person to resolve, not an endless retry.
Prove the totals. After the backlog clears, reconcile: serials issued should equal serials packed plus rejected plus returned unused. Carton and pallet contents (aggregation) should match on both sides.
Refill the pool. Once the ERP is reachable, top up the local serial pool so the line is ready for the next outage.
Local-first is not unlimited offline
It is easy to oversell offline capability. A line that can ride through a short outage cannot run disconnected forever. Practical limits include:
- Serial stock. When the local pool is empty, there is nothing to print.
- Storage health and capacity. The outbox and records must have room, and the disk must be healthy.
- Business tolerance. Finance, logistics or customers may need production data in the ERP within a certain time — for example, before goods can be shipped or invoiced.
- Change risk. The longer the line runs offline, the more can change upstream (cancelled orders, master-data updates), and the more reconciliation work builds up.
That is why each site should agree an offline operating window: how long production may continue without upstream systems, and what happens as it approaches the limit.
What should stop the line
Some conditions should stop production deliberately, whether or not the network is up:
- The serial pool is about to run out. Warn early; stop before printing would require serials that do not exist.
- Local storage is failing or full. If results cannot be recorded, they cannot be trusted later.
- The agreed offline window is reached. Continue only with an explicit, recorded decision by an authorised person.
- Camera or reject confirmation is unavailable. Without verification and confirmed rejects, unverified packs could reach cartons.
- Repeated mismatches. A sequence shift needs investigation, not more production.
- Unresolvable uncertainty. If the software cannot tell whether a serial was printed, it should void it and, if that keeps happening, stop.
A controlled stop with a clear reason is far cheaper than a day of unrecorded or duplicated serials.
A worked example: a two-hour ERP outage
At 09:00, Maison Ardent's fictional ERP goes down for unplanned maintenance. Line 3 runs at 120 perfume boxes a minute. The site's agreed offline window is 90 minutes, and the line software holds 12,000 serials — enough for about 100 minutes at full speed.
- 09:00–10:15 Production continues. Cartons are recorded locally; the outbox grows. The supervisor's screen shows the outage and backlog.
- 10:15 The line software warns that the offline window ends in 15 minutes. The production manager decides not to extend.
- 10:30 The line completes its current carton and stops in a controlled way. The team uses the time for a planned changeover.
- 11:00 The ERP returns. The outbox delivers nearly 900 carton registrations in order. Two that were in flight at 09:00 are checked against the ERP first; one had already arrived and is marked registered without resending.
- 11:10 Reconciliation balances. The pool is refilled. Production resumes.
Nothing was lost, nothing was doubled, and the stop was a decision rather than a surprise.
Planning checklist
- What pool size and offline window has the site agreed, and who can extend it?
- How are supervisors alerted, and at what thresholds?
- In what order is the backlog delivered, and how are uncertain messages resolved?
- How are business rejections surfaced and resolved?
- Who reviews reconciliation before batches are released?
- Have outages, power cuts and restarts been tested on the actual line?
For the local layer that makes all of this possible, see why edge software matters on a packaging line. For where the outage sits in the wider flow of stations, see the journey of a serialized product through a packaging line, and for how delayed events still form a coherent history once delivered, EPCIS explained through the journey of one product.