Redundancy improves availability only when the redundant channels are independent, their state is synchronized and exactly one authority controls the process. Two servers on one power supply, two gateways behind one switch or dual WAN links through one carrier duct may look redundant but retain a single failure domain. Active/standby failover also creates event gaps, duplicates and command hazards unless application state is engineered.
This guide designs and tests redundant SCADA servers, station gateways, LANs and WAN communications for MV switchgear, with special attention to command ownership, event continuity, time and return-to-service.
1. Define the availability target by function
- Local protection and breaker-failure operation must normally survive complete SCADA loss.
- Bay/station local control availability may differ from remote control availability.
- Monitoring, alarms, SOE, historian and engineering access have different recovery/loss tolerances.
- Specify maximum interruption, allowed event loss/duplication and recovery behavior per function.
- State maintenance and common-mode assumptions; availability percentages alone are insufficient.
- Define degraded modes operators can recognize and safely use.
2. Map every failure domain
| Layer | Potential common mode | Independence question |
|---|---|---|
| Power | One DC board, fuse, UPS or AC charger | Are feeds/protection/cabling independent? |
| Compute | One chassis, hypervisor, storage or cluster manager | Can one platform fault stop both nodes? |
| LAN | One switch, fiber tray, configuration or broadcast domain | Do redundant paths fail together? |
| WAN | One router, carrier POP, tower, duct or firewall | Are route and provider physically diverse? |
| Time | One GNSS antenna/clock/distributor | Can both nodes retain coherent time? |
| Software/config | Same defect, bad database or update | Can rollback/diversity contain it? |
| Operations/cyber | One credential, orchestration or attack path | Does compromise propagate to both? |
3. Redundancy patterns
| Pattern | Benefit | Main engineering issue |
|---|---|---|
| Cold standby | Low cost and cyber isolation | Manual restore time and configuration freshness |
| Warm standby | Faster start with partial state | Event/command reconciliation after takeover |
| Hot active/standby | Short automatic failover | State replication and split-brain prevention |
| Active/active | Capacity and node continuity | Ownership, duplicate events/commands and partition behavior |
| Dual independent systems | Strong failure isolation | Operator/data reconciliation and cost |
| PRP/HSR endpoint | Zero network recovery time for a single network failure | Does not duplicate server/application/power |
4. SCADA server redundancy
- Synchronize real-time database, alarm acknowledgment/shelving, users/roles, time, configuration and application version.
- Decide whether both servers acquire data or only the active server maintains field sessions.
- Prevent duplicate alarm/event insertion and preserve original source timestamps.
- Define historian buffering/forward and gap detection.
- Keep client/HMI reconnection deterministic and show active/degraded status.
- Protect against cluster quorum/split-brain failure; loss of the inter-server heartbeat must not create two command masters.
- Test service/process crash separately from complete node/power/network loss.
5. Gateway redundancy
- Both gateways need the same approved point map/SCL/protocol profile and verified database revision.
- Synchronize IEC 61850 report-buffer positions, event queues, northbound sequence state and command ownership as the design requires.
- Define whether both subscribe to IED reports; ensure enough RCB instances and no ownership stealing.
- Coordinate IEC 104/DNP3 master/outstation sessions and general interrogation/integrity after takeover.
- Deduplicate events without discarding two real source transitions.
- Never replay or repeat a command whose outcome is uncertain; first reconcile primary position.
- Alarm standby, synchronization, database or redundant-link degradation.
6. LAN redundancy
- PRP uses two separate LANs and duplicated frames; HSR sends duplicated frames around a ring/mesh of suitable nodes.
- IEC 62439-3 specifies PRP/HSR seamless behavior for a network-element failure, but endpoints/RedBoxes and engineering must conform.
- Spanning-tree or routed redundancy has nonzero recovery that must meet the application budget and be measured.
- Keep A/B switches, fibers, DC feeds, patching and configuration failure domains independent.
- Test multicast, PTP, VLAN/QoS and management under one-network failure and restoration.
- Monitor duplicate-discard and path-health counters; seamless service can hide a failed path until the second failure.
7. WAN redundancy
- Require physical route/provider/last-mile/firewall/router diversity, not only two logical circuits.
- Define active/standby or load-sharing routing and application session behavior after path change.
- Account for different latency, bandwidth, MTU, NAT/firewall state and packet reordering.
- Verify IEC 104/DNP3/TLS session reconnect, backoff, integrity/interrogation and buffered event replay.
- Prevent dual paths from creating two active control-center authorities.
- Measure failover with an open control dialogue and an event burst—not just ping.
- Provide station-local operations procedure for complete WAN loss.
8. Command ownership and split-brain prevention
- Use a deterministic active authority agreed by server, gateway and bay IED layers.
- Require quorum/witness/lease or equivalent architecture where appropriate; understand its own failure modes.
- Do not transfer an ambiguous SBO selection or queued command to the standby.
- Carry unique origin/control identifiers and detect duplicates where protocols support them.
- Block or reconcile remote control when the pair cannot agree on state.
- Protection trips/local emergency control remain independent of SCADA cluster arbitration.
- Audit ownership transfer, command before/during/after failover and final equipment state.
9. Event continuity
- Define the authoritative event sequence/source and preserve its timestamp/quality.
- Use source/event identifiers, sequence/entry information and controlled deduplication.
- Size buffers for the longest credible outage and worst simultaneous event burst.
- On takeover, acquire current state plus recover buffered events without presenting old events as new current changes.
- Detect and alarm sequence gaps, report-buffer overflow and historian gaps.
- Test clock offset between redundant nodes; failover must not reverse event order by server time.
- Reconcile acknowledgments/shelving so operators do not receive a second alarm flood.
10. Failover versus failback
Automatic failover removes a failed node; failback reintroduces a repaired node and can be equally hazardous. Validate its software/database version, clock, event position, certificates and health before synchronization. Prefer controlled failback during an operational window when automatic return offers little benefit.
- Prevent stale standby state from overwriting current active state.
- Perform staged health checks and read-only synchronization before authority transfer.
- Maintain one command owner throughout.
- Alarm excessive failover oscillation/flapping and require manual stabilization.
- Record exact outage/failover/failback times and lost/degraded functions.
11. Cybersecurity and redundancy
- Redundant nodes should not share one uncontrolled privileged credential.
- Protect synchronization/heartbeat/quorum channels and management plane.
- Replicate only validated/signed configuration; malware or bad settings can otherwise spread instantly.
- Stagger patches only with temporary availability risk assessed and regression-tested.
- Keep recoverable offline backups beyond the live replicated cluster.
- Test certificate/key renewal and expiry on active/standby links.
- Design incident isolation so one suspect node/path can be removed while local protection remains.
12. Quantitative acceptance metrics
- Maximum monitoring/alarm interruption and client reconnection time.
- Maximum remote-control unavailability and prohibition window.
- Zero unintended/duplicate command requirement.
- Allowed missing/duplicate events and time/sequence reconciliation rule.
- Maximum data age after takeover and time to revalidate current state.
- Standby synchronization lag and buffer capacity.
- One-failure and maintenance availability assumptions.
- Maximum failback interruption and manual/automatic policy.
13. FAT/SAT failure campaign
- Freeze hardware, software, database/SCL, certificates, routing and time configuration.
- Establish normal-load and event-storm baselines.
- Fail one application service, process, server, gateway, NIC, switch, LAN, router, WAN, DC feed, clock and synchronization link.
- Partition the redundant peers to test split-brain/quorum behavior.
- Fail during an idle period, event burst, buffered replay, SBO selection, command execution and historian write.
- Verify event count/order/quality/time and no unintended/duplicate command.
- Restore/fail back each component and test stale-state rejection.
- Fail the standby silently, then the active, to prove latent-failure alarming matters.
- Apply worst load with one path/node unavailable.
- Test patch/upgrade rollback and restore a failed node from controlled backup.
- At SAT, repeat physical DC, fiber, WAN and equipment-control cases.
- Archive packet captures, logs, timings, state comparisons and residual common modes.
14. Final design rule
Redundancy is a complete stateful function, not a component count. Approve it only when power, compute, network, time, software, configuration and authority failure domains are documented and the real application has proven event continuity and safe control under every single failure and restoration.
References and further reading
- IEC 62439-3:2021 — PRP and HSR seamless network redundancy
- IEC 62439-1 consolidated edition — High-availability network concepts
- IEC TR 61850-90-4:2020 — Substation network engineering
- IEC 61850-8-1 consolidated edition — MMS/GOOSE reporting and control context
- IEC 60870-5-104 consolidated edition — Telecontrol over IP
- IEC 62351-4 consolidated edition — Security for MMS and derivatives
- IEC 62443-3-3:2013 — System security requirements
Engineering note: A seamless network failover that leaves two command masters or loses the event buffer is not a successful SCADA failover.