Digital-Substation Network Failure Modes and Redundancy Verification

A complete fault-injection and evidence program for IEC 61850 network, PRP/HSR, GOOSE, SV and time-system failure modes from FAT to maintenance.

Network redundancy is proven only when the protection and control functions survive the defined failures—not when two green link LEDs are visible. A digital substation can lose a GOOSE trip through a wrong multicast filter, corrupt SV through a time error, or lose both “redundant” paths through one DC MCB. Effective verification starts with function dependencies and injects realistic failures while measuring application behavior.

This article provides an FMEA-based method for IEC 61850 station/process networks, including physical, protocol, time, configuration, cyber and maintenance failures; redundancy claims; instrumentation; FAT/SAT and operational monitoring.

1. Define the protected functions first

  • Direct GOOSE trip and breaker-failure initiation.
  • Blocking/permissive/intertrip schemes.
  • SV-fed feeder, transformer or bus protection.
  • Breaker/disconnector/earthing-switch interlocking.
  • Bus transfer, load shedding and autoreclose.
  • Station HMI/SCADA control and alarms.
  • Time-aligned fault/event recording.

For each function, state required availability, transfer/operating time, acceptable degradation, safe response and maximum repair time. “Network shall be redundant” is not an acceptance criterion.

2. Build a dependency diagram

Trace source to physical result: sensor/contact → MU/process I/O/publisher → interface → link/switch/queue/redundancy node → subscriber/interface → application → output/trip circuit → breaker. Add DC power, time source, SCL/configuration, firmware, management and human maintenance dependencies.

  • Mark components shared by Protection 1 and Protection 2.
  • Mark paths that share a tray, patch panel, fire zone or supply.
  • Identify single-attached nodes and RedBox/QuadBox dependencies.
  • Identify a common grandmaster/GNSS antenna and switch.
  • Identify common SCD templates/tool/engineer errors.
  • Map alarms and diagnostics; a fault not visible can become a latent second-fault risk.

3. Failure categories

CategoryExamplesPotential effect
PhysicalFiber break, dirty connector, port, switch, RedBox, EMI, temperatureLoss, errors, path asymmetry
PowerDC MCB, PSU, brownout, common chargerMultiple IED/network nodes lost
TrafficCongestion, microburst, storm, queue starvationGOOSE delay, SV loss/jitter
ProtocolDuplicate/out-of-order/stale frames, redundancy table issueWrong/late application data
ConfigurationVLAN, MAC/APPID, ConfRev, ExtRef, filter, PTP domainHealthy links but wrong/missing function
TimeGNSS loss, grandmaster step, asymmetry, holdover expirySV phase error, bad chronology
Cyber/humanRogue publisher, scan/storm, patch error, wrong cable/fileLoss, false action or common-mode outage

4. Physical-layer failures

  • Open one fiber/link at a time and at each meaningful location—not only at the easiest patch cord.
  • Introduce marginal optical power/dirty connector where safe to verify error/alarm thresholds.
  • Power-cycle each switch, MU network interface, RedBox and clock port.
  • Check SFP type, wavelength, connector, optical budget and diagnostics.
  • Inspect redundant routes for shared trays/patch panels.
  • Test environmental alarms and device reboot time.
  • Verify link flap cannot cause repeated GOOSE/SV loss or redundancy oscillation.

5. Power failures

A dual network fed from one auxiliary MCB is a common hidden mode. Verify each source, MCB/selectivity, converter, ride-through and grounding arrangement. Brownout/recovery can be more disruptive than a clean outage because devices reboot at different thresholds and times.

  • Trip each network and IED DC feeder under controlled conditions.
  • Test minimum/maximum auxiliary voltage and interruption/ride-through requirements.
  • Observe startup traffic, table learning, PTP recovery and outputs.
  • Confirm LAN A/B and duplicated protection use independent power where claimed.
  • Check alarm before total failure where redundant PSU modules are used.
  • Verify selective isolation so one failed node/PSU does not collapse a bus.

6. Congestion, queues and storms

  • Generate representative continuous SV, simultaneous GOOSE event bursts, MMS reports/files, PTP and management traffic.
  • Measure per-port and per-queue utilization, drops and latency.
  • Test an unauthorized high-PCP source and rate/storm containment.
  • Verify lower-priority management/file traffic cannot impair protection.
  • Verify strict-priority traffic cannot starve required time/management indefinitely.
  • Test unknown-multicast flooding/filter failure.
  • Repeat with one PRP LAN or altered HSR topology carrying the surviving traffic.

7. GOOSE failure injection

  • Drop one and bursts of retransmissions; observe final state and transfer time.
  • Delay or reorder frames; verify stNum/sqNum handling.
  • Stop publisher and verify timeAllowedToLive timeout plus safe application response.
  • Restart publisher and subscriber; verify no false trip/control.
  • Change ConfRev, dataset order, APPID, VLAN or source identity.
  • Inject test/simulation/invalid quality and verify acceptance rules.
  • Duplicate a publisher; verify detection and no ambiguous application state.
  • Measure application-to-application and complete breaker response.

8. Sampled Values failure injection

  • Introduce isolated and consecutive sample loss, duplicate and out-of-order frames.
  • Change svID, APPID, ConfRev, rate, nominal frequency or dataset channel order.
  • Apply scaling, polarity and phase errors.
  • Set invalid/questionable/test/simulation quality.
  • Remove synchronization, add time offset/drift and switch grandmaster.
  • Fail one redundant path and one MU/interface as distinct cases.
  • Verify each protection function’s block, restraint, fallback or continued-operation policy.
  • Measure operating time and security during data degradation.

9. Time-system failures

TestMeasurePass principle
GNSS/reference lossHoldover offset/drift and qualityWithin application budget for declared duration
Grandmaster changePhase/time step and protection responseNo unwanted operation; valid quality/fallback
PTP path failureOffset/asymmetry on surviving pathRequired accuracy maintained or safe degradation
Wrong/rogue masterSelection and alarmUnauthorized source cannot silently take control
RecoverySlew/step, sample alignment, event continuityControlled return without false operation

10. PRP verification

  • Verify every DANP and SAN/RedBox classification.
  • Confirm LAN A and B are not accidentally interconnected.
  • Open each A and B path/switch/interface and measure function performance.
  • Check duplicate discard and supervision node tables/counters.
  • Run each LAN alone at full required traffic.
  • Inspect independent power and routes.
  • Restore a path and check no storm, duplicate application action or PTP upset.
  • Demonstrate alarm of a failed side while service remains available.

11. HSR verification

  • Verify every DANH, SAN, RedBox and QuadBox role.
  • Open each ring link and power/remove each relevant forwarding node.
  • Measure longest surviving-path GOOSE latency and SV behavior.
  • Check per-link traffic/queues in normal and open-ring topology.
  • Test duplicate handling, node-table supervision and ring alarms.
  • Evaluate ring insertion/removal maintenance procedure.
  • Inject two selected failures to establish the boundary beyond the single-failure claim.
  • Restore the ring and verify no loop/storm or time discontinuity.

12. Configuration failures are often more dangerous

A wrong VLAN or ExtRef can affect both redundant paths identically. Hardware N-1 testing will not reveal every common configuration failure. Perform automated and independent semantic review.

  • Validate SCD schema, types, addresses, datasets and subscriptions.
  • Compare SCL, IED settings and switch configuration to one approved matrix.
  • Detect duplicate IP/MAC/APPID/VID and wrong adjacent-bay references.
  • Review multicast filters and queue maps across A/B/ring.
  • Test restore from backups on spare equipment.
  • Use semantic diff after firmware/tool round trips.
  • Require change ticket, independent review and regression evidence.

13. Cyber and human-induced failures

  • Inject only controlled lab/test traffic: rogue multicast, duplicate identity, scan and rate burst.
  • Verify access control, unused-port shutdown, logging and alerting.
  • Test a wrong patch cord/port and confirm VLAN/port policy contains it.
  • Confirm test laptop/simulation streams cannot reach unauthorized live outputs.
  • Exercise patch/firmware rollback.
  • Use two-person verification for critical SCD/network changes.
  • Test incident diagnosis without relying on one unavailable vendor engineer.

14. Measurement and evidence

  • Use calibrated time sources/test sets and state measurement uncertainty.
  • Capture on relevant publisher/subscriber and A/B/ring points.
  • Record IED events, GOOSE/SV counters, PTP state and switch port/queue data.
  • Measure application response, not ping alone.
  • Repeat enough operations to reveal outliers.
  • Synchronize test instruments or document correlation limitations.
  • Store raw captures/configurations with the signed test result.

15. FAT sequence

  1. Freeze requirements, topology, dependency/FMEA and pass limits.
  2. Validate SCD/device/switch/time configurations.
  3. Baseline all functions and traffic in healthy state.
  4. Inject every single physical/power/path failure.
  5. Test protocol, data-quality, time and configuration failures.
  6. Apply load/storm and one-failure combinations.
  7. Test restoration, reboot and backup restore.
  8. Verify alarms identify consequence and failed path.
  9. Close defects and repeat affected regression.
  10. Hash/archive the final release and raw evidence.

16. SAT and operational monitoring

  • Inspect installed physical diversity, optics, labels and supplies.
  • Compare loaded hashes/firmware with FAT.
  • Repeat critical path, power, time and primary-to-breaker tests.
  • Baseline port errors, drops, utilization, redundancy nodes, GOOSE timeouts, SV gaps and PTP offset.
  • Alarm loss of one redundant path before the second failure.
  • Trend intermittent optical/time/queue degradation.
  • Periodically proof-test dormant paths and restoration.
  • After expansion/change, redo traffic/FMEA and affected tests.
  • Reconcile field changes into the master as-built set.

17. Acceptance checklist

  • Redundancy claim names exact covered failures.
  • Dependency diagram includes process, power, time and configuration.
  • Every critical function has measurable pass criteria.
  • Normal, N-1, loaded and restoration states are tested.
  • Common-mode failures are identified and mitigated.
  • Latent path failure creates an actionable alarm.
  • GOOSE/SV/clock bad-data behavior is proven.
  • Raw packet, event, configuration and timing evidence is archived.
  • SAT proves installed routes and primary/physical outputs.
  • Periodic testing and change regression have owners.

References and further reading

Engineering note: Test failure while observing the protection function, not just the network. A seamless packet path that delivers the wrong dataset or wrong time is not a successful redundancy result.

LearnSwitchgear

Search the engineering library