Network redundancy is proven only when the protection and control functions survive the defined failures—not when two green link LEDs are visible. A digital substation can lose a GOOSE trip through a wrong multicast filter, corrupt SV through a time error, or lose both “redundant” paths through one DC MCB. Effective verification starts with function dependencies and injects realistic failures while measuring application behavior.
This article provides an FMEA-based method for IEC 61850 station/process networks, including physical, protocol, time, configuration, cyber and maintenance failures; redundancy claims; instrumentation; FAT/SAT and operational monitoring.
1. Define the protected functions first
- Direct GOOSE trip and breaker-failure initiation.
- Blocking/permissive/intertrip schemes.
- SV-fed feeder, transformer or bus protection.
- Breaker/disconnector/earthing-switch interlocking.
- Bus transfer, load shedding and autoreclose.
- Station HMI/SCADA control and alarms.
- Time-aligned fault/event recording.
For each function, state required availability, transfer/operating time, acceptable degradation, safe response and maximum repair time. “Network shall be redundant” is not an acceptance criterion.
2. Build a dependency diagram
Trace source to physical result: sensor/contact → MU/process I/O/publisher → interface → link/switch/queue/redundancy node → subscriber/interface → application → output/trip circuit → breaker. Add DC power, time source, SCL/configuration, firmware, management and human maintenance dependencies.
- Mark components shared by Protection 1 and Protection 2.
- Mark paths that share a tray, patch panel, fire zone or supply.
- Identify single-attached nodes and RedBox/QuadBox dependencies.
- Identify a common grandmaster/GNSS antenna and switch.
- Identify common SCD templates/tool/engineer errors.
- Map alarms and diagnostics; a fault not visible can become a latent second-fault risk.
3. Failure categories
| Category | Examples | Potential effect |
|---|---|---|
| Physical | Fiber break, dirty connector, port, switch, RedBox, EMI, temperature | Loss, errors, path asymmetry |
| Power | DC MCB, PSU, brownout, common charger | Multiple IED/network nodes lost |
| Traffic | Congestion, microburst, storm, queue starvation | GOOSE delay, SV loss/jitter |
| Protocol | Duplicate/out-of-order/stale frames, redundancy table issue | Wrong/late application data |
| Configuration | VLAN, MAC/APPID, ConfRev, ExtRef, filter, PTP domain | Healthy links but wrong/missing function |
| Time | GNSS loss, grandmaster step, asymmetry, holdover expiry | SV phase error, bad chronology |
| Cyber/human | Rogue publisher, scan/storm, patch error, wrong cable/file | Loss, false action or common-mode outage |
4. Physical-layer failures
- Open one fiber/link at a time and at each meaningful location—not only at the easiest patch cord.
- Introduce marginal optical power/dirty connector where safe to verify error/alarm thresholds.
- Power-cycle each switch, MU network interface, RedBox and clock port.
- Check SFP type, wavelength, connector, optical budget and diagnostics.
- Inspect redundant routes for shared trays/patch panels.
- Test environmental alarms and device reboot time.
- Verify link flap cannot cause repeated GOOSE/SV loss or redundancy oscillation.
5. Power failures
A dual network fed from one auxiliary MCB is a common hidden mode. Verify each source, MCB/selectivity, converter, ride-through and grounding arrangement. Brownout/recovery can be more disruptive than a clean outage because devices reboot at different thresholds and times.
- Trip each network and IED DC feeder under controlled conditions.
- Test minimum/maximum auxiliary voltage and interruption/ride-through requirements.
- Observe startup traffic, table learning, PTP recovery and outputs.
- Confirm LAN A/B and duplicated protection use independent power where claimed.
- Check alarm before total failure where redundant PSU modules are used.
- Verify selective isolation so one failed node/PSU does not collapse a bus.
6. Congestion, queues and storms
- Generate representative continuous SV, simultaneous GOOSE event bursts, MMS reports/files, PTP and management traffic.
- Measure per-port and per-queue utilization, drops and latency.
- Test an unauthorized high-PCP source and rate/storm containment.
- Verify lower-priority management/file traffic cannot impair protection.
- Verify strict-priority traffic cannot starve required time/management indefinitely.
- Test unknown-multicast flooding/filter failure.
- Repeat with one PRP LAN or altered HSR topology carrying the surviving traffic.
7. GOOSE failure injection
- Drop one and bursts of retransmissions; observe final state and transfer time.
- Delay or reorder frames; verify stNum/sqNum handling.
- Stop publisher and verify timeAllowedToLive timeout plus safe application response.
- Restart publisher and subscriber; verify no false trip/control.
- Change ConfRev, dataset order, APPID, VLAN or source identity.
- Inject test/simulation/invalid quality and verify acceptance rules.
- Duplicate a publisher; verify detection and no ambiguous application state.
- Measure application-to-application and complete breaker response.
8. Sampled Values failure injection
- Introduce isolated and consecutive sample loss, duplicate and out-of-order frames.
- Change svID, APPID, ConfRev, rate, nominal frequency or dataset channel order.
- Apply scaling, polarity and phase errors.
- Set invalid/questionable/test/simulation quality.
- Remove synchronization, add time offset/drift and switch grandmaster.
- Fail one redundant path and one MU/interface as distinct cases.
- Verify each protection function’s block, restraint, fallback or continued-operation policy.
- Measure operating time and security during data degradation.
9. Time-system failures
| Test | Measure | Pass principle |
|---|---|---|
| GNSS/reference loss | Holdover offset/drift and quality | Within application budget for declared duration |
| Grandmaster change | Phase/time step and protection response | No unwanted operation; valid quality/fallback |
| PTP path failure | Offset/asymmetry on surviving path | Required accuracy maintained or safe degradation |
| Wrong/rogue master | Selection and alarm | Unauthorized source cannot silently take control |
| Recovery | Slew/step, sample alignment, event continuity | Controlled return without false operation |
10. PRP verification
- Verify every DANP and SAN/RedBox classification.
- Confirm LAN A and B are not accidentally interconnected.
- Open each A and B path/switch/interface and measure function performance.
- Check duplicate discard and supervision node tables/counters.
- Run each LAN alone at full required traffic.
- Inspect independent power and routes.
- Restore a path and check no storm, duplicate application action or PTP upset.
- Demonstrate alarm of a failed side while service remains available.
11. HSR verification
- Verify every DANH, SAN, RedBox and QuadBox role.
- Open each ring link and power/remove each relevant forwarding node.
- Measure longest surviving-path GOOSE latency and SV behavior.
- Check per-link traffic/queues in normal and open-ring topology.
- Test duplicate handling, node-table supervision and ring alarms.
- Evaluate ring insertion/removal maintenance procedure.
- Inject two selected failures to establish the boundary beyond the single-failure claim.
- Restore the ring and verify no loop/storm or time discontinuity.
12. Configuration failures are often more dangerous
A wrong VLAN or ExtRef can affect both redundant paths identically. Hardware N-1 testing will not reveal every common configuration failure. Perform automated and independent semantic review.
- Validate SCD schema, types, addresses, datasets and subscriptions.
- Compare SCL, IED settings and switch configuration to one approved matrix.
- Detect duplicate IP/MAC/APPID/VID and wrong adjacent-bay references.
- Review multicast filters and queue maps across A/B/ring.
- Test restore from backups on spare equipment.
- Use semantic diff after firmware/tool round trips.
- Require change ticket, independent review and regression evidence.
13. Cyber and human-induced failures
- Inject only controlled lab/test traffic: rogue multicast, duplicate identity, scan and rate burst.
- Verify access control, unused-port shutdown, logging and alerting.
- Test a wrong patch cord/port and confirm VLAN/port policy contains it.
- Confirm test laptop/simulation streams cannot reach unauthorized live outputs.
- Exercise patch/firmware rollback.
- Use two-person verification for critical SCD/network changes.
- Test incident diagnosis without relying on one unavailable vendor engineer.
14. Measurement and evidence
- Use calibrated time sources/test sets and state measurement uncertainty.
- Capture on relevant publisher/subscriber and A/B/ring points.
- Record IED events, GOOSE/SV counters, PTP state and switch port/queue data.
- Measure application response, not ping alone.
- Repeat enough operations to reveal outliers.
- Synchronize test instruments or document correlation limitations.
- Store raw captures/configurations with the signed test result.
15. FAT sequence
- Freeze requirements, topology, dependency/FMEA and pass limits.
- Validate SCD/device/switch/time configurations.
- Baseline all functions and traffic in healthy state.
- Inject every single physical/power/path failure.
- Test protocol, data-quality, time and configuration failures.
- Apply load/storm and one-failure combinations.
- Test restoration, reboot and backup restore.
- Verify alarms identify consequence and failed path.
- Close defects and repeat affected regression.
- Hash/archive the final release and raw evidence.
16. SAT and operational monitoring
- Inspect installed physical diversity, optics, labels and supplies.
- Compare loaded hashes/firmware with FAT.
- Repeat critical path, power, time and primary-to-breaker tests.
- Baseline port errors, drops, utilization, redundancy nodes, GOOSE timeouts, SV gaps and PTP offset.
- Alarm loss of one redundant path before the second failure.
- Trend intermittent optical/time/queue degradation.
- Periodically proof-test dormant paths and restoration.
- After expansion/change, redo traffic/FMEA and affected tests.
- Reconcile field changes into the master as-built set.
17. Acceptance checklist
- Redundancy claim names exact covered failures.
- Dependency diagram includes process, power, time and configuration.
- Every critical function has measurable pass criteria.
- Normal, N-1, loaded and restoration states are tested.
- Common-mode failures are identified and mitigated.
- Latent path failure creates an actionable alarm.
- GOOSE/SV/clock bad-data behavior is proven.
- Raw packet, event, configuration and timing evidence is archived.
- SAT proves installed routes and primary/physical outputs.
- Periodic testing and change regression have owners.
References and further reading
- IEC 62439-3:2021 — PRP and HSR seamless redundancy
- IEC TR 61850-10-3:2022 — Functional testing of IEC 61850 systems
- IEC 61850-10 edition 2.1 — Conformance testing and performance methods
- IEC TS 60255-216-1:2025 — Functional/interoperability protection testing
- IEC 61850-8-1 consolidated edition — GOOSE/MMS
- IEC 61850-9-2 consolidated edition — Sampled values
- IEC/IEEE 61850-9-3:2016 — PTP power profile
- IEC 62351-6:2020 — IEC 61850 communication security
Engineering note: Test failure while observing the protection function, not just the network. A seamless packet path that delivers the wrong dataset or wrong time is not a successful redundancy result.