PRP and HSR can deliver a valid frame despite one network failure without waiting for a topology-recovery protocol, but “zero recovery time” is not zero latency, zero risk or guaranteed delivery under every failure. Both methods duplicate traffic and depend on correct node behavior, addressing, supervision, bandwidth, power, physical diversity and configuration. A common switch supply, multicast storm or duplicated SCL error can defeat a design whose marketing diagram shows two paths.
This IEC 62439-3 engineering guide explains PRP/HSR operation, DAN/SAN/RedBox/QuadBox roles, traffic and latency calculations, PTP and IEC 61850 behavior, failure coverage, testing, cybersecurity and practical selection for MV digital substations.
1. What seamless redundancy means
A redundancy-capable source sends duplicate copies over disjoint paths. The destination accepts the first valid copy and discards the duplicate. If one path fails, the other copy can still arrive; the application therefore need not wait for RSTP or another reconvergence process. The claim applies to covered single failures under a correctly engineered network.
- There is still normal forwarding, serialization and processing latency.
- A frame can still be lost if both paths or a common endpoint fail.
- Congestion, corruption, wrong VLAN/filtering or application overload can affect both copies.
- Seamless delivery may hide a failed path; supervision and alarm handling are mandatory.
- Network redundancy does not duplicate sensors, merging units, relay applications, process outputs, DC supplies or breaker coils.
2. PRP operation
Parallel Redundancy Protocol (PRP) uses two independent LANs, commonly called LAN A and LAN B. A double-attached PRP node (DANP) connects to both and sends a copy of each frame through each LAN. The destination DANP uses redundancy information to accept the first copy and discard the second.
- LAN A and LAN B should be independent in switches, power, fibers and routes; interconnecting them incorrectly destroys the parallel architecture.
- Ordinary single-attached nodes (SANs) can reside on one LAN but are not themselves protected against loss of that LAN.
- A RedBox can provide redundant attachment for SAN equipment or connect redundancy domains, subject to its ports, forwarding and failure behavior.
- Each LAN must independently carry the full required traffic because, after one failure, only one valid copy remains.
- PRP usually requires more switch hardware/cabling but gives a clear physical A/B model and avoids ring-forwarding accumulation.
3. HSR operation
High-availability Seamless Redundancy (HSR) connects double-attached HSR nodes (DANH) in a ring. A source sends copies in both ring directions. Nodes forward frames according to HSR rules and the destination accepts one while duplicate handling prevents endless circulation.
- A single link or intervening node failure can leave an alternate direction around the ring.
- Every ring node participates in forwarding, so node processing and link load affect the complete ring.
- Traffic passes through multiple hops; worst-path latency grows with ring size and node implementation.
- All duplicated/forwarded multicast traffic, including SV and GOOSE, must be included in per-link capacity.
- SAN devices require a RedBox or other designed interface.
- QuadBox functions can connect rings; they introduce important capacity, configuration and common-mode considerations.
4. PRP versus HSR
| Criterion | PRP | HSR |
|---|---|---|
| Topology | Two independent LANs | Dual-attached ring |
| Infrastructure | Usually two switch fabrics | Ring nodes forward traffic; fewer central switches may be possible |
| Traffic location | Full traffic on each LAN | Copies travel/are forwarded around ring segments |
| Latency scaling | Path depends on LAN switch hops | Potentially accumulates with ring hops |
| Expansion | Add to both independent LANs | Ring insertion affects links/path and outage procedure |
| Fault domain clarity | Clear A/B physical domains if built independently | One ring; node/link/QuadBox behavior is central |
| Typical fit | Station backbone and high independence | Bay/process rings and compact distributed layouts |
Neither protocol is universally superior. Select from physical layout, critical functions, node support, hop count, traffic, maintenance and common-mode requirements.
5. Node terminology and engineering significance
- DANP: PRP double-attached node that creates/handles duplicates on two LANs.
- DANH: HSR double-attached node that sends/forwards/handles ring traffic.
- SAN: ordinary single-attached node with no native seamless redundancy.
- RedBox: redundancy box that gives SANs access to a PRP/HSR domain or performs approved interconnection.
- QuadBox: function used to interconnect HSR rings through multiple ports/paths.
- VDAN: virtual double-attached representation enabled through a RedBox arrangement.
The bill of materials must state the exact role and redundancy behavior of every IED, switch and box. “Two Ethernet ports” does not prove PRP or HSR support; ports may be an internal switch, separate station/process interfaces or a non-seamless redundancy option.
6. Duplicate handling and supervision
Nodes maintain information used to recognize duplicate frames and supervise redundant paths. Implementation resource limits, aging/forget behavior, restarts and unusual traffic rates must be considered. Duplicate discard should occur before the application sees two GOOSE events or SV samples.
- Verify duplicate-discard behavior for unicast, multicast, GOOSE, SV and time traffic used by the project.
- Monitor supervision frames/node tables and expose LAN/ring faults to station alarms.
- Define alarm delay to avoid chatter yet reveal a lost path promptly.
- Test a device reboot and table relearning without duplicate application actions.
- Check maximum supported nodes/entries/traffic and behavior under table exhaustion.
- Provide diagnostics that identify which LAN, ring direction, port or node is impaired.
7. Bandwidth calculation
Build a traffic matrix using actual on-wire frame sizes and rates. Include Ethernet overhead as required by the calculation method, SV streams, stable and event GOOSE, MMS disturbance uploads, PTP, supervision, network management, security logging and future margin.
- PRP: calculate LAN A and B independently. Each must carry the complete duplicated source set and meet performance if the other is unavailable.
- HSR: calculate every link/direction for the actual source/destination/multicast forwarding behavior. Do not divide aggregate traffic by two and assume even distribution.
- Include simultaneous fault-event retransmissions and disturbance-report transfers.
- Evaluate switch/node egress queues, not only nominal link utilization.
- Allow for multicast flooding if filtering/snooping fails or is unsupported.
- Test overload/policing thresholds; storm control must not discard valid protection multicast.
For a stream, a first-order bit-rate estimate is frame rate × on-wire bits per frame. The engineering model must then apply duplication and every forwarding path. Use measured captures to confirm assumptions, while recognizing that captures can omit physical-layer overhead or duplicates depending on measurement point.
8. Latency and jitter
- PRP delivery normally follows the faster surviving A/B path; path asymmetry affects which copy wins.
- HSR copy arrival depends on direction and hop count; largest ring and failed-link topology may create the longest valid path.
- Serialization, store-and-forward/cut-through, queues, RedBox/QuadBox and end-node processing all contribute.
- Test the application-to-application GOOSE transfer time and SV data behavior under normal, one-failure and loaded conditions.
- Measure distributions/outliers, not one best result.
- Confirm duplicate arrivals do not add subscriber CPU load or jitter beyond specification.
- Include recovery to fully redundant operation; reconnection should not cause storms, loops or time steps.
9. Failure coverage—and exclusions
| Failure | Expected coverage when correctly designed | Important limitation |
|---|---|---|
| One PRP LAN link/switch | Other LAN copy remains | Only if LANs and endpoint interfaces are independent/healthy |
| One HSR ring link/node path | Opposite direction remains | Ring segmentation/multiple faults may isolate nodes |
| DAN interface port | Other interface/path can remain | Shared device CPU/power/configuration remains |
| End IED/MU failure | Not covered by network duplication | Requires independent application/device/source |
| Common DC supply | Not covered | Separate power and protection required |
| Wrong VLAN/SCL on both paths | Not covered | Independent review and functional test |
| Storm/overload common to domain | May affect both copies | Capacity, containment, filtering and cyber monitoring |
| Two simultaneous path failures | Not generally covered by single-failure claim | Analyze topology-specific combinations |
10. Physical independence
- Separate PRP A/B switches, power feeders, DC MCBs, patch panels and fiber routes.
- Avoid both fibers in one tray, duct, fire zone or connector module where independence is claimed.
- For HSR, arrange node power/maintenance so one intervention cannot open two adjacent segments.
- Check RedBox/QuadBox power and placement; a single box can become a shared choke point.
- Use clear A/B or ring-direction labels and color/port conventions.
- Include grounding, EMC, bend radius, connector cleanliness and optical budget.
- Inspect as-built routes—drawings alone do not prove diversity.
11. PTP and time behavior
PRP/HSR duplication does not automatically provide correct PTP behavior. End devices, clocks, RedBoxes/QuadBoxes and switches must support the selected PTP power profile and redundancy architecture. Duplicate timing messages, path asymmetry and grandmaster selection can affect time quality.
- Document clock message paths and which devices are transparent/boundary/ordinary clocks.
- Verify PRP A/B handling and HSR forwarding for event/general messages.
- Measure time offset during each link/node failure and restoration.
- Test grandmaster failover separately from network-path failure.
- Define MU/IED holdover and protection response to unsynchronized samples.
- Monitor grandmaster identity, time quality, path delay and offset at critical publishers/subscribers.
12. IEC 61850 multicast engineering
- Allocate unique GOOSE/SV multicast MAC, APPID, VLAN and priority.
- Ensure both PRP LANs or every required HSR segment carry each publisher/subscriber group.
- Validate multicast filtering/registration after switch, RedBox and IED reboot.
- Confirm duplicate discard happens below the subscriber application.
- Test GOOSE stNum/sqNum and SV smpCnt diagnostics during path failures.
- Capture at multiple points; one capture port can see only one copy or alter traffic visibility.
- Verify test/simulation streams cannot escape into operational subscribers.
13. Cybersecurity considerations
- Redundancy duplicates malicious or malformed traffic as readily as valid traffic.
- Segment protection domains and control access to both LANs/ring.
- Harden switch, RedBox, clock and IED management; change defaults and use least privilege.
- Detect unauthorized publishers, duplicate identities, node-table anomalies and abnormal multicast rates.
- Coordinate IEC 62351 controls with device support, keys/certificates and performance tests.
- Rate controls must distinguish invalid storms from legitimate simultaneous protection events.
- Back up configurations and maintain tested rollback after firmware/security updates.
14. FAT program
- Freeze the topology, node roles, address/VLAN plan, SCD and device firmware/configuration.
- Verify every device is truly DANP/DANH/SAN/RedBox/QuadBox as designed.
- Measure normal duplicate arrival, discard, latency and supervision.
- Open each PRP LAN link/switch/port and each HSR link/node path one at a time.
- Fail RedBoxes/QuadBoxes, power supplies and redundant endpoint ports.
- Measure GOOSE transfer and SV/protection behavior during and after each failure.
- Apply realistic simultaneous SV, GOOSE bursts, MMS/report and management traffic.
- Test time synchronization, grandmaster changeover and path asymmetry.
- Inject duplicate, out-of-order, corrupted/lost and abnormal-rate frames with controlled tools.
- Verify alarms for latent path failure, event logs and operator diagnosis.
- Test restoration/reconnection and confirm no loop, storm or unwanted application action.
- Archive packet captures, calibrated timing, configurations and acceptance evidence.
15. SAT and maintenance
- Inspect physical A/B/ring routes, labels, supplies, optics and port mapping.
- Compare all loaded files/firmware with the FAT release.
- Repeat every single installed path failure while critical functions are monitored.
- Prove supervision alarms reach the correct HMI/maintenance system.
- Baseline node tables, path counters, port errors, time and transfer performance.
- Periodically proof-test each path; seamless service can hide a failed side indefinitely.
- After adding a bay/node, recalculate link/queue load and worst-path latency.
- After firmware, VLAN, SCL, time or cyber changes, perform impact-based end-to-end regression.
- Keep compatible spares and tested device/configuration restore procedures.
16. Selection checklist
- Required single-failure coverage and independence are explicitly stated.
- All devices’ native redundancy modes and limitations are verified.
- Topology matches physical station/bay layout and expansion plan.
- Per-link/queue traffic and worst-path latency meet SV/GOOSE requirements.
- PTP profile and clock behavior are proven through redundancy events.
- Common power, route, box, configuration and cyber failures are mitigated.
- Supervision exposes latent loss of either path.
- Operations can isolate, add, replace and restore nodes safely.
- FAT/SAT covers failure and restoration with real applications.
- Lifecycle tools, training, spares and support are funded.
References and further reading
- IEC 62439-3:2021 with corrigendum — PRP and HSR
- IEC 61850-8-1 consolidated edition — MMS and GOOSE mapping
- IEC 61850-9-2 consolidated edition — Sampled values
- IEC 61850-5 consolidated edition — Communication requirements
- IEC 61850-10 consolidated through Amendment 1:2025 — Conformance testing
- IEC TR 61850-10-3:2022 — Functional testing
- IEC 62351-6:2020 — IEC 61850 communication security
Engineering note: State redundancy claims precisely: “no application interruption for the specified single network failure” is testable. “Zero downtime network” is not. The accepted design must show which failures are covered, which remain common and how a latent failed path becomes visible.