An alarm is not every bit that changes state. An alarm tells an operator about an abnormal condition that requires a defined response within a useful time. Status shows current condition; an event records a transition; sequence-of-events (SOE) data preserves time order for diagnosis. Mixing all four produces alarm floods, hidden trip causes and unreliable post-fault analysis.
This practical guide develops alarm, indication and SOE design for MV switchgear LVCs. It covers philosophy, priority, first-out, latching/acknowledgement/reset, signal conditioning, device health, VT/DC/network alarms, timestamps, IEC 61850 quality, HMI text, alarm floods, maintenance shelving, cybersecurity, FAT/SAT and performance review. IEC 62682:2022 is written for process industries, but its lifecycle principles are useful when contractually adapted to power-system operations.
Executive conclusions
- Classify every point as command, status, event, alarm, protection operation or diagnostic before assigning HMI priority.
- An alarm must have a credible abnormal condition, consequence, operator response and response time. If no action exists, use an event/status or engineering diagnostic.
- Priority follows consequence and available response time—not which engineering department owns the signal.
- Capture protection trip cause and breaker operation at the source with synchronized timestamps; gateway scan order is not authoritative SOE.
- Use first-out/trigger information without hiding subsequent causes. Preserve full event history and oscillography.
- Define normal, asserted, de-energized, open-wire and communication-failed states for each alarm.
- Apply debounce, pickup/dropout delay, deadband and hysteresis only with a documented physical/time basis; excessive filtering destroys event order.
- Acknowledge means the operator has seen an alarm; it does not mean the condition is normal. Reset must not erase an active hazard.
- Monitor DC supply, trip circuit, VT circuits, IED self-supervision, time source and network paths as distinct availability layers.
- FAT/SAT must test alarm occurrence, text, priority, audible/visual behavior, timestamp, quality, latching, acknowledgement, reset, loss and restoration end-to-end.
1. Status, event, alarm and SOE
| Data type | Purpose | Example |
|---|---|---|
| Status/indication | Shows current condition | Breaker open, spring charged, Local mode |
| Event | Records a transition for history/diagnosis | 52a changed, setting group selected |
| Alarm | Calls for operator response to abnormal condition | Trip circuit failed, DC low, protection IED failed |
| Protection operation | Records protection start/trip and affected zone | Bus differential trip, feeder 50/51 pickup |
| SOE record | Time-ordered event list with source timestamp/quality | Relay trip → coil energized → 52a open → current zero |
| Diagnostic | Supports engineering/maintenance | GOOSE subscriber timeout counter, fan runtime |
The same physical input can feed several uses. Breaker 52a is a status, its transition is an event, an implausible 52a/52b combination can become an alarm, and its timestamp helps SOE analysis.
2. Alarm philosophy and lifecycle
- Define objectives, roles, priorities, display/annunciation and performance targets.
- Identify candidate alarms from protection, operating, safety, asset and availability studies.
- Rationalize each alarm: cause, consequence, operator action, response time, priority, suppression and documentation.
- Implement signal, logic, text, HMI, historian and SOE mapping under configuration control.
- Verify in FAT/SAT and train operators with realistic scenarios.
- Monitor nuisance, standing, chattering and flood performance.
- Maintain, modify and periodically review as the plant/network changes.
IEC 62682 provides a structured alarm-management lifecycle for process facilities. Power utilities should adapt it to protection speed, substation staffing, remote operation, safety rules and grid-code obligations rather than copy process-industry metrics blindly.
3. Rationalization fields
- unique alarm ID and equipment/function reference;
- initiating physical/logical signal and normal state;
- abnormal cause(s) and operating state in which it is relevant;
- consequence if no action is taken;
- required operator action and maximum useful response time;
- priority and basis;
- pickup/dropout delay, deadband and latching;
- acknowledge/reset/shelve/suppress rules;
- alarm text, help text and navigation;
- SOE timestamp source, resolution/accuracy and quality;
- dependencies: breaker mode, maintenance, VT selection, communications;
- FAT/SAT stimulus and acceptance result.
4. Priority based on consequence and time
A practical matrix combines consequence severity with time available for response. Use a limited number of priorities so operators can distinguish them.
| Priority concept | Typical use | Presentation |
|---|---|---|
| Critical/emergency | Immediate safety/system consequence requiring urgent action | Distinct audible/visual treatment and clear response |
| High | Protection/control availability degraded with short response window | Prominent annunciation and escalation |
| Medium | Abnormal condition needing timely operational/maintenance action | Standard alarm presentation |
| Low/advisory | Useful non-urgent abnormal condition | Limited annunciation; not allowed to flood |
| Event/diagnostic | No immediate operator response | Event log/maintenance view, not alarm banner |
A breaker protection trip is important but often already complete before an operator can act. The post-trip alarm should identify the tripped equipment/cause and required response; it is not automatically the same priority as a developing condition that can still be prevented.
5. Alarm text that supports action
- Begin with unique equipment/location, then abnormal condition: “11 kV Feeder F12 — Trip circuit 1 failed.”
- Use one controlled vocabulary for failed, unavailable, low, open, mismatch and communication loss.
- Do not use ambiguous text such as “Relay alarm” or “CB problem.”
- State channel/source where redundancy exists: Main 1, DC1, LAN A, VT bus 2.
- Separate cause from consequence; “IED watchdog failed” is not “protection unavailable” unless the logic proves it.
- Provide operator help: verification, safe first action, escalation and relevant drawing/procedure.
- Use consistent active/return-to-normal wording and multilingual governance where required.
- Keep raw IEC 61850 object names available for engineering, but present human-readable operational text.
6. First-out, trip cause and sequence
First-out logic captures the first initiating cause in a cascade. It is useful for transformer lockout, bus trips and complex interlocks, but it must not discard subsequent protection operations.
- Latch the first cause at the source or deterministic logic layer with synchronized time.
- Record every other start/trip/output as independent SOE events.
- Do not reset first-out automatically on breaker opening if investigation requires it.
- Define reset authority and confirm initiating conditions are clear.
- Map lockout 86 operation separately from the protection element that operated it.
- Store fault report/oscillography/COMTRADE references with the event if the system supports linkage.
- Keep breaker-failure initiation, retrip and backtrip distinguishable.
7. Latching, acknowledgement and reset
| Action | Meaning | Must not do |
|---|---|---|
| Acknowledge | Operator has seen/accepted responsibility | Hide an active alarm or restore equipment |
| Return to normal | Initiating condition cleared | Erase history |
| Reset | Clear a latched indication/logic when safe | Override an active cause or auto-close equipment |
| Shelve | Temporarily suppress presentation under authorization | Permanently disable or remove history |
| Suppress by design/state | Alarm not meaningful in declared operating state | Mask unexpected failures outside that state |
Local annunciator “lamp reset” and protection/86 reset are different operations. Keep safety/protection resets deliberate, access-controlled and visible in SOE.
8. Signal conditioning and chatter
- Use hardware/logic debounce for contact bounce, with time short enough to preserve required sequence.
- Use analogue deadband/hysteresis for temperature/voltage thresholds.
- Use pickup/dropout delay based on physical persistence and operator usefulness.
- Do not delay trip-circuit failure or DC loss so long that redundancy remains unknowingly degraded.
- For 52a/52b mismatch, allow normal mechanism transition but alarm a persistent contradictory state.
- Detect and flag chattering inputs; do not merely extend the filter until the symptom disappears.
- Timestamp at the earliest reliable layer after approved input filtering.
9. SOE timestamp architecture
Useful SOE requires a stated time chain:
- Record where the timestamp is generated and its resolution/accuracy.
- Use source timestamps rather than gateway scan/arrival time for protection sequencing.
- Monitor time synchronization quality, grandmaster/source and holdover.
- Preserve IEC 61850 quality and timestamp attributes through gateways.
- Indicate unsynchronized/invalid time in the event record; never silently display it as precise.
- Account for input debounce and auxiliary-contact mechanical timing when comparing events.
- Maintain UTC/local-time and daylight-saving presentation rules without altering stored sequence.
10. Breaker and mechanism indications
- breaker open (52b) and closed (52a), with intermediate/mismatch logic;
- truck/service/test/disconnected position and invalid position;
- earthing-switch position and disagreement;
- spring charged/mechanism ready and charging timeout;
- trip coil/TCS 1 and 2 healthy;
- close circuit/anti-pumping/close supply availability as required;
- low gas/density/pressure or vacuum/mechanism diagnostics where provided;
- operation counter, motor runtime and wear/maintenance threshold.
Do not treat one auxiliary contact as proof of primary isolation. Position indications support control/diagnosis but safe work uses the approved isolation and verification procedure.
11. Protection, VT and DC alarms
| Layer | Examples | Operational meaning |
|---|---|---|
| Protection operation | Element start/trip, 86, BF | Fault clearing and affected equipment |
| Protection health | IED watchdog, settings/config error, input board fault | Function availability degraded |
| Measuring circuit | VT fuse/MCB failure, CT circuit alarm, sensor fail | Security/dependability of voltage/current functions |
| Station DC | Low/high DC, charger fail, earth fault, branch MCB | Trip/control availability risk |
| Communications/time | GOOSE timeout, LAN path fail, clock unsynchronized | Digital function/SOE quality degraded |
| Environment | High temperature, fan fail, humidity/condensation | Equipment life/availability risk |
Alarm the layer actually proven. One LAN A failure on a healthy PRP pair means redundancy degraded, not “protection failed.” Conversely, do not label a healthy IED as complete protection health if its CT circuit or trip coil is open.
12. IEC 61850 quality and events
- Map semantic data objects/logical nodes rather than opaque register names.
- Carry quality fields such as validity, source/test and operator blocking according to the engineering design.
- Do not generate a process alarm from invalid/stale data without distinguishing communication/data quality.
- Handle test/simulation values so they cannot appear as live plant alarms.
- Use buffered reporting/logs as required to retain events through communication outages.
- Define duplicate/event-sequence handling through gateways.
- Record configuration/SCL revision so event meanings remain traceable.
13. Alarm floods, standing and nuisance alarms
- Flood: many alarms arrive faster than an operator can interpret; use state-based suppression, root-cause grouping and consequence review.
- Standing: active for long periods; repair, rationalize or formally manage—not normalize it.
- Chattering: repeats active/clear; fix contact, threshold, EMC or process cause.
- Fleeting: clears before response; retain event/SOE and decide whether alarm presentation helps.
- Duplicate: same condition from IED, PLC and gateway; establish one operational alarm with supporting diagnostics.
- Consequential: predictable cascade after a trip; present root trip clearly while preserving detailed events.
14. Local annunciation and remote HMI
- Define which critical alarms need local lamp/LED, audible annunciator and/or remote control-centre presentation.
- Use lamp test without changing process logic.
- Keep colour/flash/horn patterns limited and governed by the operator standard.
- Provide local indication when communication loss prevents remote alarms.
- Ensure alarm supply remains available during the failure being reported.
- Preserve panel labels/LED mappings with IED configuration revisions.
- For unmanned substations, define escalation, notification and loss-of-channel monitoring.
15. FAT/SAT test matrix
- Stimulate each physical input and internal diagnostic at the source.
- Verify active/clear text, equipment ID, priority, colour/audible and help/action.
- Check normal, failed supply, open wire, invalid quality and communication-loss states.
- Measure pickup/dropout delay, debounce and timestamp from a common time reference.
- Test first-out plus preservation of subsequent events.
- Verify latch, acknowledge, return to normal and reset as independent states.
- Test breaker 52a/52b travel and persistent mismatch.
- Test DC, TCS, VT failure, IED watchdog, LAN path and time-source alarms.
- Disconnect HMI/gateway/network and confirm buffering/restoration/no duplicates.
- Apply test/simulation flags and prove they cannot be mistaken for live alarms.
- Run a realistic trip sequence and reconstruct it from SOE/oscillography.
- Remove shelves/forces/blocks and record final database/configuration revision.
16. Performance review
- most frequent and chattering alarms;
- standing alarms by duration;
- alarm floods per trip/disturbance;
- priority distribution and operator response evidence;
- shelved/suppressed alarms and expiry;
- time-quality/communication gaps;
- misleading text or duplicate sources found in incident reviews;
- maintenance completion for protection-availability alarms.
Common mistakes
- Turning every status transition into an alarm.
- Assigning highest priority to every protection-related point.
- Using “Relay alarm” without device/channel/cause.
- Treating acknowledge as reset or condition clearance.
- Timestamping at a slow gateway and claiming millisecond SOE.
- Filtering inputs so heavily that event order becomes false.
- Showing invalid communication data as a process alarm.
- Suppressing a flood without preserving root and detailed events.
- Leaving test/shelved alarms active after commissioning.
- Testing HMI bits without stimulating the physical source.
Official standards and primary references
- IEC 62682:2022 — current alarm-management lifecycle principles for process industries, adaptable where specified for utility operations.
- IEC 60255-1:2022 — current common requirements for protection equipment, including monitoring/indication interfaces.
- IEC 61850-7-4 consolidated edition — standard logical-node/data-object semantics for utility automation information.
- IEC 61850-8-1 consolidated edition — MMS/GOOSE mapping, reports and event communication.
- IEC 61850-5:2013+AMD1:2022 — current communication performance and function/device model requirements.
- IEC 61850-6:2009+AMD1:2018+AMD2:2024 — current SCL configuration description requirements.
- IEC 62271-1:2017+AMD1:2021 — current common specifications for MV/HV switchgear indication and auxiliary/control equipment.
Engineering note: The goal is not more alarms; it is reliable operator understanding. Preserve detailed events for engineers while presenting only actionable abnormal conditions as alarms.