An alarm requires timely operator action; a status or event only records information. Treating every relay bit as an alarm overwhelms operators during the exact bus trip, DC failure or communication disturbance when prioritization matters most. Substation alarm management therefore needs a lifecycle: philosophy, identification, rationalization, design, implementation, operation, monitoring and change.
IEC 62682 is written for process industries, but its alarm-management principles provide a strong framework when adapted to power-system operating responsibilities and protection/SOE requirements.
1. Alarm, event, status and diagnostic
| Class | Purpose | Example |
|---|---|---|
| Alarm | Notify abnormal condition requiring operator response | Trip circuit failed; breaker failed to open |
| Event/SOE | Record state transition for sequence/fault analysis | Protection start and trip output |
| Status | Show current equipment/system state | Breaker open; spring charged |
| Maintenance diagnostic | Support technical troubleshooting | IED internal warning or port counter |
| Security event | Record/alert cyber-relevant activity | Failed login or configuration change |
The same source may create an SOE record and an alarm, but these are separate functions. A protection trip is an event; the resulting feeder outage or breaker failure may require a specific operator alarm.
2. Alarm philosophy
- Define what qualifies as an alarm and who must respond.
- Set priority methodology from consequence and maximum allowable response time.
- Define text, colors/sounds, acknowledgment, shelving, suppression and latching.
- State handling of bad quality, communication loss, test/simulation and maintenance.
- Define flood/standing/chattering alarm criteria and review metrics.
- Separate station operator, control-center dispatcher, protection engineer and cyber/SOC responsibilities.
- Control additions/changes through rationalization and testing.
3. Rationalize every alarm
| Rationalization field | Required decision |
|---|---|
| Cause | What physical/system condition creates it? |
| Consequence | What happens without response? |
| Operator action | Specific action, owner and escalation |
| Time | Maximum useful response time |
| Priority | Derived consistently from consequence/time |
| Setpoint/delay | Value, debounce/on/off delay and justification |
| Suppression | Parent/state/maintenance conditions and safety |
| Evidence | FAT/SAT stimulus, expected text/priority/action |
4. Priority design
- Use few, clearly distinguished priority levels.
- Assign from operator consequence and time—not equipment price, vendor default or signal source.
- Reserve highest priority for rare conditions needing immediate action.
- Protection operation may be high consequence but already automatic; prioritize the operator’s remaining required action.
- Communication alarms should identify the lost operational function, not create hundreds of identical point alarms.
- Validate priority distribution; if everything is high, nothing is prioritized.
- Document separate notification/escalation paths outside the HMI only where justified.
5. Alarm text and HMI context
- Use station/voltage/bay/equipment/condition in concise, consistent text.
- Distinguish trip, breaker failure, lockout, IED failure, source communication loss and bad data quality.
- Provide cause/action help without burying the primary message.
- Show current value/state, quality, source time and related topology.
- Acknowledgment means the operator saw the alarm; it does not clear the physical condition.
- Use return-to-normal messages and latching consistently.
- Test translations and truncation on every operator display.
6. Flood prevention by design
- Correct chattering contacts, unstable thresholds and communication retry storms at source.
- Use justified debounce/on/off delay without hiding fast protection/SOE events.
- Group dependent symptoms under a root-cause alarm while keeping forensic events.
- For IED/gateway loss, suppress or mark child points bad rather than alarming every point separately.
- Apply deadband/rate handling to non-actionable analog activity.
- Stagger integrity scans/reconnect to avoid recovery floods.
- Size servers/gateways/HMI queues for credible bus/DC/common-network event bursts.
7. State-based suppression
- Suppress only when the alarm is not meaningful in a known operating state.
- Examples: feeder measurement low when breaker open; mechanism unavailable while breaker withdrawn for maintenance.
- Do not suppress protective trips, lockout or safety-critical conditions merely because equipment is out of service.
- Use fresh, valid state; invalid topology should not silently suppress alarms.
- Record suppression reason and keep an audit/event trail.
- Test entry/exit timing so alarms do not flash or remain hidden after state change.
8. Shelving and out-of-service handling
- Shelving is an authorized, temporary operator action—not permanent deletion.
- Require reason, owner and expiry; show shelved alarms prominently.
- Limit which priorities/types may be shelved.
- Alarm/notify on shelf expiry or excessive duration.
- Out-of-service should follow maintenance/work-permit status and restore checks.
- Retain event history and assess whether an independent notification is still required.
9. Standing, stale and bad-quality alarms
- Review long-standing alarms daily/periodically by responsibility and risk.
- Bad quality is not automatically the same alarm as the process value; identify source failure.
- Separate IED offline, gateway mapping failure, WAN loss and HMI client loss.
- After communication restoration, revalidate state before clearing quality alarms.
- Prevent repeated alarm/return cycles during unstable links through appropriate state logic.
- Do not discard old source-time events that arrive during reconnect; mark/present them correctly.
10. Performance monitoring
- Alarm rate per operator/time and peak flood rate.
- Top frequent/chattering alarms and repeated returns.
- Standing alarms and duration.
- Shelved/suppressed/out-of-service population and expiry.
- Priority distribution and acknowledgment/response time.
- Flood episodes linked to initiating cause and lost/missed operator actions.
- Changes/additions and post-change performance.
- Use metrics to improve sources/rationalization, not punish operators.
11. Cybersecurity and alarm integrity
- Protect alarm configuration, priorities, suppression logic and audit records.
- Control who may acknowledge, shelve, inhibit or alter alarms.
- Monitor unexpected alarm disabling, configuration change and time manipulation.
- Ensure cyber alarms reach an accountable recipient without flooding substation operators.
- Preserve alarms locally during WAN/SOC loss and reconcile later.
- Test denial-of-service/event-flood resilience and logging capacity.
12. FAT/SAT and lifecycle
- Freeze the rationalized master alarm database and HMI build.
- Stimulate every alarm: active, acknowledge, return, latch/reset and bad quality.
- Verify text, priority, sound/color, source time, help and operator action.
- Test delay/debounce, state suppression, shelving, expiry and out-of-service.
- Generate IED/network/DC/bus event floods and measure system/operator usability.
- Fail gateway/WAN/server and verify buffering/recovery without duplicate flood.
- Test roles/audit and unauthorized alarm-configuration/shelving attempts.
- At SAT, prove physical inputs and operating procedures with representative staff.
- Monitor post-commissioning metrics and rationalize nuisance/standing alarms.
- Regression-test every alarm/logic/database change.
13. Practical substation alarm grouping
- Immediate system response: breaker failure, busbar/arc operation, unexpected energization or loss of critical bus supply.
- Protection/control impaired: trip circuit, DC, IED self-supervision, protection channel or time dependency that removes required coverage.
- Equipment attention: mechanism/drive, pressure/density, temperature or condition threshold with a defined maintenance response time.
- Communications: aggregate the lost operational function at IED, station or WAN level; retain child-point quality without creating hundreds of alarms.
- Security/configuration: unauthorized access/change, disabled protection/reporting or unexpected test/bypass state routed to the responsible operator/SOC.
- Advisory/status: keep routine position, settings group and normal mode visible but not alarmed unless an abnormal condition requires action.
Priority within each group still follows consequence and response time. A low-voltage panel-heater failure and a duplicated DC supply failure are not equivalent merely because both are “auxiliary supply” signals.
References and further reading
- IEC 62682:2022 — Alarm-system management principles and lifecycle
- IEC 61850-8-1 consolidated edition — Reporting and event communication
- IEC 61850-7-3 consolidated edition — Data/quality semantics
- IEC 60870-5-104 consolidated edition — Telecontrol alarms/events
- IEC 62443-3-3:2013 — IACS system security requirements
- IEC 62351-5:2023 — Secure IEC 60870-5 derivatives
Engineering note: Keep complete SOE evidence, but present only conditions requiring action as alarms; forensic completeness and operator usability are complementary, not competing goals.