Alert description
This alert indicates that an O&M operation (such as adding a node, restarting, or replacing) has been performed on a node in the LogService cluster.
Alert principle
The following table lists the key parameters involved in the monitoring logic of this alert.
Parameter |
Value |
|---|---|
| Monitoring Metrics | None (for event alerts) |
| Metric description | Reported proactively by OCP upon completion of maintenance operations on the LogService node |
| Monitoring Expression | Event Trigger (expression_type: Event) |
| Metric Collection | operation_name、operation_result、operator、task_result、host |
| Metric Source | OCP Maintenance Events |
| Collection Cycle | Real-time Event Trigger |
Rule information
Monitoring Metrics |
Default Threshold |
Duration |
Detection Cycle |
Elimination Cycle |
|---|---|---|---|---|
| - | - | 0 Seconds | 0 Seconds | 5 Minutes |
Default status: is_enabled: 0 (disabled by default, can be enabled as needed)
Alert information
Alert Trigger Method |
Alert Level |
Scope |
|---|---|---|
| Based on O&M Events | Reminder | Host (LogService) |
Alert template
Alert overview
- Template: ${alarm_target} ${operation_result}
- Example: logservice_cluster=my_ls_cluster:host=11.124.9.47 succeeded
Alert details
- Template: LogService cluster: ${logservice_cluster}, node: ${host}, alert: ${operation_name} ${operation_result}, operator: ${operator}, result: ${task_result}.
- Example: LogService cluster: my_ls_cluster, node: 11.124.9.47, alert: Restart node succeeded, operator: admin, result: Successful.
Alert recovery
- Auto-resolve (auto_resolved: 1), with a resolution cycle of 5 minutes.
Impact on the system
Node-level O&M may cause the node to be temporarily unavailable; the node may be in an abnormal state if the operation fails.
Possible causes
- The administrator performed an O&M operation on a node of the LogService cluster in the OCP console.
Solution
- Confirm whether it is an on-plan node maintenance
- Check the node process status and monitoring metrics after the operation is completed
- If the operation fails, troubleshoot and resolve the issue based on the task details.
