Applicability
This topic applies only to OceanBase Database Enterprise Edition. OceanBase Database Community Edition does not support the arbitration service feature.
Overview
The Arbitration Service supports collecting monitoring metrics via the Prometheus protocol. It has a built-in standard GET /metrics interface, which provides metrics such as the CPU, memory, and disk usage of the arbserver, as well as information about the currently managed clusters, tenants, LSs, elections, configuration versions, schema versions, and RPC latency.
The Arbitration Service exposes monitoring metrics through its built-in HTTP Server. You can use Prometheus Server to capture and store this metric data, and integrate with Grafana for metric visualization and Alertmanager for alerts. This feature does not involve the installation and configuration of Prometheus, Grafana, or Alertmanager; you must deploy and configure these external components yourself.
By using the Prometheus monitoring metrics from the Arbitration Service, you can gain real-time insights into the resource usage of arbservers (CPU, memory, disk usage), understand the number of clusters, tenants, and LSs currently being managed, monitor the leader status and election frequency of each LS, track changes in configuration and schema versions, and analyze the distribution of RPC latency across network transmission and server-side processing phases.
Prerequisites
- The Arbitration Service has been deployed in the current cluster. For detailed deployment instructions, see Deploy an OceanBase cluster with two replicas and the Arbitration Service.
- Prometheus Server has been deployed in the external environment, and network connectivity between Prometheus Server and the arbserver is ensured.
- (Optional) If visualization or alerts are required, Grafana and Alertmanager must also be deployed.
Configure Prometheus collection
Configure the metric collection port
The Arbitration Service controls whether to enable the metric collection port via the parameter prometheus_metrics_port. There are two ways to enable this service:
Method 1:
Add the port parameter to the startup configuration in the Arbitration Service to enable the metric collection interface. Example:
-o "ob_startup_mode=arbitration,prometheus_metrics_port=2886"
For the complete startup command, see Deploy an OceanBase cluster with two replicas and the Arbitration Service by using the command line.
When using the default configuration, you do not need to explicitly specify the port, as the arbserver attempts to listen on 2886. To disable the metric collection interface, set this parameter to 0. Example:
-o "ob_startup_mode=arbitration,prometheus_metrics_port=0"
Method 2:
Modify the port during cluster runtime using alter system set, and the change takes effect after restarting the arbserver process.
Modify the port using
alter system set.alter system set prometheus_metrics_port = 2887;The change takes effect after restarting the Arbitration Service process.
The Arbitration Service is a lightweight observer process started in a special mode. To restart a node process, see Restart a node.
If you want to disable the metric collection interface at this point, changing the prometheus_metrics_port number to 0 still requires a restart for the change to take effect.
Behavior description
The behavior of the prometheus_metrics_port parameter varies depending on the configuration scenario.
- No port explicitly configured: Exposes
/metricsusing the default port2886. - Configured as another positive number: The arbserver changes to listen on the specified port after restart.
- Configured as
0: The metric HTTP Server is not started, and/metricsis inaccessible. - Port modified but not restarted: The original port is used until the arbserver restarts.
- Port is occupied or listening fails: The metric service fails to initialize and shuts down automatically. All metrics stop updating, but this does not affect the normal startup of arbserver. The system logs a WARN-level listening error (
disable Prometheus monitoring). After the port conflict is resolved, you must restart arbserver to resume metric collection. - One arbserver serves multiple clusters: The endpoint only needs to be captured once; LS metrics distinguish between clusters, tenants, and LS by labels.
- arbserver restart: Manages objects and versioned metrics are reloaded from the persistent state.
Verify metric interface availability
After configuration, you can use the curl command to verify whether the metric interface is accessible.
Execute the following command from a Prometheus Server or a machine within the same collection network:
curl -fsS http://<arbserver_ip>:2886/metrics
A valid response includes # HELP, # TYPE, and the arb_* metric data. A sample response is as follows:
# HELP arb_cluster_count number of clusters managed
# TYPE arb_cluster_count gauge
arb_cluster_count 1
Configure scraping
Add a scrape job in the Prometheus configuration file, prometheus.yml, to capture arbserver metrics. The configuration example is as follows:
scrape_configs:
- job_name: oceanbase-arbserver
metrics_path: /metrics
scrape_interval: 15s
static_configs:
- targets:
- <arbserver_ip>:2886
labels:
service: arbserver
The parameters are described as follows:
Parameter |
Required |
Description |
|---|---|---|
job_name |
Yes | The name of the scrape job, which can be customized. It is recommended to use oceanbase-arbserver for quick identification in Prometheus. |
metrics_path |
No | The metric exposure path, which defaults to /metrics, is consistent with the built-in interface of arbserver and usually does not require modification. |
scrape_interval |
No | The scraping interval, which is the time interval between each metric data collection by Prometheus. It is recommended to set it to 15s. A value that is too short increases the load on the arbserver, while a value that is too long affects the real-time performance of monitoring. |
targets |
Yes | The access address of the arbserver, in the format of <arbserver_ip>:<prometheus_metrics_port>, where <prometheus_metrics_port> is the value of the arbserver configuration item prometheus_metrics_port, which defaults to 2886. If there are multiple arbserver instances, add the address of each instance to this list. |
labels |
No | Append custom labels to all metrics for this scrape job. In the example, the service: arbserver label is added, which facilitates filtering by service name when querying in Grafana or using PromQL. |
If there are multiple arbserver instances, add each instance's address to the targets list, or integrate with an existing service discovery mechanism of your deployment platform. For detailed format information on Prometheus scraping configuration, see Prometheus Configuration.
Verify the scraping status
After configuring scraping, follow these steps to verify whether Prometheus has successfully scraped arbserver metrics.
On the Prometheus Server, use the
prometheususer to check whether the configuration file syntax is correct.promtool check config <prometheus.yml>If the configuration file syntax is correct,
SUCCESSis displayed in the output.On the Prometheus Server, reload or restart the Prometheus service for the configuration to take effect.
If Prometheus has hot reloading enabled, any user with network access can trigger a configuration reload via the HTTP API:
curl -X POST http://<prometheus_ip>:9090/-/reloadYou can also directly restart the Prometheus service as the
rootuser or withsudoprivileges:sudo systemctl restart prometheus
View the log
sudo journalctl -u prometheus -fto confirm thatCompleted loading of configuration fileis displayed and no errors are reported.On any machine where the Prometheus web interface is accessible, confirm that the target status of arbserver is
UP.Note
upis a metric automatically generated by Prometheus for each scrape target. It is not anarb_*metric exposed by arbserver itself. A status of UP indicates that the Prometheus server successfully initiated an HTTP request to the arbserver's exposer and received a well-formed response.Use a browser to access
http://<prometheus_ip>:9090/targets, findoceanbase-arbserverin the target list, and confirm its status isUP.You can also query the target status via command on any machine where the Prometheus API is accessible. Example:
curl -s http://<prometheus_ip>:9090/api/v1/targets | grep -o '"health":"[^"]*"'
On any machine where the Prometheus API is accessible, query the
upmetric to confirm that arbserver has been successfully scraped. Example:curl -s "http://<prometheus_ip>:9090/api/v1/query?query=up%7Bjob%3D%22oceanbase-arbserver%22%7D"If the
valuein the returned result is1, the scrape is normal; if it is0, the scrape failed. You need to check the network connectivity and port configuration between Prometheus and arbserver.On any machine where the Prometheus API is accessible, query business metrics such as
arb_cluster_countandarb_ls_has_leaderto confirm that arbserver metrics are being collected normally. Example:curl -s "http://<prometheus_ip>:9090/api/v1/query?query=arb_cluster_count"You can also list all metrics with the
arb_prefix to confirm that metrics are being collected completely. Example:curl -s "http://<prometheus_ip>:9090/api/v1/label/__name__/values" | grep '^arb_'
Metric description
arbserver currently provides 13 logical metrics, divided into three categories: resources and management scale, LS status, and RPC latency.
Resources and management scale
Metrics |
Type |
Meaning |
Update method |
|---|---|---|---|
arb_memory_usage_percent |
Gauge | Percentage of physical memory occupied by the arbserver process's VmRSS | Every 10 seconds |
arb_cpu_usage_percent |
Gauge | CPU usage of the arbserver process | Every 10 seconds, only the baseline is established for the first sample. |
arb_disk_usage_percent |
Gauge | File system usage of the arbserver data directory | Every 10 seconds |
arb_cluster_count |
Gauge | Number of clusters managed | Lifecycle event updates |
arb_tenant_count |
Gauge | Number of tenants managed | Lifecycle event updates |
arb_ls_count |
Gauge | Total managed LSs | Lifecycle event updates |
LS status
The following metrics all have the cluster_id, tenant_id, and ls_id labels to distinguish different clusters, tenants, and log streams.
Metrics |
Type |
Meaning |
Update method |
|---|---|---|---|
arb_ls_has_leader |
Gauge | Whether the LS has a leader from the perspective of arbitration; 1 indicates yes, 0 indicates no. |
Every 2 seconds |
arb_ls_election_count |
Counter | Number of elections observed by the arbitration side | Election event update |
arb_ls_config_version_proposal_id |
Gauge | The proposal id of LogConfigVersion |
Config event update |
arb_ls_config_version_config_seq |
Gauge | The config sequence of LogConfigVersion |
Config event update |
arb_ls_mode_version |
Gauge | The mode version of LogModeMeta |
Mode event update |
RPC latency
Metrics |
Type |
Label |
Meaning |
|---|---|---|---|
arb_rpc_process_latency_seconds |
Histogram | None | Queuing and processing delay of RPC on the arbserver side |
arb_rpc_transport_latency_seconds |
Histogram | None | Transmission delay from the client to the arbserver |
Histogram-type metrics are expanded into _bucket, _sum, and _count under /metrics. The default buckets are as follows:
100us, 500us, 1ms, 5ms, 10ms, 50ms,
100ms, 500ms, 1s, 5s, 10s, +Inf
Note
The two RPC Histogram metrics currently do not have pcode, cluster, or tenant labels, indicating the global aggregation result of the arbserver. For information about how to query Histogram metrics, see Prometheus Histograms and Summaries.
Common query examples
You can directly execute the following query statements in the query box of the Prometheus web interface, or use them to query through the Prometheus HTTP API.
Query the arbserver scrape status:
up{job="oceanbase-arbserver"}Query the current leaderless LS:
arb_ls_has_leader == 0Query the number of elections within 10 minutes:
increase(arb_ls_election_count[10m])Query the server-side RPC P99 latency:
histogram_quantile( 0.99, sum by (le) (rate(arb_rpc_process_latency_seconds_bucket[5m])) )Query the network transport RPC P99 latency:
histogram_quantile( 0.99, sum by (le) (rate(arb_rpc_transport_latency_seconds_bucket[5m])) )
Limitations
- Currently, you cannot switch the metric collection port online via SQL. The
prometheus_metrics_portparameter is statically effective. After modification, you must restart the arbserver for the change to take effect. - Currently, HTTPS and API authentication are not supported. The
/metricsinterface uses plaintext HTTP, without authentication or TLS encryption. - Currently, listening on an IPv6 address is not supported. It listens on the fixed IPv4 address
0.0.0.0. If observer is deployed using an IPv6 address, Prometheus will not be able to access the/metricsinterface via the IPv6 address, resulting in no data for arbitration service monitoring. - Currently, there is no automatic discovery mechanism for arbserver. You need to manually configure static targets in Prometheus, or integrate with the existing service discovery mechanism of your deployment platform.
- Currently, pre-configured Grafana Dashboards and Alertmanager alert rules are not provided. You need to create and configure them yourself in Grafana and Alertmanager.
- Currently, only three resource metrics are provided: CPU usage, memory usage, and disk usage. This does not include comprehensive host network and file system monitoring. To monitor metrics such as packet loss, retransmissions, and read-only file systems, it is recommended to use node_exporter.
- Currently, arbserver metrics do not include observer log replication lag or FPAXOS status information, such as follower lag, Q1/Q2, or Reconfirm waiting members.
- Currently, the two RPC Histogram metrics only provide global aggregated values and do not support latency analysis by RPC type, cluster, or tenant.
Security and capacity considerations
The
/metricsinterface listens on all IPv4 network interfaces and has no authentication or TLS encryption. It is recommended to use security groups, firewalls, container networks, or reverse proxies to allow only Prometheus to access this interface over the network, avoiding direct exposure to untrusted networks.When running multiple arbserver instances on the same machine, each instance must use a different port. After modifying a port, update the Prometheus target configuration accordingly.
Resource metrics are sampled every 10 seconds, and leader status refreshes every 2 seconds. The actual refresh delay for monitoring data is the sum of the arbserver sampling interval and the Prometheus
scrape_interval. It is recommended to set theforparameter in alert rules based on the duration of the status to avoid triggering alerts immediately due to fluctuations from a single sample.Each LS generates five sets of metrics tagged with
cluster_id,tenant_id, andls_id. In large-scale LS scenarios, evaluate the scrape response size, the number of Prometheus series, storage costs, and query overhead. If necessary, appropriately increase thescrape_intervalto reduce the collection frequency.
