實(shí)時(shí)作業(yè)要保證7 x 24運(yùn)行校焦,除了要在業(yè)務(wù)邏輯和編碼上下功夫之外,好的監(jiān)控系統(tǒng)也是必不可少的统倒。Flink支持多種匯報(bào)監(jiān)控指標(biāo)(metrics)的reporter斟湃,如JMX、SLF4J檐薯、InfluxDB、Prometheus等注暗。
Prometheus+Grafana是目前相當(dāng)流行的監(jiān)控+可視化一體方案坛缕,容易上手。下圖示出Prometheus及其周邊組件組成的生態(tài)系統(tǒng)捆昏。
本文簡(jiǎn)述整個(gè)配置流程宠页。
Prometheus+PushGateway安裝配置
為什么還要使用PushGateway呢遍烦?從上面的圖可以發(fā)現(xiàn)供填,Prometheus在正常情況下是采用拉模式從產(chǎn)生metric的作業(yè)或者exporter(比如專(zhuān)門(mén)監(jiān)控主機(jī)的NodeExporter)拉取監(jiān)控?cái)?shù)據(jù)粘捎。但是我們要監(jiān)控的是Flink on YARN作業(yè),想要讓Prometheus自動(dòng)發(fā)現(xiàn)作業(yè)的提交蓬痒、結(jié)束以及自動(dòng)拉取數(shù)據(jù)顯然是比較困難的亲轨。PushGateway就是一個(gè)中轉(zhuǎn)組件讯嫂,通過(guò)配置Flink on YARN作業(yè)將metric推到PushGateway,Prometheus再?gòu)腜ushGateway拉取就可以了。
首先將與Prometheus監(jiān)控相關(guān)的JAR包拷貝到Flink的lib目錄下。
cd /opt/flink-1.9.0
cp opt/flink-metrics-prometheus-1.9.0.jar lib
來(lái)到Prometheus官網(wǎng)的下載頁(yè)面阅嘶,分別下載本體以及PushGateway組件魂迄。也可以直接wget到服務(wù)器上熊昌。
wget -c https://github.com/prometheus/prometheus/releases/download/v2.12.0/prometheus-2.12.0.linux-amd64.tar.gz
wget -c https://github.com/prometheus/pushgateway/releases/download/v0.9.1/pushgateway-0.9.1.linux-amd64.tar.gz
解壓之后推溃,編輯Prometheus的配置文件prometheus.yml硬萍。主要是修改監(jiān)控間隔买羞,以及添加PushGateway的監(jiān)控配置漠嵌。Prometheus的默認(rèn)端口是9090几晤,PushGateway的是9091憾朴。
global:
scrape_interval: 60s # Set the scrape interval to every 15 seconds. Default is every 1 minute.
evaluation_interval: 60s # Evaluate rules every 15 seconds. The default is every 1 minute.
# scrape_timeout is set to the global default (10s).
# Alertmanager configuration
alerting:
alertmanagers:
- static_configs:
- targets:
# - alertmanager:9093
# A scrape configuration containing exactly one endpoint to scrape:
# Here it's Prometheus itself.
scrape_configs:
# The job name is added as a label `job=<job_name>` to any timeseries scraped from this config.
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']
labels:
instance: prometheus
- job_name: 'pushgateway'
static_configs:
- targets: ['my-flink002:9091']
labels:
instance: pushgateway
然后用nohup的方式啟動(dòng)PushGateway和Prometheus鸡岗。
nohup ./pushgateway --web.listen-address :9091 > /var/log/pushgateway.log 2>&1 &
nohup ./prometheus --config.file=prometheus.yml > /var/log/prometheus.log 2>&1 &
瀏覽器訪問(wèn)Prometheus所在節(jié)點(diǎn)的9090端口炮姨,查看Targets一項(xiàng)俄认。如果兩個(gè)項(xiàng)目都為UP狀態(tài)岂贩,就是配置好了抹镊。
還得記得編輯flink-conf.yaml俊嗽,在其中加上Flink與PushGateway集成的參數(shù)。
metrics.reporter.promgateway.class: org.apache.flink.metrics.prometheus.PrometheusPushGatewayReporter
# 這里寫(xiě)PushGateway的主機(jī)名與端口號(hào)
metrics.reporter.promgateway.host: my-flink002
metrics.reporter.promgateway.port: 9091
# Flink metric在前端展示的標(biāo)簽(前綴)與隨機(jī)后綴
metrics.reporter.promgateway.jobName: flink-metrics-ppg
metrics.reporter.promgateway.randomJobNameSuffix: true
metrics.reporter.promgateway.deleteOnShutdown: false
Grafana安裝配置
從Grafana官網(wǎng)直接下載二進(jìn)制包非常慢,所以要借助清華的鏡像源箭养。在/etc/yum.repos.d目錄下新建grafana.repo,寫(xiě)入如下內(nèi)容潘酗。
[grafana]
name=grafana
baseurl=https://mirrors.tuna.tsinghua.edu.cn/grafana/yum/rpm
repo_gpgcheck=0
enabled=1
gpgcheck=0
然后用yum安裝祭衩。
yum makecache
yum -y install grafana
Grafana的配置文件位于/etc/grafana/grafana.ini精算,但是我們可以不用改,直接啟動(dòng)服務(wù)蹂喻。
service grafana-server start
Grafana默認(rèn)跑在3000端口赤嚼,用瀏覽器訪問(wèn)之,然后點(diǎn)擊Create your first data source添加Prometheus數(shù)據(jù)源果录。具體的填寫(xiě)方法如下圖所示。
運(yùn)行示例程序來(lái)產(chǎn)生一些監(jiān)控?cái)?shù)據(jù)。
nc -l 17777
/opt/flink-1.9.0/bin/flink run \
--detached \
--jobmanager yarn-cluster \
--yarnname "example-socket-window-wordcount" \
/opt/flink-1.9.0/examples/streaming/SocketWindowWordCount.jar \
--port 17777
最后添加一個(gè)用于查看監(jiān)控圖表的dashboard抡诞。點(diǎn)擊New Dashboard->Add Query按鈕源哩,就可以直接看到Flink下的各個(gè)metric赊锚,還可以添加多個(gè)metric资昧。
選完之后格带,在上面的圖表中就可以看到監(jiān)控?cái)?shù)據(jù)的折線了撤缴。