Prometheus 서버 모니터링 구축하기 (2) - 설치부터 SMS 알림까지

2026. 9. 30. 00:59ㆍ기타

모든 모니터링 도구는 같은 EC2 인스턴스 안에 설치하고, 서로 로컬 통신만 하게 하였습니다.

 

공통 설치 패턴

Exporter, Prometheus, Alertmanager는 모두 아래와 같은 패턴으로 설치했습니다.

  1. /opt에 다운로드하고 압축 해제
  2. 실행 전용 계정 생성
  3. 실행 파일만 /usr/local/bin으로 복사, 소유권 변경
  4. 설정 파일은 /etc/도구이름/, 데이터는 /var/lib/도구이름/ 에 저장
  5. systemd 서비스로 등록

 

파일 위치 정리

위치 설명
/opt 설치 및 압축을 푸는 작업 공간으로 사용
/usr/local/bin 직접 설치한 프로그램의 표준 위치로 경로 없이 실행 가능
/etc/... 리눅스 설정 파일의 표준 위치
/var/lib/... 실행 중 생기는 데이터의 표준 위치

 

 

전용 계정을 만드는 이유

패턴을 보면 각 프로그램 별로 실행 전용 계정을 생성합니다.

이는 최소 권한 원칙때문인데, Node Exporter는 CPU/메모리 등의 수치를 읽기만 하는 프로그램입니다.

이걸 root로 실행하는데 취약점이 발견된다면 공격자는 root 권한을 그대로 얻을 수 있습니다.

따라서, 로그인도 안되는 전용 계정으로 실행하면 피해가 최소화될 수 있습니다.

 

 

 

 

systemd

원래라면 서버에 들어가서 프로그램을 실행시켜야 합니다.

그런데 이 방식이라면 아래와 같은 문제가 있습니다.

  • SSH 연결을 끊는 순간 프로그램도 같이 종료
  • 서버 재부팅시에 다시 들어가서 수동으로 실행해야 함

이러한 이유로 24시간 떠있어야 하는 모니터링 도구는 systemd 서비스로 등록해야 합니다.

systemd는 이런 백그라운드 프로그램들을 리눅스가 대신 관리해 주는 시스템입니다.

"이 프로그램은 이렇게 실행해"라는 명세서를 등록해 놓으면 그다음부터는 systemd가 알아서 명세서에 따라서 관리를 합니다.

 

.service 파일

방금 말했던 명세서가 .service 파일입니다.

/etc/systemd/system/ 밑에 .service로 텍스트 파일을 만들어두면 systemd가 그 내용을 읽어서 프로그램을 관리합니다.

 

.servcie 파일 예시

cat > /etc/systemd/system/node_exporter.service << 'EOF'
[Unit]
Description=Node Exporter
After=network.target

[Service]
User=node_exporter
Group=node_exporter
Type=simple
ExecStart=/usr/local/bin/node_exporter --web.listen-address=127.0.0.1:9100

[Install]
WantedBy=multi-user.target
EOF

 

여기서 cat > 파일 << 'EOF' ... EOF는 여러 줄짜리 텍스트를 한 번에 파일로 저장하는 셸 문법입니다.

EOF(End Of File)로 시작해서 다시 EOF가 나올 때까지의 내용을 그대로 파일에 씁니다.

즉 위 명령은 "node_exporter.service라는 파일을 만들고, 그 안에 아래 내용을 그대로 넣어라"는 뜻이 됩니다.

 

.service 파일 내용을 뜯어보면

  • [Unit]: 이 서비스가 뭔지, 언제 시작할지에 대한 정보
    After=network.target은 "네트워크가 뜬 다음에 시작해라"는 뜻
    (예시의 Node Exporter는 네트워크로 요청을 받아야 하니 네트워크보다 먼저 뜨면 의미가 없기 때문)
  • [Service]: 실제로 어떻게 실행할지.
    • User, Group: 추후에 만들 전용 계정으로 실행
    • Type=simple: 가장 기본적인 실행 방식 (프로그램이 뜨면 바로 "실행 중"으로 간주)
    • ExecStart: 실제 실행할 명령어. --web.listen-address=127.0.0.1:9100으로 9100 포트에서, 로컬(127.0.0.1)에서만 접근 가능하게 제한
  • [Install]: WantedBy=multi-user.target은 "서버가 정상적으로 부팅을 마친 시점에 이 서비스를 켜라"라고 선언해 두는 부분
    이 자체가 자동 실행을 켜는 건 아니고, 다른 명령어인 systemctl enable이 실제로 자동 시작을 등록한다.

 

 

 

Node Exporter 설치 - OS 리소스 수집

Step 1. 다운로드 및 압축 해제

https://github.com/prometheus/node_exporter/releases

git repo에서 최신 버전을 확인한 다음 wget, tar 명령어를 통해서 설치 및 압축해제 해줍니다.

cd /opt
wget https://github.com/prometheus/node_exporter/releases/download/v1.12.1/node_exporter-1.12.1.linux-amd64.tar.gz
tar xvf node_exporter-1.12.0.linux-amd64.tar.gz

 

 

Step 2. 전용 계정 생성 + 실행 파일 배치

useradd --no-create-home --shell /sbin/nologin node_exporter
cp /opt/node_exporter-1.12.1.linux-amd64/node_exporter /usr/local/bin/
chown node_exporter:node_exporter /usr/local/bin/node_exporter
  • --no-create-home: 실행용이므로 홈 디렉토리 생성X
  • --shell /sbin/nologin: 이 계정으로는 로그인 불가능

Step 3. systemd 서비스 등록

cat > /etc/systemd/system/node_exporter.service << 'EOF'
[Unit]
Description=Node Exporter
After=network.target

[Service]
User=node_exporter
Group=node_exporter
Type=simple
ExecStart=/usr/local/bin/node_exporter --web.listen-address=127.0.0.1:9100

[Install]
WantedBy=multi-user.target
EOF

systemctl daemon-reload      # 새 서비스 파일을 systemd가 인식하게
systemctl start node_exporter
systemctl enable node_exporter
  • systemctl daemon-reload: 방금 만든 서비스 파일을 systemd가 인식하도록 다시 읽어 들이는 명령어
  • systemctl start: 서비스 실행
  • systemctl enable: 재부팅 시 자동 시작되도록 등록

Step 4. 확인

systemctl status node_exporter                          # active (running) 확인
curl -s http://127.0.0.1:9100/metrics | head -n 20      # 메트릭이 나오는지

 

정상적으로 실행 중이라면 이렇게 표시됩니다.

 

 

 

 

Blackbox Exporter - 앱의 작동여부 확인

Step 1. 설치 (공통 패턴과 동일)

cd /opt
wget https://github.com/prometheus/blackbox_exporter/releases/download/v0.28.0/blackbox_exporter-0.28.0.linux-amd64.tar.gz
tar xvf blackbox_exporter-0.28.0.linux-amd64.tar.gz

useradd --no-create-home --shell /sbin/nologin blackbox_exporter
cp /opt/blackbox_exporter-0.28.0.linux-amd64/blackbox_exporter /usr/local/bin/
chown blackbox_exporter:blackbox_exporter /usr/local/bin/blackbox_exporter

mkdir /etc/blackbox_exporter
  • mkdir /etc/blackbox_exporter : Blackbox Exporter는 "어떤 방식으로 체크할지"를 정의한 설정 파일(blackbox.yml)을 읽어서 동작함 따라서, 이 파일을 저장할 전용 폴더를 /etc 아래에 만듦

 

Step 2. 설정 파일 생성

cat > /etc/blackbox_exporter/blackbox.yml << 'EOF'
modules:
  http_alive:
    prober: http
    timeout: 5s
    http:
      valid_status_codes: [200, 301, 302, 400, 401, 403, 404]
      method: GET
      preferred_ip_protocol: ip4
EOF
  • valid_status_codes : 이 목록의 응답이 오면 "살아있음". 연결 실패나 타임아웃이면 "다운".
  • timeout: 5s : 5초 안에 응답이 없으면 실패

기본 설정 파일은 각종 예시 모듈(grpc, ssh, postgresql...)이 들어있는 샘플입니다.

기본 http_2xx 모듈은 200번 대만 정상으로 보기 때문에 우선 응답이 있다면 앱이 살아있다고 판단하기 위해 위처럼 설정해 줬습니다.

 

Step 3. systemd 서비스 등록

 

cat > /etc/systemd/system/blackbox_exporter.service << 'EOF'
[Unit]
Description=Blackbox Exporter
After=network.target

[Service]
User=blackbox_exporter
Group=blackbox_exporter
Type=simple
ExecStart=/usr/local/bin/blackbox_exporter --config.file=/etc/blackbox_exporter/blackbox.yml --web.listen-address=127.0.0.1:9115

[Install]
WantedBy=multi-user.target
EOF

systemctl daemon-reload
systemctl start blackbox_exporter
systemctl enable blackbox_exporter

 

Step 4. 확인

curl -s "http://127.0.0.1:9115/probe?target=http://127.0.0.1:[서비스 port]&module=http_alive" | grep probe_success
# probe_success 1   ← 1이면 살아있음, 0이면 다운

 

"이 대상을 이 모듈로 체크해 줘."라고 Blackbox가 대상에게 요청을 보내고 결과를 알려줍니다.

나중에는 이 요청을 Prometheus가 자동으로 보냅니다.

 

 

 

 

Prometheus - 수집/저장/규칙 검사

Step 1. 설치

cd /opt
wget https://github.com/prometheus/prometheus/releases/download/v3.14.0/prometheus-3.14.0.linux-amd64.tar.gz
tar xvf prometheus-3.14.0.linux-amd64.tar.gz

useradd --no-create-home --shell /sbin/nologin prometheus
cp /opt/prometheus-3.14.0.linux-amd64/prometheus /usr/local/bin/
cp /opt/prometheus-3.14.0.linux-amd64/promtool /usr/local/bin/
chown prometheus:prometheus /usr/local/bin/prometheus /usr/local/bin/promtool

mkdir /etc/prometheus          # 설정
mkdir /var/lib/prometheus      # 수집한 데이터(시계열 DB)가 쌓이는 곳
chown prometheus:prometheus /var/lib/prometheus

 

Step 2. 알람 규칙 파일 생성

cat > /etc/prometheus/alert_rules.yml << 'EOF'
groups:
  - name: service_alerts
    rules:
      - alert: AppDown
        expr: probe_success == 0
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: "앱 다운 감지"
          description: "{{ $labels.instance }} 이(가) 1분 이상 응답하지 않습니다."
EOF

 

어떤 조건일 때 알림을 발동할지 정의합니다.

  • expr: probe_success가 0이면 앱 다운이므로 알림 발동
  • for: 조건이 1분 동안 지속되어야 발동, 순간적인 깜빡임으로 잘못된 알림이 가는 걸 방지

Step 3. 메인 설정 파일 생성

cat > /etc/prometheus/prometheus.yml << 'EOF'
global: # 전역설정
  scrape_interval: 1m # Exporter에서 얼마나 자주 긁어올지.
  evaluation_interval: 1m # 저장된 값을 알림 규칙과 얼마나 자주 비교할지.

rule_files: # 알림규칙 파일 연결
  - /etc/prometheus/alert_rules.yml #해당 경로의 파일을 규칙파일로 지정

alerting: # 규칙 위반시 알림을 어디로 보낼지 설정
  alertmanagers:
    - static_configs:
        - targets: ['127.0.0.1:9093'] # 추후에 Alertmanager를 해당 포트에 띄움

scrape_configs: # 긁어올 대상 목록
  - job_name: 'prometheus' # 수집 작업의 이름
    static_configs:
      - targets: ['127.0.0.1:9090'] # 긁어올 대상의 주소. 
      # Prometheus도 하나의 프로그램이니 스스로 잘 돌고있는지 확인하기 위해 자기 상태 모니터링

  - job_name: 'node' # Node Exporter
    static_configs:
      - targets: ['127.0.0.1:9100']

  - job_name: 'blackbox_http' # Blackbox Exporter
    metrics_path: /probe # 요청 경로
    params:
      module: [http_alive] # /probe로 요청시 module: http_alive 파라미터를 자동으로 붙여줌
    static_configs:
      - targets: ['http://127.0.0.1:5011']
        labels: {alias: '운영 서버'}
    relabel_configs: # 요청 변환 (실제 요청은 5011이 아니라 9115로 가야함)
      - source_labels: [__address__]
        target_label: __param_target
        # 원래 target에 적은 주소(__address__, 5011포트)을 __param_target이라는 값으로 복사한다.
        # 이렇게하면 요청시 target=http://127.0.0.1:5011 파라미터로 붙음
      - source_labels: [__param_target]
        target_label: instance
      - target_label: __address__
        replacement: 127.0.0.1:9115
        # 진짜 요청을 보낼주소(__address__)를 강제로 9115포트로 덮어씌움
EOF

# 실제 요청은 아래와 같이 된다.
# GET http://127.0.0.1:9115/probe?target=http://127.0.0.1:5011&module=http_alive

promtool check config /etc/prometheus/prometheus.yml # prometheus.yml이 문법에 맞는지 검사

 

Step 4. 서비스 등록

cat > /etc/systemd/system/prometheus.service << 'EOF'
[Unit]
Description=Prometheus
After=network.target

[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/prometheus \
  --config.file=/etc/prometheus/prometheus.yml \
  --storage.tsdb.path=/var/lib/prometheus/ \ # 수집한 데이터가 쌓이는 위치 지정
  --web.listen-address=127.0.0.1:9090

[Install]
WantedBy=multi-user.target
EOF

systemctl daemon-reload
systemctl start prometheus
systemctl enable prometheus

 

 

 

Alertmanager - 알림 전달

Step 1. 설치

cd /opt
wget https://github.com/prometheus/alertmanager/releases/download/v0.34.0/alertmanager-0.34.0.linux-amd64.tar.gz
tar xvf alertmanager-0.34.0.linux-amd64.tar.gz

useradd --no-create-home --shell /sbin/nologin alertmanager
cp /opt/alertmanager-0.34.0.linux-amd64/alertmanager /usr/local/bin/
cp /opt/alertmanager-0.34.0.linux-amd64/amtool /usr/local/bin/
chown alertmanager:alertmanager /usr/local/bin/alertmanager /usr/local/bin/amtool

mkdir /etc/alertmanager
mkdir /var/lib/alertmanager      # 알림 상태(누구에게 언제 보냈는지) 기록
chown alertmanager:alertmanager /var/lib/alertmanager

 

Step 2. 설정 파일 생성

cat > /etc/alertmanager/alertmanager.yml << 'EOF'
global:
  resolve_timeout: 5m # 문제 해결 감지 시간. 5분 동안 그 문제가 다시 감지되지 않으면 해결됬다고 판단

route: # 알림 처리 방법 정의
  receiver: 'webhook-server' # receivers에 정의된 것중 webhook-server에 보내라는 의미
  group_by: ['alertname'] # 알림을 묶는 기준. 여러개의 서버가 동시에 다운되면 묶어서 한 번에 전달
  group_wait: 30s # 알림 대기 시간. 첫 알림 발생 후 30초를 기다렸다가 발송(동시에 다운되도 한 번만 보내기 위함)
  group_interval: 5m
  repeat_interval: 1h # 문제가 해결되지 않은 경우, 같은 알림 발송 주기

receivers: # 알림을 보낼곳 설정
  - name: 'webhook-server'
    webhook_configs:
      - url: 'http://127.0.0.1:5001/alert'
        send_resolved: true # 문제 해결됐을 때도 보낼지 여부
EOF

amtool check-config /etc/alertmanager/alertmanager.yml # 문법 검사

 

Step 3. 서비스 파일 생성

cat > /etc/systemd/system/alertmanager.service << 'EOF'
[Unit]
Description=Alertmanager
After=network.target

[Service]
User=alertmanager
Group=alertmanager
Type=simple
ExecStart=/usr/local/bin/alertmanager \
  --config.file=/etc/alertmanager/alertmanager.yml \
  --storage.path=/var/lib/alertmanager/ \
  --web.listen-address=127.0.0.1:9093

[Install]
WantedBy=multi-user.target
EOF

systemctl daemon-reload
systemctl start alertmanager
systemctl enable alertmanager

 

Alertmanager의 요청을 받을 웹훅서버는 원하는 방식대로 만들면 됩니다.

이 글에서는 웹훅 서버를 만드는 방법에 대해서는 넘어가도록 하겠습니다.