← 모든 글

Liveness와 Readiness Probe 구성 - 헬스체크와 자동 복구

Liveness·Readiness·Startup Probe의 역할과 설정 방법을 학습용 구성 예제로 정리합니다

CKA 준비 · 6 / 8편

이 글의 목차

학습 예시 안내: 본문의 상황과 출력은 개념 설명을 위한 예시로 정리했습니다. 당시 실행 여부는 확인되지 않았으며, 현재 환경에서의 클러스터 실습도 별도 검증이 필요합니다.

🎯 학습 목표

이 글을 통해 다음을 학습할 수 있습니다:

  • Pod 헬스체크의 필요성과 Probe의 역할
  • Liveness, Readiness, Startup Probe의 차이점과 활용법
  • 다양한 Probe 타입(HTTP, TCP, Command) 구성 방법
  • Probe 설정 최적화와 트러블슈팅 기법
  • CKA 준비 과정에서 연습할 Probe 관련 문제 패턴

📝 앞선 개념과 이번 예제

지난 글에서는 Pod 리소스 요청과 제한을 다뤘습니다. 이번에는 리소스가 충분해도 애플리케이션이 요청을 처리하지 못하는 상황을 가정합니다.

예제에서 가정한 환경

# 기존 리소스 제한이 적용된 Deployment 확인
kubectl get deployments
kubectl describe deployment web-app

# Pod 상태와 리소스 사용량 확인
kubectl get pods -o wide
kubectl top pods

문제 상황 시뮬레이션

프로세스는 실행 중이지만 애플리케이션이 준비되지 않았거나 응답하지 않는 상황을 가정해 봅니다:

# 웹 애플리케이션 테스트
kubectl run test-client --image=curlimages/curl --rm -it -- sh
# 컨테이너 내에서: curl http://web-app-service

가정한 상황의 문제점:

  1. 애플리케이션이 시작되지 않았는데 트래픽 수신: Pod는 Running 상태지만 실제로는 아직 준비되지 않음
  2. 장애가 발생해도 자동 복구 안됨: 애플리케이션이 hang 상태여도 Pod는 그대로 유지
  3. 시작 시간이 긴 애플리케이션 문제: 데이터베이스 연결 등으로 시작이 늦어지는 경우
  4. 좀비 프로세스 문제: 프로세스는 살아있지만 요청을 처리하지 못하는 상태
# 문제 상황 재현을 위한 테스트 Pod 생성
kubectl run slow-app --image=nginx --command -- sh -c "sleep 30 && nginx -g 'daemon off;'"

# Pod 상태 확인
kubectl get pod slow-app
# STATUS가 Running이지만 실제로는 nginx가 아직 시작되지 않음

# 서비스 연결 시도
kubectl expose pod slow-app --port=80
kubectl run test --image=curlimages/curl --rm -it --command -- curl http://slow-app
# 연결 실패 또는 타임아웃 발생

이런 상태를 구분하기 위해 Probe를 이용한 헬스체크와 자동 복구 설정을 살펴보겠습니다.

💡 Probe의 종류와 역할 이해하기

기본 개념 정리

쿠버네티스는 세 가지 종류의 Probe를 제공합니다:

Liveness Probe (생존 확인)

  • 컨테이너가 정상적으로 동작하고 있는지 확인
  • 연속 실패가 failureThreshold에 도달하면 해당 컨테이너를 종료하고 재시작 정책을 적용
  • 용도: 데드락, 무한루프 등 복구 가능한 장애 감지

Readiness Probe (준비 상태 확인)

  • Pod가 트래픽을 받을 준비가 되었는지 확인
  • 실패 시 → Service 엔드포인트에서 제거
  • 용도: 초기화 완료, 의존성 연결 확인

Startup Probe (시작 확인)

  • 컨테이너가 시작을 완료했는지 확인
  • 연속 실패가 failureThreshold에 도달하면 해당 컨테이너를 종료하고 재시작 정책을 적용
  • 용도: 시작 시간이 긴 애플리케이션 처리
시작 → Startup Probe 성공 (설정한 경우)
           ├─ Readiness Probe: 트래픽을 받을 준비 확인
           └─ Liveness Probe: 컨테이너 복구 필요 여부 확인

Startup이 성공한 뒤 Readiness와 Liveness는 각각 독립적으로 동작합니다. Readiness가 성공해야 Liveness가 시작되는 순서는 아닙니다. Probe 공식 문서

첫 번째 Probe 설정

간단한 HTTP Probe 구성 예제입니다:

# basic-probe-pod.yaml
apiVersion: v1
kind: Pod
metadata:
  name: web-with-probes
spec:
  containers:
  - name: web
    image: nginx
    ports:
    - containerPort: 80
    # Readiness Probe: 트래픽 받을 준비 확인
    readinessProbe:
      httpGet:
        path: /
        port: 80
      initialDelaySeconds: 5    # 5초 후 시작
      periodSeconds: 10         # 10초마다 확인
    # Liveness Probe: 정상 동작 확인
    livenessProbe:
      httpGet:
        path: /
        port: 80
      initialDelaySeconds: 30   # 30초 후 시작
      periodSeconds: 30         # 30초마다 확인
      timeoutSeconds: 5         # 5초 타임아웃
      failureThreshold: 3       # 3번 실패 시 재시작
# Pod 생성 및 상태 확인
kubectl apply -f basic-probe-pod.yaml

# Probe 상태 모니터링
watch kubectl describe pod web-with-probes

중요한 관찰 포인트:

  • Ready 상태 변화 추적
  • Conditions 섹션에서 PodReadyCondition 확인
  • Events에서 Probe 실행 결과 확인

🔍 Probe 타입별 상세 실습

1. HTTP GET Probe

가장 일반적으로 사용되는 HTTP 기반 헬스체크입니다:

# http-probe-app.yaml
apiVersion: v1
kind: Pod
metadata:
  name: http-probe-app
spec:
  containers:
  - name: app
    image: httpd:2.4
    ports:
    - containerPort: 80
    readinessProbe:
      httpGet:
        path: /
        port: 80
        httpHeaders:           # 필요시 커스텀 헤더 추가
        - name: Custom-Header
          value: Health-Check
      initialDelaySeconds: 5
      periodSeconds: 10
      successThreshold: 1      # 1번 성공하면 Ready
      failureThreshold: 3      # 3번 실패하면 Not Ready
    livenessProbe:
      httpGet:
        path: /                # 기본 이미지가 제공하는 경로
        port: 80
      initialDelaySeconds: 30
      periodSeconds: 20
      timeoutSeconds: 10
      failureThreshold: 3

헬스체크 엔드포인트 테스트:

# Pod 생성
kubectl apply -f http-probe-app.yaml

# 수동으로 헬스체크 엔드포인트 테스트
kubectl port-forward pod/http-probe-app 8080:80
# 다른 터미널에서 확인 (종료할 때 port-forward 터미널에서 Ctrl+C)
curl http://127.0.0.1:8080/

2. TCP Socket Probe

HTTP가 아닌 TCP 서비스의 경우 사용합니다:

# tcp-probe-app.yaml
apiVersion: v1
kind: Pod
metadata:
  name: tcp-probe-app
spec:
  containers:
  - name: redis
    image: redis:6.2
    ports:
    - containerPort: 6379
    readinessProbe:
      tcpSocket:
        port: 6379             # Redis 포트 연결 확인
      initialDelaySeconds: 10
      periodSeconds: 5
    livenessProbe:
      tcpSocket:
        port: 6379
      initialDelaySeconds: 30
      periodSeconds: 10
      failureThreshold: 3
# Redis Pod 생성 및 테스트
kubectl apply -f tcp-probe-app.yaml

# TCP 연결 수동 테스트
kubectl exec -it tcp-probe-app -- redis-cli ping
# PONG 응답 확인

3. Command Probe

커스텀 명령어를 실행하여 상태를 확인합니다:

# command-probe-app.yaml
apiVersion: v1
kind: Pod
metadata:
  name: command-probe-app
spec:
  containers:
  - name: app
    image: busybox
    command: ["/bin/sh"]
    args: ["-c", "while true; do sleep 30; done"]
    readinessProbe:
      exec:
        command:
        - cat
        - /tmp/ready          # 파일 존재 여부로 준비 상태 확인
      initialDelaySeconds: 10
      periodSeconds: 5
    livenessProbe:
      exec:
        command:
        - /bin/sh
        - -c
        - "[ $(ps aux | grep -v grep | grep sleep | wc -l) -gt 0 ]"  # 프로세스 존재 확인
      initialDelaySeconds: 30
      periodSeconds: 15
# Pod 생성
kubectl apply -f command-probe-app.yaml

# Ready 상태로 만들기 위해 파일 생성
kubectl exec -it command-probe-app -- touch /tmp/ready

# Pod 상태 변화 확인
kubectl get pod command-probe-app

🚀 Startup Probe로 느린 시작 애플리케이션 처리

시작 시간이 긴 애플리케이션을 위한 Startup Probe 실습입니다:

# slow-startup-app.yaml
apiVersion: v1
kind: Pod
metadata:
  name: slow-startup-app
spec:
  containers:
  - name: slow-app
    image: nginx
    command: ["/bin/sh"]
    args: ["-c", "echo 'Starting up...' && sleep 60 && nginx -g 'daemon off;'"]
    ports:
    - containerPort: 80
    # Startup Probe: 최초 30초 지연 후 주기 10초·연속 실패 기준 12회
    startupProbe:
      httpGet:
        path: /
        port: 80
      initialDelaySeconds: 30
      periodSeconds: 10
      failureThreshold: 12     # 연속 12회 실패하면 컨테이너 종료
    # Readiness Probe: Startup 완료 후 작동
    readinessProbe:
      httpGet:
        path: /
        port: 80
      periodSeconds: 5
    # Liveness Probe: Startup 완료 후 작동
    livenessProbe:
      httpGet:
        path: /
        port: 80
      periodSeconds: 10
      failureThreshold: 3
# 느린 시작 애플리케이션 생성
kubectl apply -f slow-startup-app.yaml

# 시작 과정 모니터링
watch kubectl describe pod slow-startup-app

중요한 동작 방식:

  • Startup Probe가 성공할 때까지 Liveness/Readiness Probe는 비활성화
  • Startup Probe가 실패 임계치에 도달하면 해당 컨테이너 종료 후 재시작 정책 적용
  • Startup Probe 성공 후 일반적인 Probe 동작 시작

💡 실전 시나리오: 다양한 애플리케이션별 Probe 설정

아래는 설정 구조 예시입니다. Spring Boot JAR·Node.js 코드·헬스체크 경로가 포함된 애플리케이션 이미지를 직접 준비해야 합니다. 기본 런타임 이미지만으로 해당 파일이나 API가 생기지는 않습니다. 이 글의 이미지 태그는 원문에 있던 값이며, 현재 환경에서의 실행은 별도 확인이 필요합니다.

1. 웹 애플리케이션 (Spring Boot)

# spring-boot-app.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: spring-boot-app
spec:
  replicas: 3
  selector:
    matchLabels:
      app: spring-boot
  template:
    metadata:
      labels:
        app: spring-boot
    spec:
      containers:
      - name: app
        image: openjdk:11-jre-slim
        command: ["java"]
        args: ["-jar", "/app/spring-boot-app.jar"]
        ports:
        - containerPort: 8080
        # 시작이 느린 Spring Boot 앱을 위한 Startup Probe
        startupProbe:
          httpGet:
            path: /actuator/health
            port: 8080
          initialDelaySeconds: 30
          periodSeconds: 10
          failureThreshold: 30    # 최대 5분 대기
        # 의존성 준비 확인
        readinessProbe:
          httpGet:
            path: /actuator/health/readiness
            port: 8080
          periodSeconds: 10
          successThreshold: 1
          failureThreshold: 3
        # 애플리케이션 데드락 감지
        livenessProbe:
          httpGet:
            path: /actuator/health/liveness
            port: 8080
          periodSeconds: 20
          timeoutSeconds: 5
          failureThreshold: 3
        resources:
          requests:
            memory: "512Mi"
            cpu: "250m"
          limits:
            memory: "1Gi"
            cpu: "500m"

2. 데이터베이스 (PostgreSQL)

# postgres-app.yaml
apiVersion: v1
kind: Pod
metadata:
  name: postgres-db
spec:
  containers:
  - name: postgres
    image: postgres:13
    env:
    - name: POSTGRES_PASSWORD
      value: "mypassword"
    - name: POSTGRES_DB
      value: "mydb"
    ports:
    - containerPort: 5432
    # 데이터베이스 초기화 완료 확인
    startupProbe:
      exec:
        command:
        - /bin/sh
        - -c
        - "pg_isready -U postgres -d mydb"
      initialDelaySeconds: 30
      periodSeconds: 10
      failureThreshold: 30
    # 연결 가능 상태 확인
    readinessProbe:
      exec:
        command:
        - /bin/sh
        - -c
        - "pg_isready -U postgres -d mydb"
      periodSeconds: 10
    # 데이터베이스 프로세스 확인
    livenessProbe:
      exec:
        command:
        - /bin/sh
        - -c
        - "pg_isready -U postgres"
      periodSeconds: 30
      failureThreshold: 3

3. 마이크로서비스 API

# microservice-api.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: user-api
spec:
  replicas: 2
  selector:
    matchLabels:
      app: user-api
  template:
    metadata:
      labels:
        app: user-api
    spec:
      containers:
      - name: api
        image: node:16-alpine
        command: ["node"]
        args: ["server.js"]
        ports:
        - containerPort: 3000
        env:
        - name: DB_HOST
          value: "postgres-db"
        # API 서버 시작 확인
        startupProbe:
          httpGet:
            path: /health
            port: 3000
          initialDelaySeconds: 10
          periodSeconds: 5
          failureThreshold: 12
        # 의존성 서비스 연결 확인
        readinessProbe:
          httpGet:
            path: /health/ready    # DB 연결 등 의존성 확인
            port: 3000
          periodSeconds: 10
          successThreshold: 1
          failureThreshold: 3
        # API 응답 확인
        livenessProbe:
          httpGet:
            path: /health/live     # 기본적인 API 응답 확인
            port: 3000
          periodSeconds: 15
          timeoutSeconds: 5
          failureThreshold: 3

🚨 Probe 실패 상황과 문제 해결

일반적인 실패 시나리오 실습

1. Readiness Probe 실패 시뮬레이션

# readiness-fail-test.yaml
apiVersion: v1
kind: Pod
metadata:
  name: readiness-fail-test
  labels:
    app: readiness-fail-test
spec:
  containers:
  - name: web
    image: nginx
    ports:
    - containerPort: 80
    readinessProbe:
      httpGet:
        path: /nonexistent     # 존재하지 않는 경로
        port: 80
      periodSeconds: 5
      failureThreshold: 2
# Pod 생성 및 Service 생성
kubectl apply -f readiness-fail-test.yaml
kubectl expose pod readiness-fail-test --port=80 --name=fail-test-service

# Service 엔드포인트 확인
kubectl get endpoints fail-test-service
# ENDPOINTS 컬럼이 비어있음을 확인

# Pod 상태 확인
kubectl describe pod readiness-fail-test
# Ready: False 상태 확인

2. Liveness Probe 실패 시뮬레이션

# liveness-fail-test.yaml
apiVersion: v1
kind: Pod
metadata:
  name: liveness-fail-test
spec:
  containers:
  - name: app
    image: busybox
    command: ["/bin/sh"]
    args: ["-c", "touch /tmp/healthy && sleep 30 && rm /tmp/healthy && sleep 3600"]
    livenessProbe:
      exec:
        command:
        - cat
        - /tmp/healthy
      initialDelaySeconds: 5
      periodSeconds: 10
      failureThreshold: 2
# Probe를 포함한 위 매니페스트로 처음부터 생성
kubectl apply -f liveness-fail-test.yaml

# 컨테이너 재시작 모니터링
watch kubectl get pod liveness-fail-test
# RESTARTS 컬럼 증가 확인

이미 실행 중인 단독 Pod의 Probe는 일반적인 edit·patch로 추가하거나 변경할 수 없습니다. 수정한 매니페스트로 Pod를 다시 만들거나, Deployment의 Pod 템플릿을 수정해 교체해야 합니다. Pod 변경 제약

디버깅과 문제 해결

Probe 실패 원인 분석:

# Pod 이벤트 확인
kubectl describe pod <pod-name>

# 상세한 이벤트 로그
kubectl get events --field-selector involvedObject.name=<pod-name> --sort-by=.metadata.creationTimestamp

# 컨테이너 로그 확인
kubectl logs <pod-name> -c <container-name>

# 이전 컨테이너 로그 (재시작된 경우)
kubectl logs <pod-name> -c <container-name> --previous

일반적인 문제와 해결책:

문제 원인 해결책
Readiness 계속 실패 의존성 서비스 미준비 initialDelaySeconds 증가, 의존성 확인
Liveness 간헐적 실패 일시적 부하 timeoutSeconds, failureThreshold 조정
Startup 타임아웃 시작 시간 부족 failureThreshold 증가, periodSeconds 조정
너무 잦은 재시작 너무 엄격한 설정 각 threshold와 timeout 값 완화

🎯 Probe 설정 최적화 가이드

권장 설정 값

다음은 시간·횟수만 발췌한 조정 예시입니다. 실제 적용 시 각 Probe에 httpGet·tcpSocket·exec 등의 검사 방식도 넣어야 하며, 시작 대기 시간에는 initialDelaySeconds도 함께 고려해야 합니다.

일반적인 웹 애플리케이션:

startupProbe:
  initialDelaySeconds: 10
  periodSeconds: 10
  failureThreshold: 30      # 최대 5분

readinessProbe:
  initialDelaySeconds: 0    # Startup 완료 후 즉시 시작
  periodSeconds: 10
  successThreshold: 1
  failureThreshold: 3       # 30초 내 복구 기회

livenessProbe:
  initialDelaySeconds: 0
  periodSeconds: 30         # 자주 확인할 필요 없음
  timeoutSeconds: 5
  failureThreshold: 3       # 90초 내 복구 기회

데이터베이스:

startupProbe:
  initialDelaySeconds: 30   # DB 초기화 시간 고려
  periodSeconds: 10
  failureThreshold: 60      # 최대 10분

readinessProbe:
  periodSeconds: 10
  timeoutSeconds: 5
  failureThreshold: 3

livenessProbe:
  periodSeconds: 60         # DB는 덜 자주 확인
  timeoutSeconds: 10
  failureThreshold: 3

성능 최적화 고려사항

# Probe 설정이 클러스터에 미치는 영향 분석
kubectl top nodes
kubectl top pods --all-namespaces

# 특정 Pod의 Probe 실행 빈도 계산
# Readiness: 10초마다, Liveness: 30초마다
# = 1분에 총 8번의 헬스체크 요청

리소스 사용량 최적화:

  • periodSeconds: 너무 자주 확인하지 않도록 조정
  • timeoutSeconds: 네트워크 지연 고려
  • failureThreshold: 일시적 장애 허용도 고려

🎯 CKA 대비 연습 패턴

연습할 문제 유형

1. 기존 Pod/Deployment에 Probe 추가

# 문제: nginx Deployment에 liveness probe 추가
kubectl edit deployment nginx-deployment

# 또는 patch 명령 사용
kubectl patch deployment nginx-deployment -p '{"spec":{"template":{"spec":{"containers":[{"name":"nginx","livenessProbe":{"httpGet":{"path":"/","port":80},"initialDelaySeconds":30,"periodSeconds":10}}]}}}}'

2. 문제가 있는 Probe 설정 수정

# Probe 실패 원인 찾기
kubectl describe pod <pod-name>
kubectl get events

# Deployment가 관리하는 Pod는 템플릿에서 Probe 수정
kubectl edit deployment <deployment-name>
# 단독 Pod는 파일에서 수정 후 다시 생성

3. 다양한 타입의 Probe 구성

# Probe 구성을 연습하는 패턴
readinessProbe:
  httpGet:              # 또는 tcpSocket, exec
    path: /health
    port: 8080
  initialDelaySeconds: 5
  periodSeconds: 10

livenessProbe:
  httpGet:
    path: /
    port: 80
  initialDelaySeconds: 30
  periodSeconds: 30

4. 트러블슈팅 문제

  • Pod가 Ready 상태가 되지 않는 이유 찾기
  • Service에서 트래픽을 받지 못하는 Pod 문제 해결
  • 계속 재시작되는 Pod 원인 분석

시간 단축 팁

# 빠른 Probe 설정이 포함된 Pod 생성
kubectl run nginx --image=nginx --port=80 \
  --dry-run=client -o yaml > pod.yaml

# 편집기에서 spec.containers의 해당 컨테이너 아래에 Probe 추가
# 파일 끝에 단순히 덧붙이면 잘못된 위치에 들어갈 수 있음
vi pod.yaml

# Pod 생성
kubectl apply -f pod.yaml

자주 사용하는 Probe 명령어:

# Probe 상태 빠른 확인
kubectl get pods -o wide
kubectl describe pod <name> | grep -A 10 Conditions

# Probe 실패 이벤트 확인
kubectl get events --field-selector reason=Unhealthy

📊 모니터링과 알림 설정

Probe 상태 모니터링

# 클러스터 전체 Pod 상태 확인
kubectl get pods --all-namespaces -o wide

# Running이 아닌 Pod 필터링 (Ready=False와는 다른 조건)
kubectl get pods --all-namespaces --field-selector=status.phase!=Running

# Probe 실패 이벤트 모니터링
kubectl get events --all-namespaces --field-selector reason=Unhealthy -w

커스텀 헬스체크 엔드포인트 개발

실제 애플리케이션에서 사용할 수 있는 헬스체크 예제:

// Node.js Express 애플리케이션 예제
const express = require('express');
const app = express();

let isReady = false;
let isHealthy = true;

// 시작 시 의존성 초기화
setTimeout(() => {
  // DB 연결, 캐시 워밍업 등
  isReady = true;
}, 10000);  // 10초 후 준비 완료

// Liveness Probe 엔드포인트
app.get('/health/live', (req, res) => {
  if (isHealthy) {
    res.status(200).send('OK');
  } else {
    res.status(503).send('Service Unavailable');
  }
});

// Readiness Probe 엔드포인트
app.get('/health/ready', (req, res) => {
  if (isReady && isHealthy) {
    res.status(200).send('Ready');
  } else {
    res.status(503).send('Not Ready');
  }
});

// Startup Probe 엔드포인트
app.get('/health/startup', (req, res) => {
  if (isReady) {
    res.status(200).send('Started');
  } else {
    res.status(503).send('Starting');
  }
});

app.listen(3000);

📚 필수 명령어 정리

Probe 관련 핵심 명령어

# Pod 상태와 Probe 정보 확인
kubectl describe pod <pod-name>
kubectl get pod <pod-name> -o yaml

# Probe 실패 이벤트 확인
kubectl get events --field-selector involvedObject.name=<pod-name>
kubectl get events --field-selector reason=Unhealthy

# 실시간 Pod 상태 모니터링
watch kubectl get pods
kubectl get pods -w

# 특정 조건의 Pod 찾기
kubectl get pods --field-selector=status.phase=Pending
kubectl get pods -o jsonpath='{.items[?(@.status.containerStatuses[*].ready==false)].metadata.name}'

# Service 엔드포인트 확인
kubectl get endpoints <service-name>
kubectl describe service <service-name>

Probe 설정 템플릿

# 기본 HTTP Probe 템플릿
readinessProbe:
  httpGet:
    path: /health/ready
    port: 8080
  initialDelaySeconds: 5
  periodSeconds: 10
  successThreshold: 1
  failureThreshold: 3

livenessProbe:
  httpGet:
    path: /health/live
    port: 8080
  initialDelaySeconds: 30
  periodSeconds: 30
  timeoutSeconds: 5
  failureThreshold: 3

---
# 별도 예제: TCP Probe 템플릿
readinessProbe:
  tcpSocket:
    port: 6379
  initialDelaySeconds: 10
  periodSeconds: 5

---
# 별도 예제: Command Probe 템플릿
livenessProbe:
  exec:
    command:
    - /bin/sh
    - -c
    - "pg_isready -U postgres"
  periodSeconds: 30
  failureThreshold: 3

🔄 고급 Probe 활용 패턴

1. 다단계 헬스체크

복잡한 애플리케이션에서는 여러 단계의 헬스체크를 구성할 수 있습니다:

아래 /health/* 경로는 직접 구현해야 합니다. 기본 nginx 이미지에는 없으므로, 이 파일은 애플리케이션에 맞게 바꾸는 구성 예시입니다.

# multi-tier-health-check.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: multi-tier-app
spec:
  replicas: 2
  selector:
    matchLabels:
      app: multi-tier-app
  template:
    metadata:
      labels:
        app: multi-tier-app
    spec:
      containers:
      - name: app
        image: nginx  # 실제로는 복잡한 애플리케이션
        ports:
        - containerPort: 80
        # 1단계: 기본 서비스 시작 확인
        startupProbe:
          httpGet:
            path: /health/startup
            port: 80
          initialDelaySeconds: 10
          periodSeconds: 5
          failureThreshold: 24    # 2분 대기
        # 2단계: 의존성과 캐시 준비 확인
        readinessProbe:
          httpGet:
            path: /health/ready
            port: 80
            httpHeaders:
            - name: X-Health-Check
              value: "readiness"
          periodSeconds: 10
          successThreshold: 2     # 2번 연속 성공해야 Ready
          failureThreshold: 3
        # 3단계: 지속적인 건강 상태 모니터링
        livenessProbe:
          httpGet:
            path: /health/live
            port: 80
            httpHeaders:
            - name: X-Health-Check
              value: "liveness"
          periodSeconds: 30
          timeoutSeconds: 10
          failureThreshold: 3

2. 조건부 Probe (Init Container 활용)

의존성 서비스가 준비된 후에만 메인 컨테이너를 시작하는 패턴:

# conditional-startup.yaml
apiVersion: v1
kind: Pod
metadata:
  name: conditional-startup-app
spec:
  initContainers:
  # 의존성 서비스 대기
  - name: wait-for-db
    image: busybox
    command: 
    - /bin/sh
    - -c
    - |
      echo "Waiting for database..."
      until nc -z postgres-service 5432; do
        echo "Database not ready, sleeping..."
        sleep 2
      done
      echo "Database is ready!"
  - name: wait-for-cache
    image: busybox
    command:
    - /bin/sh
    - -c
    - |
      echo "Waiting for cache..."
      until nc -z redis-service 6379; do
        echo "Cache not ready, sleeping..."
        sleep 2
      done
      echo "Cache is ready!"
  containers:
  - name: app
    image: myapp:latest
    ports:
    - containerPort: 8080
    # 의존성이 모두 준비된 상태에서 시작하므로 빠른 Probe 가능
    readinessProbe:
      httpGet:
        path: /health
        port: 8080
      initialDelaySeconds: 5
      periodSeconds: 5
    livenessProbe:
      httpGet:
        path: /health
        port: 8080
      initialDelaySeconds: 30
      periodSeconds: 20

3. 연속 성공 후 트래픽에 참여하는 패턴

successThreshold: 3은 세 번 연속 성공한 뒤 Ready로 전환하는 조건입니다. 트래픽 비율을 조금씩 늘리는 기능은 아니며, 가중치 조정은 별도의 라우팅 구성이 필요합니다.

# gradual-traffic-increase.yaml
apiVersion: v1
kind: Service
metadata:
  name: gradual-service
spec:
  selector:
    app: gradual-app
  ports:
  - port: 80
    targetPort: 8080
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: gradual-app
spec:
  replicas: 3
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 1
      maxSurge: 1
  selector:
    matchLabels:
      app: gradual-app
  template:
    metadata:
      labels:
        app: gradual-app
    spec:
      containers:
      - name: app
        image: myapp:latest
        ports:
        - containerPort: 8080
        readinessProbe:
          httpGet:
            path: /health/ready
            port: 8080
          initialDelaySeconds: 10
          periodSeconds: 5
          successThreshold: 3      # 3번 연속 성공 후 트래픽 수신
          failureThreshold: 2      # 빠른 트래픽 차단
        livenessProbe:
          httpGet:
            path: /health/live
            port: 8080
          initialDelaySeconds: 60
          periodSeconds: 30
          failureThreshold: 5      # 관대한 실패 허용

🛠️ 실전 문제 해결 시나리오

아래 장애와 로그는 진단 방법을 설명하기 위한 가정입니다. 실제 장애를 관찰하거나 해결한 결과로 제시하는 것은 아닙니다.

시나리오 1: 간헐적 Readiness Probe 실패

문제 상황:

# 서비스가 간헐적으로 응답하지 않음
kubectl get pods -l app=unstable-app
# 일부 Pod가 Ready 상태와 Not Ready 상태를 반복

문제 분석:

# 이벤트 확인
kubectl get events --field-selector involvedObject.name=unstable-app-xxx

# 로그 분석
kubectl logs unstable-app-xxx --tail=100

# 리소스 사용량 확인
kubectl top pod unstable-app-xxx

해결 방법:

# 더 관대한 Readiness Probe 설정
readinessProbe:
  httpGet:
    path: /health
    port: 8080
  periodSeconds: 10
  timeoutSeconds: 10       # 타임아웃 증가
  successThreshold: 2      # 2번 연속 성공 필요
  failureThreshold: 5      # 5번 실패 허용

시나리오 2: 메모리 누수로 인한 반복적 재시작

문제 분석:

# 재시작 횟수 확인
kubectl get pods -l app=memory-leak-app
# RESTARTS 컬럼이 계속 증가

# 메모리 사용량 트렌드 확인
kubectl top pods -l app=memory-leak-app --containers

# 상세 이벤트 확인
kubectl describe pod memory-leak-app-xxx
# Reason: OOMKilled 이벤트 확인

임시 해결책:

메모리 제한 상향은 누수 수정을 대신하지 않습니다. Liveness 간격을 늘려도 커널의 OOM 종료를 막을 수는 없습니다.

# 리소스 제한 증가 (근본 해결은 아님)
resources:
  limits:
    memory: "2Gi"      # 기존 1Gi에서 증가
  requests:
    memory: "512Mi"

# Liveness Probe 간격 조정
livenessProbe:
  httpGet:
    path: /health
    port: 8080
  periodSeconds: 60      # 더 긴 간격으로 설정
  failureThreshold: 5    # 더 관대하게 설정

시나리오 3: 데이터베이스 연결 풀 고갈

문제 상황:

# 애플리케이션이 간헐적으로 응답하지 않음
kubectl logs app-pod | grep "connection"
# "Connection pool exhausted" 에러 발견

헬스체크 개선 (연결 풀 고갈 원인 수정은 별도):

# 애플리케이션별 헬스체크 엔드포인트 개선
readinessProbe:
  httpGet:
    path: /health/db-connection    # DB 연결 상태 전용 체크
    port: 8080
  periodSeconds: 15               # 체크 빈도 조정
  timeoutSeconds: 5
  failureThreshold: 2

livenessProbe:
  httpGet:
    path: /health/basic           # 기본적인 프로세스 체크만
    port: 8080
  periodSeconds: 45               # DB에 부담을 주지 않도록
  failureThreshold: 3

🎭 고급 패턴: Circuit Breaker와 연동

Circuit Breaker 상태를 반영한 Probe

의존 서비스 장애는 Readiness 판단에 반영하고, Liveness는 애플리케이션 자체의 복구 불가능한 상태를 확인하도록 설계합니다. 외부 장애만으로 모든 복제본이 반복 재시작되지 않도록 구분해야 합니다.

# circuit-breaker-aware.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: resilient-app
spec:
  replicas: 3
  selector:
    matchLabels:
      app: resilient-app
  template:
    metadata:
      labels:
        app: resilient-app
    spec:
      containers:
      - name: app
        image: resilient-app:latest
        ports:
        - containerPort: 8080
        env:
        - name: CIRCUIT_BREAKER_ENABLED
          value: "true"
        readinessProbe:
          httpGet:
            path: /health/ready
            port: 8080
            httpHeaders:
            - name: X-Check-Dependencies
              value: "true"
          periodSeconds: 10
          successThreshold: 1
          failureThreshold: 3
        livenessProbe:
          httpGet:
            path: /health/live
            port: 8080
          periodSeconds: 30
          failureThreshold: 5

📈 성능 최적화와 베스트 프랙티스

1. Probe 오버헤드 최소화

# 클러스터 전체 Probe 실행 빈도 계산
kubectl get pods --all-namespaces -o jsonpath='{range .items[*]}{.metadata.name}{" "}{.spec.containers[*].readinessProbe.periodSeconds}{" "}{.spec.containers[*].livenessProbe.periodSeconds}{"\n"}{end}' | head -20

# 높은 빈도의 Probe 식별
kubectl get pods --all-namespaces -o yaml | grep -A 3 -B 3 "periodSeconds: [1-5]"

2. 효율적인 헬스체크 엔드포인트 설계

# Python Flask 예제 - 효율적인 헬스체크
from flask import Flask, jsonify
import time
import threading

app = Flask(__name__)

# 캐시된 상태 정보
health_cache = {
    'last_check': 0,
    'status': 'unknown',
    'cache_duration': 30  # 30초 캐시
}

def check_dependencies():
    """실제 의존성 체크 (무거운 작업)"""
    # DB 연결, 외부 API 호출 등
    return True

@app.route('/health/live')
def liveness():
    """가벼운 프로세스 체크만"""
    return jsonify({'status': 'alive'}), 200

@app.route('/health/ready')
def readiness():
    """캐시를 활용한 효율적인 의존성 체크"""
    current_time = time.time()
    
    # 캐시가 유효한 경우
    if current_time - health_cache['last_check'] < health_cache['cache_duration']:
        if health_cache['status'] == 'ready':
            return jsonify({'status': 'ready'}), 200
        else:
            return jsonify({'status': 'not ready'}), 503
    
    # 캐시 갱신 필요
    try:
        if check_dependencies():
            health_cache['status'] = 'ready'
            health_cache['last_check'] = current_time
            return jsonify({'status': 'ready'}), 200
        else:
            health_cache['status'] = 'not ready'
            health_cache['last_check'] = current_time
            return jsonify({'status': 'not ready'}), 503
    except Exception as e:
        health_cache['status'] = 'error'
        health_cache['last_check'] = current_time
        return jsonify({'status': 'error', 'message': str(e)}), 503

3. 환경별 Probe 설정 관리

아래 ConfigMap만 생성한다고 Pod의 Probe 설정이 바뀌지는 않습니다. Helm·Kustomize 또는 별도 생성 단계에서 값을 Pod 템플릿에 반영해야 합니다.

# ConfigMap으로 환경별 설정 관리
apiVersion: v1
kind: ConfigMap
metadata:
  name: probe-config
data:
  development.yaml: |
    probes:
      readiness:
        initialDelaySeconds: 5
        periodSeconds: 5
        failureThreshold: 3
      liveness:
        initialDelaySeconds: 30
        periodSeconds: 30
        failureThreshold: 3
  
  production.yaml: |
    probes:
      readiness:
        initialDelaySeconds: 10
        periodSeconds: 10
        failureThreshold: 5
      liveness:
        initialDelaySeconds: 60
        periodSeconds: 60
        failureThreshold: 5

🚀 다음 학습 계획

다음에는 여러 애플리케이션과 팀을 위한 네임스페이스 격리와 리소스 관리 방법을 살펴볼 계획입니다.

다음 주제들

  • 네임스페이스를 통한 리소스 격리와 다중 테넌시 - 팀별, 환경별 격리
  • Ingress를 통한 외부 트래픽 관리와 라우팅
  • 스토리지와 PersistentVolume 관리
  • RBAC과 보안 설정
  • 모니터링과 로깅 시스템 구축

개인 학습 목표

  • 다양한 애플리케이션 유형별 최적 Probe 패턴 숙달
  • 대규모 클러스터에서의 효율적인 헬스체크 설계 경험
  • 장애 상황에서의 빠른 진단과 복구 능력 향상
  • 모니터링과 알림을 통한 proactive한 운영 체계 구축

Probe를 구성할 때는 시작 완료, 요청 처리 준비, 재시작이 필요한 상태를 구분하는 것이 핵심입니다. 각 검사 경로와 실패 조건은 실제 애플리케이션 동작에 맞춰 구현하고 검증해야 합니다.

다음 글에서는 팀별·환경별 리소스를 나누는 네임스페이스 활용 방법을 다룹니다.


🔗 참고 자료:


yellow.log · 2025.08.11에 작성 · 2026.09.29에 수정

CKA 준비 전체 8편

  1. 1.Pod 생성과 관리 - 쿠버네티스의 기본 단위 이해하기
  2. 2.ReplicaSet과 Deployment로 Pod 관리하기 - 확장성과 가용성 확보
  3. 3.Service를 통한 Pod 네트워킹 - 안정적인 접근 경로 만들기
  4. 4.ConfigMap과 Secret으로 설정 관리하기 - 환경별 설정 분리
  5. 5.Pod 리소스 제한과 요청 설정 - 안정적인 클러스터 운영을 위한 리소스 관리
  6. 6.Liveness와 Readiness Probe 구성 - 헬스체크와 자동 복구읽는 중
  7. 7.네임스페이스를 통한 리소스 격리와 다중 테넌시 - 팀별, 환경별 격리
  8. 8.Ingress를 통한 외부 트래픽 관리와 라우팅 - 클러스터의 똑똑한 관문 구축