학습 예시 안내: 본문의 상황과 출력은 개념 설명을 위한 예시로 정리했습니다. 당시 실행 여부는 확인되지 않았으며, 현재 환경에서의 클러스터 실습도 별도 검증이 필요합니다.
🎯 학습 목표
이 글을 통해 다음을 학습할 수 있습니다:
- Pod 헬스체크의 필요성과 Probe의 역할
- Liveness, Readiness, Startup Probe의 차이점과 활용법
- 다양한 Probe 타입(HTTP, TCP, Command) 구성 방법
- Probe 설정 최적화와 트러블슈팅 기법
- CKA 준비 과정에서 연습할 Probe 관련 문제 패턴
📝 앞선 개념과 이번 예제
지난 글에서는 Pod 리소스 요청과 제한을 다뤘습니다. 이번에는 리소스가 충분해도 애플리케이션이 요청을 처리하지 못하는 상황을 가정합니다.
예제에서 가정한 환경
# 기존 리소스 제한이 적용된 Deployment 확인
kubectl get deployments
kubectl describe deployment web-app
# Pod 상태와 리소스 사용량 확인
kubectl get pods -o wide
kubectl top pods
문제 상황 시뮬레이션
프로세스는 실행 중이지만 애플리케이션이 준비되지 않았거나 응답하지 않는 상황을 가정해 봅니다:
# 웹 애플리케이션 테스트
kubectl run test-client --image=curlimages/curl --rm -it -- sh
# 컨테이너 내에서: curl http://web-app-service
가정한 상황의 문제점:
- 애플리케이션이 시작되지 않았는데 트래픽 수신: Pod는 Running 상태지만 실제로는 아직 준비되지 않음
- 장애가 발생해도 자동 복구 안됨: 애플리케이션이 hang 상태여도 Pod는 그대로 유지
- 시작 시간이 긴 애플리케이션 문제: 데이터베이스 연결 등으로 시작이 늦어지는 경우
- 좀비 프로세스 문제: 프로세스는 살아있지만 요청을 처리하지 못하는 상태
# 문제 상황 재현을 위한 테스트 Pod 생성
kubectl run slow-app --image=nginx --command -- sh -c "sleep 30 && nginx -g 'daemon off;'"
# Pod 상태 확인
kubectl get pod slow-app
# STATUS가 Running이지만 실제로는 nginx가 아직 시작되지 않음
# 서비스 연결 시도
kubectl expose pod slow-app --port=80
kubectl run test --image=curlimages/curl --rm -it --command -- curl http://slow-app
# 연결 실패 또는 타임아웃 발생
이런 상태를 구분하기 위해 Probe를 이용한 헬스체크와 자동 복구 설정을 살펴보겠습니다.
💡 Probe의 종류와 역할 이해하기
기본 개념 정리
쿠버네티스는 세 가지 종류의 Probe를 제공합니다:
Liveness Probe (생존 확인)
- 컨테이너가 정상적으로 동작하고 있는지 확인
- 연속 실패가
failureThreshold에 도달하면 해당 컨테이너를 종료하고 재시작 정책을 적용 - 용도: 데드락, 무한루프 등 복구 가능한 장애 감지
Readiness Probe (준비 상태 확인)
- Pod가 트래픽을 받을 준비가 되었는지 확인
- 실패 시 → Service 엔드포인트에서 제거
- 용도: 초기화 완료, 의존성 연결 확인
Startup Probe (시작 확인)
- 컨테이너가 시작을 완료했는지 확인
- 연속 실패가
failureThreshold에 도달하면 해당 컨테이너를 종료하고 재시작 정책을 적용 - 용도: 시작 시간이 긴 애플리케이션 처리
시작 → Startup Probe 성공 (설정한 경우)
├─ Readiness Probe: 트래픽을 받을 준비 확인
└─ Liveness Probe: 컨테이너 복구 필요 여부 확인
Startup이 성공한 뒤 Readiness와 Liveness는 각각 독립적으로 동작합니다. Readiness가 성공해야 Liveness가 시작되는 순서는 아닙니다. Probe 공식 문서
첫 번째 Probe 설정
간단한 HTTP Probe 구성 예제입니다:
# basic-probe-pod.yaml
apiVersion: v1
kind: Pod
metadata:
name: web-with-probes
spec:
containers:
- name: web
image: nginx
ports:
- containerPort: 80
# Readiness Probe: 트래픽 받을 준비 확인
readinessProbe:
httpGet:
path: /
port: 80
initialDelaySeconds: 5 # 5초 후 시작
periodSeconds: 10 # 10초마다 확인
# Liveness Probe: 정상 동작 확인
livenessProbe:
httpGet:
path: /
port: 80
initialDelaySeconds: 30 # 30초 후 시작
periodSeconds: 30 # 30초마다 확인
timeoutSeconds: 5 # 5초 타임아웃
failureThreshold: 3 # 3번 실패 시 재시작
# Pod 생성 및 상태 확인
kubectl apply -f basic-probe-pod.yaml
# Probe 상태 모니터링
watch kubectl describe pod web-with-probes
중요한 관찰 포인트:
Ready상태 변화 추적Conditions섹션에서PodReadyCondition확인Events에서 Probe 실행 결과 확인
🔍 Probe 타입별 상세 실습
1. HTTP GET Probe
가장 일반적으로 사용되는 HTTP 기반 헬스체크입니다:
# http-probe-app.yaml
apiVersion: v1
kind: Pod
metadata:
name: http-probe-app
spec:
containers:
- name: app
image: httpd:2.4
ports:
- containerPort: 80
readinessProbe:
httpGet:
path: /
port: 80
httpHeaders: # 필요시 커스텀 헤더 추가
- name: Custom-Header
value: Health-Check
initialDelaySeconds: 5
periodSeconds: 10
successThreshold: 1 # 1번 성공하면 Ready
failureThreshold: 3 # 3번 실패하면 Not Ready
livenessProbe:
httpGet:
path: / # 기본 이미지가 제공하는 경로
port: 80
initialDelaySeconds: 30
periodSeconds: 20
timeoutSeconds: 10
failureThreshold: 3
헬스체크 엔드포인트 테스트:
# Pod 생성
kubectl apply -f http-probe-app.yaml
# 수동으로 헬스체크 엔드포인트 테스트
kubectl port-forward pod/http-probe-app 8080:80
# 다른 터미널에서 확인 (종료할 때 port-forward 터미널에서 Ctrl+C)
curl http://127.0.0.1:8080/
2. TCP Socket Probe
HTTP가 아닌 TCP 서비스의 경우 사용합니다:
# tcp-probe-app.yaml
apiVersion: v1
kind: Pod
metadata:
name: tcp-probe-app
spec:
containers:
- name: redis
image: redis:6.2
ports:
- containerPort: 6379
readinessProbe:
tcpSocket:
port: 6379 # Redis 포트 연결 확인
initialDelaySeconds: 10
periodSeconds: 5
livenessProbe:
tcpSocket:
port: 6379
initialDelaySeconds: 30
periodSeconds: 10
failureThreshold: 3
# Redis Pod 생성 및 테스트
kubectl apply -f tcp-probe-app.yaml
# TCP 연결 수동 테스트
kubectl exec -it tcp-probe-app -- redis-cli ping
# PONG 응답 확인
3. Command Probe
커스텀 명령어를 실행하여 상태를 확인합니다:
# command-probe-app.yaml
apiVersion: v1
kind: Pod
metadata:
name: command-probe-app
spec:
containers:
- name: app
image: busybox
command: ["/bin/sh"]
args: ["-c", "while true; do sleep 30; done"]
readinessProbe:
exec:
command:
- cat
- /tmp/ready # 파일 존재 여부로 준비 상태 확인
initialDelaySeconds: 10
periodSeconds: 5
livenessProbe:
exec:
command:
- /bin/sh
- -c
- "[ $(ps aux | grep -v grep | grep sleep | wc -l) -gt 0 ]" # 프로세스 존재 확인
initialDelaySeconds: 30
periodSeconds: 15
# Pod 생성
kubectl apply -f command-probe-app.yaml
# Ready 상태로 만들기 위해 파일 생성
kubectl exec -it command-probe-app -- touch /tmp/ready
# Pod 상태 변화 확인
kubectl get pod command-probe-app
🚀 Startup Probe로 느린 시작 애플리케이션 처리
시작 시간이 긴 애플리케이션을 위한 Startup Probe 실습입니다:
# slow-startup-app.yaml
apiVersion: v1
kind: Pod
metadata:
name: slow-startup-app
spec:
containers:
- name: slow-app
image: nginx
command: ["/bin/sh"]
args: ["-c", "echo 'Starting up...' && sleep 60 && nginx -g 'daemon off;'"]
ports:
- containerPort: 80
# Startup Probe: 최초 30초 지연 후 주기 10초·연속 실패 기준 12회
startupProbe:
httpGet:
path: /
port: 80
initialDelaySeconds: 30
periodSeconds: 10
failureThreshold: 12 # 연속 12회 실패하면 컨테이너 종료
# Readiness Probe: Startup 완료 후 작동
readinessProbe:
httpGet:
path: /
port: 80
periodSeconds: 5
# Liveness Probe: Startup 완료 후 작동
livenessProbe:
httpGet:
path: /
port: 80
periodSeconds: 10
failureThreshold: 3
# 느린 시작 애플리케이션 생성
kubectl apply -f slow-startup-app.yaml
# 시작 과정 모니터링
watch kubectl describe pod slow-startup-app
중요한 동작 방식:
- Startup Probe가 성공할 때까지 Liveness/Readiness Probe는 비활성화
- Startup Probe가 실패 임계치에 도달하면 해당 컨테이너 종료 후 재시작 정책 적용
- Startup Probe 성공 후 일반적인 Probe 동작 시작
💡 실전 시나리오: 다양한 애플리케이션별 Probe 설정
아래는 설정 구조 예시입니다. Spring Boot JAR·Node.js 코드·헬스체크 경로가 포함된 애플리케이션 이미지를 직접 준비해야 합니다. 기본 런타임 이미지만으로 해당 파일이나 API가 생기지는 않습니다. 이 글의 이미지 태그는 원문에 있던 값이며, 현재 환경에서의 실행은 별도 확인이 필요합니다.
1. 웹 애플리케이션 (Spring Boot)
# spring-boot-app.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: spring-boot-app
spec:
replicas: 3
selector:
matchLabels:
app: spring-boot
template:
metadata:
labels:
app: spring-boot
spec:
containers:
- name: app
image: openjdk:11-jre-slim
command: ["java"]
args: ["-jar", "/app/spring-boot-app.jar"]
ports:
- containerPort: 8080
# 시작이 느린 Spring Boot 앱을 위한 Startup Probe
startupProbe:
httpGet:
path: /actuator/health
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
failureThreshold: 30 # 최대 5분 대기
# 의존성 준비 확인
readinessProbe:
httpGet:
path: /actuator/health/readiness
port: 8080
periodSeconds: 10
successThreshold: 1
failureThreshold: 3
# 애플리케이션 데드락 감지
livenessProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
periodSeconds: 20
timeoutSeconds: 5
failureThreshold: 3
resources:
requests:
memory: "512Mi"
cpu: "250m"
limits:
memory: "1Gi"
cpu: "500m"
2. 데이터베이스 (PostgreSQL)
# postgres-app.yaml
apiVersion: v1
kind: Pod
metadata:
name: postgres-db
spec:
containers:
- name: postgres
image: postgres:13
env:
- name: POSTGRES_PASSWORD
value: "mypassword"
- name: POSTGRES_DB
value: "mydb"
ports:
- containerPort: 5432
# 데이터베이스 초기화 완료 확인
startupProbe:
exec:
command:
- /bin/sh
- -c
- "pg_isready -U postgres -d mydb"
initialDelaySeconds: 30
periodSeconds: 10
failureThreshold: 30
# 연결 가능 상태 확인
readinessProbe:
exec:
command:
- /bin/sh
- -c
- "pg_isready -U postgres -d mydb"
periodSeconds: 10
# 데이터베이스 프로세스 확인
livenessProbe:
exec:
command:
- /bin/sh
- -c
- "pg_isready -U postgres"
periodSeconds: 30
failureThreshold: 3
3. 마이크로서비스 API
# microservice-api.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: user-api
spec:
replicas: 2
selector:
matchLabels:
app: user-api
template:
metadata:
labels:
app: user-api
spec:
containers:
- name: api
image: node:16-alpine
command: ["node"]
args: ["server.js"]
ports:
- containerPort: 3000
env:
- name: DB_HOST
value: "postgres-db"
# API 서버 시작 확인
startupProbe:
httpGet:
path: /health
port: 3000
initialDelaySeconds: 10
periodSeconds: 5
failureThreshold: 12
# 의존성 서비스 연결 확인
readinessProbe:
httpGet:
path: /health/ready # DB 연결 등 의존성 확인
port: 3000
periodSeconds: 10
successThreshold: 1
failureThreshold: 3
# API 응답 확인
livenessProbe:
httpGet:
path: /health/live # 기본적인 API 응답 확인
port: 3000
periodSeconds: 15
timeoutSeconds: 5
failureThreshold: 3
🚨 Probe 실패 상황과 문제 해결
일반적인 실패 시나리오 실습
1. Readiness Probe 실패 시뮬레이션
# readiness-fail-test.yaml
apiVersion: v1
kind: Pod
metadata:
name: readiness-fail-test
labels:
app: readiness-fail-test
spec:
containers:
- name: web
image: nginx
ports:
- containerPort: 80
readinessProbe:
httpGet:
path: /nonexistent # 존재하지 않는 경로
port: 80
periodSeconds: 5
failureThreshold: 2
# Pod 생성 및 Service 생성
kubectl apply -f readiness-fail-test.yaml
kubectl expose pod readiness-fail-test --port=80 --name=fail-test-service
# Service 엔드포인트 확인
kubectl get endpoints fail-test-service
# ENDPOINTS 컬럼이 비어있음을 확인
# Pod 상태 확인
kubectl describe pod readiness-fail-test
# Ready: False 상태 확인
2. Liveness Probe 실패 시뮬레이션
# liveness-fail-test.yaml
apiVersion: v1
kind: Pod
metadata:
name: liveness-fail-test
spec:
containers:
- name: app
image: busybox
command: ["/bin/sh"]
args: ["-c", "touch /tmp/healthy && sleep 30 && rm /tmp/healthy && sleep 3600"]
livenessProbe:
exec:
command:
- cat
- /tmp/healthy
initialDelaySeconds: 5
periodSeconds: 10
failureThreshold: 2
# Probe를 포함한 위 매니페스트로 처음부터 생성
kubectl apply -f liveness-fail-test.yaml
# 컨테이너 재시작 모니터링
watch kubectl get pod liveness-fail-test
# RESTARTS 컬럼 증가 확인
이미 실행 중인 단독 Pod의 Probe는 일반적인 edit·patch로 추가하거나 변경할 수 없습니다. 수정한 매니페스트로 Pod를 다시 만들거나, Deployment의 Pod 템플릿을 수정해 교체해야 합니다. Pod 변경 제약
디버깅과 문제 해결
Probe 실패 원인 분석:
# Pod 이벤트 확인
kubectl describe pod <pod-name>
# 상세한 이벤트 로그
kubectl get events --field-selector involvedObject.name=<pod-name> --sort-by=.metadata.creationTimestamp
# 컨테이너 로그 확인
kubectl logs <pod-name> -c <container-name>
# 이전 컨테이너 로그 (재시작된 경우)
kubectl logs <pod-name> -c <container-name> --previous
일반적인 문제와 해결책:
| 문제 | 원인 | 해결책 |
|---|---|---|
| Readiness 계속 실패 | 의존성 서비스 미준비 | initialDelaySeconds 증가, 의존성 확인 |
| Liveness 간헐적 실패 | 일시적 부하 | timeoutSeconds, failureThreshold 조정 |
| Startup 타임아웃 | 시작 시간 부족 | failureThreshold 증가, periodSeconds 조정 |
| 너무 잦은 재시작 | 너무 엄격한 설정 | 각 threshold와 timeout 값 완화 |
🎯 Probe 설정 최적화 가이드
권장 설정 값
다음은 시간·횟수만 발췌한 조정 예시입니다. 실제 적용 시 각 Probe에 httpGet·tcpSocket·exec 등의 검사 방식도 넣어야 하며, 시작 대기 시간에는 initialDelaySeconds도 함께 고려해야 합니다.
일반적인 웹 애플리케이션:
startupProbe:
initialDelaySeconds: 10
periodSeconds: 10
failureThreshold: 30 # 최대 5분
readinessProbe:
initialDelaySeconds: 0 # Startup 완료 후 즉시 시작
periodSeconds: 10
successThreshold: 1
failureThreshold: 3 # 30초 내 복구 기회
livenessProbe:
initialDelaySeconds: 0
periodSeconds: 30 # 자주 확인할 필요 없음
timeoutSeconds: 5
failureThreshold: 3 # 90초 내 복구 기회
데이터베이스:
startupProbe:
initialDelaySeconds: 30 # DB 초기화 시간 고려
periodSeconds: 10
failureThreshold: 60 # 최대 10분
readinessProbe:
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 3
livenessProbe:
periodSeconds: 60 # DB는 덜 자주 확인
timeoutSeconds: 10
failureThreshold: 3
성능 최적화 고려사항
# Probe 설정이 클러스터에 미치는 영향 분석
kubectl top nodes
kubectl top pods --all-namespaces
# 특정 Pod의 Probe 실행 빈도 계산
# Readiness: 10초마다, Liveness: 30초마다
# = 1분에 총 8번의 헬스체크 요청
리소스 사용량 최적화:
periodSeconds: 너무 자주 확인하지 않도록 조정timeoutSeconds: 네트워크 지연 고려failureThreshold: 일시적 장애 허용도 고려
🎯 CKA 대비 연습 패턴
연습할 문제 유형
1. 기존 Pod/Deployment에 Probe 추가
# 문제: nginx Deployment에 liveness probe 추가
kubectl edit deployment nginx-deployment
# 또는 patch 명령 사용
kubectl patch deployment nginx-deployment -p '{"spec":{"template":{"spec":{"containers":[{"name":"nginx","livenessProbe":{"httpGet":{"path":"/","port":80},"initialDelaySeconds":30,"periodSeconds":10}}]}}}}'
2. 문제가 있는 Probe 설정 수정
# Probe 실패 원인 찾기
kubectl describe pod <pod-name>
kubectl get events
# Deployment가 관리하는 Pod는 템플릿에서 Probe 수정
kubectl edit deployment <deployment-name>
# 단독 Pod는 파일에서 수정 후 다시 생성
3. 다양한 타입의 Probe 구성
# Probe 구성을 연습하는 패턴
readinessProbe:
httpGet: # 또는 tcpSocket, exec
path: /health
port: 8080
initialDelaySeconds: 5
periodSeconds: 10
livenessProbe:
httpGet:
path: /
port: 80
initialDelaySeconds: 30
periodSeconds: 30
4. 트러블슈팅 문제
- Pod가 Ready 상태가 되지 않는 이유 찾기
- Service에서 트래픽을 받지 못하는 Pod 문제 해결
- 계속 재시작되는 Pod 원인 분석
시간 단축 팁
# 빠른 Probe 설정이 포함된 Pod 생성
kubectl run nginx --image=nginx --port=80 \
--dry-run=client -o yaml > pod.yaml
# 편집기에서 spec.containers의 해당 컨테이너 아래에 Probe 추가
# 파일 끝에 단순히 덧붙이면 잘못된 위치에 들어갈 수 있음
vi pod.yaml
# Pod 생성
kubectl apply -f pod.yaml
자주 사용하는 Probe 명령어:
# Probe 상태 빠른 확인
kubectl get pods -o wide
kubectl describe pod <name> | grep -A 10 Conditions
# Probe 실패 이벤트 확인
kubectl get events --field-selector reason=Unhealthy
📊 모니터링과 알림 설정
Probe 상태 모니터링
# 클러스터 전체 Pod 상태 확인
kubectl get pods --all-namespaces -o wide
# Running이 아닌 Pod 필터링 (Ready=False와는 다른 조건)
kubectl get pods --all-namespaces --field-selector=status.phase!=Running
# Probe 실패 이벤트 모니터링
kubectl get events --all-namespaces --field-selector reason=Unhealthy -w
커스텀 헬스체크 엔드포인트 개발
실제 애플리케이션에서 사용할 수 있는 헬스체크 예제:
// Node.js Express 애플리케이션 예제
const express = require('express');
const app = express();
let isReady = false;
let isHealthy = true;
// 시작 시 의존성 초기화
setTimeout(() => {
// DB 연결, 캐시 워밍업 등
isReady = true;
}, 10000); // 10초 후 준비 완료
// Liveness Probe 엔드포인트
app.get('/health/live', (req, res) => {
if (isHealthy) {
res.status(200).send('OK');
} else {
res.status(503).send('Service Unavailable');
}
});
// Readiness Probe 엔드포인트
app.get('/health/ready', (req, res) => {
if (isReady && isHealthy) {
res.status(200).send('Ready');
} else {
res.status(503).send('Not Ready');
}
});
// Startup Probe 엔드포인트
app.get('/health/startup', (req, res) => {
if (isReady) {
res.status(200).send('Started');
} else {
res.status(503).send('Starting');
}
});
app.listen(3000);
📚 필수 명령어 정리
Probe 관련 핵심 명령어
# Pod 상태와 Probe 정보 확인
kubectl describe pod <pod-name>
kubectl get pod <pod-name> -o yaml
# Probe 실패 이벤트 확인
kubectl get events --field-selector involvedObject.name=<pod-name>
kubectl get events --field-selector reason=Unhealthy
# 실시간 Pod 상태 모니터링
watch kubectl get pods
kubectl get pods -w
# 특정 조건의 Pod 찾기
kubectl get pods --field-selector=status.phase=Pending
kubectl get pods -o jsonpath='{.items[?(@.status.containerStatuses[*].ready==false)].metadata.name}'
# Service 엔드포인트 확인
kubectl get endpoints <service-name>
kubectl describe service <service-name>
Probe 설정 템플릿
# 기본 HTTP Probe 템플릿
readinessProbe:
httpGet:
path: /health/ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 10
successThreshold: 1
failureThreshold: 3
livenessProbe:
httpGet:
path: /health/live
port: 8080
initialDelaySeconds: 30
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 3
---
# 별도 예제: TCP Probe 템플릿
readinessProbe:
tcpSocket:
port: 6379
initialDelaySeconds: 10
periodSeconds: 5
---
# 별도 예제: Command Probe 템플릿
livenessProbe:
exec:
command:
- /bin/sh
- -c
- "pg_isready -U postgres"
periodSeconds: 30
failureThreshold: 3
🔄 고급 Probe 활용 패턴
1. 다단계 헬스체크
복잡한 애플리케이션에서는 여러 단계의 헬스체크를 구성할 수 있습니다:
아래 /health/* 경로는 직접 구현해야 합니다. 기본 nginx 이미지에는 없으므로, 이 파일은 애플리케이션에 맞게 바꾸는 구성 예시입니다.
# multi-tier-health-check.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: multi-tier-app
spec:
replicas: 2
selector:
matchLabels:
app: multi-tier-app
template:
metadata:
labels:
app: multi-tier-app
spec:
containers:
- name: app
image: nginx # 실제로는 복잡한 애플리케이션
ports:
- containerPort: 80
# 1단계: 기본 서비스 시작 확인
startupProbe:
httpGet:
path: /health/startup
port: 80
initialDelaySeconds: 10
periodSeconds: 5
failureThreshold: 24 # 2분 대기
# 2단계: 의존성과 캐시 준비 확인
readinessProbe:
httpGet:
path: /health/ready
port: 80
httpHeaders:
- name: X-Health-Check
value: "readiness"
periodSeconds: 10
successThreshold: 2 # 2번 연속 성공해야 Ready
failureThreshold: 3
# 3단계: 지속적인 건강 상태 모니터링
livenessProbe:
httpGet:
path: /health/live
port: 80
httpHeaders:
- name: X-Health-Check
value: "liveness"
periodSeconds: 30
timeoutSeconds: 10
failureThreshold: 3
2. 조건부 Probe (Init Container 활용)
의존성 서비스가 준비된 후에만 메인 컨테이너를 시작하는 패턴:
# conditional-startup.yaml
apiVersion: v1
kind: Pod
metadata:
name: conditional-startup-app
spec:
initContainers:
# 의존성 서비스 대기
- name: wait-for-db
image: busybox
command:
- /bin/sh
- -c
- |
echo "Waiting for database..."
until nc -z postgres-service 5432; do
echo "Database not ready, sleeping..."
sleep 2
done
echo "Database is ready!"
- name: wait-for-cache
image: busybox
command:
- /bin/sh
- -c
- |
echo "Waiting for cache..."
until nc -z redis-service 6379; do
echo "Cache not ready, sleeping..."
sleep 2
done
echo "Cache is ready!"
containers:
- name: app
image: myapp:latest
ports:
- containerPort: 8080
# 의존성이 모두 준비된 상태에서 시작하므로 빠른 Probe 가능
readinessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 30
periodSeconds: 20
3. 연속 성공 후 트래픽에 참여하는 패턴
successThreshold: 3은 세 번 연속 성공한 뒤 Ready로 전환하는 조건입니다. 트래픽 비율을 조금씩 늘리는 기능은 아니며, 가중치 조정은 별도의 라우팅 구성이 필요합니다.
# gradual-traffic-increase.yaml
apiVersion: v1
kind: Service
metadata:
name: gradual-service
spec:
selector:
app: gradual-app
ports:
- port: 80
targetPort: 8080
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: gradual-app
spec:
replicas: 3
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 1
maxSurge: 1
selector:
matchLabels:
app: gradual-app
template:
metadata:
labels:
app: gradual-app
spec:
containers:
- name: app
image: myapp:latest
ports:
- containerPort: 8080
readinessProbe:
httpGet:
path: /health/ready
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
successThreshold: 3 # 3번 연속 성공 후 트래픽 수신
failureThreshold: 2 # 빠른 트래픽 차단
livenessProbe:
httpGet:
path: /health/live
port: 8080
initialDelaySeconds: 60
periodSeconds: 30
failureThreshold: 5 # 관대한 실패 허용
🛠️ 실전 문제 해결 시나리오
아래 장애와 로그는 진단 방법을 설명하기 위한 가정입니다. 실제 장애를 관찰하거나 해결한 결과로 제시하는 것은 아닙니다.
시나리오 1: 간헐적 Readiness Probe 실패
문제 상황:
# 서비스가 간헐적으로 응답하지 않음
kubectl get pods -l app=unstable-app
# 일부 Pod가 Ready 상태와 Not Ready 상태를 반복
문제 분석:
# 이벤트 확인
kubectl get events --field-selector involvedObject.name=unstable-app-xxx
# 로그 분석
kubectl logs unstable-app-xxx --tail=100
# 리소스 사용량 확인
kubectl top pod unstable-app-xxx
해결 방법:
# 더 관대한 Readiness Probe 설정
readinessProbe:
httpGet:
path: /health
port: 8080
periodSeconds: 10
timeoutSeconds: 10 # 타임아웃 증가
successThreshold: 2 # 2번 연속 성공 필요
failureThreshold: 5 # 5번 실패 허용
시나리오 2: 메모리 누수로 인한 반복적 재시작
문제 분석:
# 재시작 횟수 확인
kubectl get pods -l app=memory-leak-app
# RESTARTS 컬럼이 계속 증가
# 메모리 사용량 트렌드 확인
kubectl top pods -l app=memory-leak-app --containers
# 상세 이벤트 확인
kubectl describe pod memory-leak-app-xxx
# Reason: OOMKilled 이벤트 확인
임시 해결책:
메모리 제한 상향은 누수 수정을 대신하지 않습니다. Liveness 간격을 늘려도 커널의 OOM 종료를 막을 수는 없습니다.
# 리소스 제한 증가 (근본 해결은 아님)
resources:
limits:
memory: "2Gi" # 기존 1Gi에서 증가
requests:
memory: "512Mi"
# Liveness Probe 간격 조정
livenessProbe:
httpGet:
path: /health
port: 8080
periodSeconds: 60 # 더 긴 간격으로 설정
failureThreshold: 5 # 더 관대하게 설정
시나리오 3: 데이터베이스 연결 풀 고갈
문제 상황:
# 애플리케이션이 간헐적으로 응답하지 않음
kubectl logs app-pod | grep "connection"
# "Connection pool exhausted" 에러 발견
헬스체크 개선 (연결 풀 고갈 원인 수정은 별도):
# 애플리케이션별 헬스체크 엔드포인트 개선
readinessProbe:
httpGet:
path: /health/db-connection # DB 연결 상태 전용 체크
port: 8080
periodSeconds: 15 # 체크 빈도 조정
timeoutSeconds: 5
failureThreshold: 2
livenessProbe:
httpGet:
path: /health/basic # 기본적인 프로세스 체크만
port: 8080
periodSeconds: 45 # DB에 부담을 주지 않도록
failureThreshold: 3
🎭 고급 패턴: Circuit Breaker와 연동
Circuit Breaker 상태를 반영한 Probe
의존 서비스 장애는 Readiness 판단에 반영하고, Liveness는 애플리케이션 자체의 복구 불가능한 상태를 확인하도록 설계합니다. 외부 장애만으로 모든 복제본이 반복 재시작되지 않도록 구분해야 합니다.
# circuit-breaker-aware.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: resilient-app
spec:
replicas: 3
selector:
matchLabels:
app: resilient-app
template:
metadata:
labels:
app: resilient-app
spec:
containers:
- name: app
image: resilient-app:latest
ports:
- containerPort: 8080
env:
- name: CIRCUIT_BREAKER_ENABLED
value: "true"
readinessProbe:
httpGet:
path: /health/ready
port: 8080
httpHeaders:
- name: X-Check-Dependencies
value: "true"
periodSeconds: 10
successThreshold: 1
failureThreshold: 3
livenessProbe:
httpGet:
path: /health/live
port: 8080
periodSeconds: 30
failureThreshold: 5
📈 성능 최적화와 베스트 프랙티스
1. Probe 오버헤드 최소화
# 클러스터 전체 Probe 실행 빈도 계산
kubectl get pods --all-namespaces -o jsonpath='{range .items[*]}{.metadata.name}{" "}{.spec.containers[*].readinessProbe.periodSeconds}{" "}{.spec.containers[*].livenessProbe.periodSeconds}{"\n"}{end}' | head -20
# 높은 빈도의 Probe 식별
kubectl get pods --all-namespaces -o yaml | grep -A 3 -B 3 "periodSeconds: [1-5]"
2. 효율적인 헬스체크 엔드포인트 설계
# Python Flask 예제 - 효율적인 헬스체크
from flask import Flask, jsonify
import time
import threading
app = Flask(__name__)
# 캐시된 상태 정보
health_cache = {
'last_check': 0,
'status': 'unknown',
'cache_duration': 30 # 30초 캐시
}
def check_dependencies():
"""실제 의존성 체크 (무거운 작업)"""
# DB 연결, 외부 API 호출 등
return True
@app.route('/health/live')
def liveness():
"""가벼운 프로세스 체크만"""
return jsonify({'status': 'alive'}), 200
@app.route('/health/ready')
def readiness():
"""캐시를 활용한 효율적인 의존성 체크"""
current_time = time.time()
# 캐시가 유효한 경우
if current_time - health_cache['last_check'] < health_cache['cache_duration']:
if health_cache['status'] == 'ready':
return jsonify({'status': 'ready'}), 200
else:
return jsonify({'status': 'not ready'}), 503
# 캐시 갱신 필요
try:
if check_dependencies():
health_cache['status'] = 'ready'
health_cache['last_check'] = current_time
return jsonify({'status': 'ready'}), 200
else:
health_cache['status'] = 'not ready'
health_cache['last_check'] = current_time
return jsonify({'status': 'not ready'}), 503
except Exception as e:
health_cache['status'] = 'error'
health_cache['last_check'] = current_time
return jsonify({'status': 'error', 'message': str(e)}), 503
3. 환경별 Probe 설정 관리
아래 ConfigMap만 생성한다고 Pod의 Probe 설정이 바뀌지는 않습니다. Helm·Kustomize 또는 별도 생성 단계에서 값을 Pod 템플릿에 반영해야 합니다.
# ConfigMap으로 환경별 설정 관리
apiVersion: v1
kind: ConfigMap
metadata:
name: probe-config
data:
development.yaml: |
probes:
readiness:
initialDelaySeconds: 5
periodSeconds: 5
failureThreshold: 3
liveness:
initialDelaySeconds: 30
periodSeconds: 30
failureThreshold: 3
production.yaml: |
probes:
readiness:
initialDelaySeconds: 10
periodSeconds: 10
failureThreshold: 5
liveness:
initialDelaySeconds: 60
periodSeconds: 60
failureThreshold: 5
🚀 다음 학습 계획
다음에는 여러 애플리케이션과 팀을 위한 네임스페이스 격리와 리소스 관리 방법을 살펴볼 계획입니다.
다음 주제들
- 네임스페이스를 통한 리소스 격리와 다중 테넌시 - 팀별, 환경별 격리
- Ingress를 통한 외부 트래픽 관리와 라우팅
- 스토리지와 PersistentVolume 관리
- RBAC과 보안 설정
- 모니터링과 로깅 시스템 구축
개인 학습 목표
- 다양한 애플리케이션 유형별 최적 Probe 패턴 숙달
- 대규모 클러스터에서의 효율적인 헬스체크 설계 경험
- 장애 상황에서의 빠른 진단과 복구 능력 향상
- 모니터링과 알림을 통한 proactive한 운영 체계 구축
Probe를 구성할 때는 시작 완료, 요청 처리 준비, 재시작이 필요한 상태를 구분하는 것이 핵심입니다. 각 검사 경로와 실패 조건은 실제 애플리케이션 동작에 맞춰 구현하고 검증해야 합니다.
다음 글에서는 팀별·환경별 리소스를 나누는 네임스페이스 활용 방법을 다룹니다.
🔗 참고 자료:
- Kubernetes 공식 문서 - Configure Liveness, Readiness and Startup Probes
- Kubernetes 공식 문서 - Pod Lifecycle
- Kubernetes Best Practices - Health Checks
yellow.log · 2025.08.11에 작성 · 2026.09.29에 수정
CKA 준비 전체 8편
- 1.Pod 생성과 관리 - 쿠버네티스의 기본 단위 이해하기
- 2.ReplicaSet과 Deployment로 Pod 관리하기 - 확장성과 가용성 확보
- 3.Service를 통한 Pod 네트워킹 - 안정적인 접근 경로 만들기
- 4.ConfigMap과 Secret으로 설정 관리하기 - 환경별 설정 분리
- 5.Pod 리소스 제한과 요청 설정 - 안정적인 클러스터 운영을 위한 리소스 관리
- 6.Liveness와 Readiness Probe 구성 - 헬스체크와 자동 복구읽는 중
- 7.네임스페이스를 통한 리소스 격리와 다중 테넌시 - 팀별, 환경별 격리
- 8.Ingress를 통한 외부 트래픽 관리와 라우팅 - 클러스터의 똑똑한 관문 구축