Introduce myself
Hermaeus Mora ·
플랫폼 엔지니어 (2025.01 – Present)
ML 모델 서빙 플랫폼 및 Kubernetes 클러스터 구축·운영 총괄
-
컴퓨팅 비용 76% 절감 — API 서버 VM 20+대를 워커 노드 5+대로 통합하고 유휴 리소스를 집약. 웹 서버 VM 10+대를 CDN으로 마이그레이션 (8k)
-
모델 학습 비용 50% 절감 — SageMaker HyperPod에 Spot + Karpenter 적용, HyperPod·KEDA·Karpenter로 scale-to-zero 추론 구현. (월 GPU 비용 3.6k)

-
장애 대응 시간 84% 단축 — LGTM(Loki·Grafana·Tempo·Mimir) + Alloy 풀스택 관측성 표준화, Mimir·Loki Ruler → Alertmanager → Slack 24/7 온콜 체계 구축. 노드/프로세스 장애 시 self-heal로 MTTR 30분 → 5분.
-
배포 시간 88% 단축, 파이프라인 100% 자동화 — 클러스터 워크로드는 GitLab Runner·ArgoCD·Image Updater·Reloader·ExternalDNS로, 엣지 워크로드는 Ansible·Teleport·Vault Agent로. 배포·롤백 40분 → 5분.
-
단일 NCP 계정 → AWS 멀티 어카운트 구조로 마이그레이션 — 거버넌스·컴플라이언스 표준화, 장애 전파 범위 최소화, GPU 쿼터 상향, 비용 가시성 확보.
-
전 구간 정적 시크릿 제거 — Vault 동적 자격증명(DB/RabbitMQ), AWS Secrets Manager + External Secrets 자동 갱신, EKS Pod Identity, IAM Identity Center 기반 멀티 계정 접근 일원화, Keycloak + Istio로 마이크로서비스 간 인증 중앙화. 하드코딩·장기 크리덴셜 0건.
-
VPN을 Teleport DB/노드 프록시로 대체 — 사용자별 DB 로그 감사.
-
Istio ztunnel mTLS — 클러스터 내부 통신 암호화.
-
기존 프로세스를 Kafka(Strimzi) 기반 이벤트 드리븐 방식으로 변경. Outbox(CDC) 및 Saga(Choreography) 패턴을 도입하여 데이터 일관성 향상.
-
Redis Sentinel·CloudNativePG·Strimzi·RabbitMQ 스테이트풀 워크로드 가용성 99.9% 달성.
백엔드 (2022.03 – 2024.04)
DAU 200만+ 미디어 스트리밍 서비스 운영
- AWS API Gateway 엔드포인트별 Elasticache 캐싱 최적화로 평균 응답 속도 63%+ 향상
- AWS Lambda 동시성 프로비저닝으로 피크 타임 에러율 27%+ 감소
백엔드 (2019.9 – 2022.03)
공공기관용 온프레미스 메신저 운영 및 개발
- 조직도 기능을 메신저에 통합하여 대전/광주/부산/강원교육청 등에 납품
- 10+ 고객사 베어메탈 환경에서 메신저 서버 및 DB, Redis, Rabbitmq 운영
Platform Engineer (2025.01 – Present)
Owned design and operation of the ML serving platform and Kubernetes clusters
- Cut compute spend 76% — consolidated 20+ API server VMs into 5+ worker nodes to reclaim idle capacity, and migrated 10+ web server VMs to CDN (8k/mo).
- Cut model training cost 50% — applied Spot + Karpenter on SageMaker HyperPod, and implemented scale-to-zero inference with HyperPod, KEDA, and Karpenter (GPU spend 3.6k/mo).
- Cut incident response time 84% — standardized full-stack observability on LGTM (Loki, Grafana, Tempo, Mimir) + Alloy, with Mimir/Loki Ruler → Alertmanager → Slack for 24/7 on-call. Node and process failures now self-heal: MTTR 30min → 5min.
- Cut deployment time 88% with a fully automated pipeline — GitLab Runner, ArgoCD, Image Updater, Reloader, and ExternalDNS for cluster workloads; Ansible, Teleport, and Vault Agent for edge. Deploy/rollback 40min → 5min.
- Sustained 99.9% availability on stateful workloads (Redis Sentinel, CloudNativePG, Strimzi, RabbitMQ).
- Migrated from a single NCP account to a multi-account AWS architecture — standardizing governance and compliance, containing blast radius, unlocking higher GPU quotas, and gaining per-account cost visibility.
- Eliminated static secrets platform-wide — Vault dynamic credentials (DB/RabbitMQ), AWS Secrets Manager + External Secrets for automatic rotation, EKS Pod Identity, unified multi-account access via IAM Identity Center, and centralized service-to-service auth with Keycloak + Istio. Zero hardcoded or long-lived credentials.
- Replaced VPN with Teleport DB/node proxies, enabling per-user database audit logs.
- Encrypted all intra-cluster traffic with Istio ztunnel mTLS.
- Re-architected batch processes into a Kafka (Strimzi) event-driven model, applying Outbox (CDC) and Saga (choreography) patterns to guarantee consistency across services.
Backend Developer (2022.03 – 2024.04)
Operated a media streaming service with 2M+ DAU
- Improved average response time 63%+ through per-endpoint ElastiCache caching optimization on AWS API Gateway.
- Reduced peak-time error rate 27%+ via AWS Lambda provisioned concurrency.