월 1,200명이 하루 평균 3만, 최대 10만 건의 query를 던지는 전사 분석 환경을 Trino와 Iceberg 중심으로 다시 세운 여정이다. 고정 cluster의 전면 장애와 수동 복구에서 출발해, 탄력성·고가용성·heavy query 방어·대용량 log 최적화까지 단계별로 어떻게 풀었는지 따라간다.
핵심 포인트- Trino를 Kubernetes로 이관해 KEDA cron 선제 scaling과 Prometheus 기반 후행 scaling, preStop graceful shutdown을 결합하고, coordinator 병목은 Trino Gateway가 active query 수로 BI·OLAP cluster에 분산하며 core time만 blue·green 이중화로 매일 교대
- Connector metadata 확장으로 partition 조건 없는 full scan을 실행 전 차단하고 resource group·session property로 동시 실행·memory·scan을 제한, upgrade는 staging query pump·replay와 단일 cluster canary 1주 관찰 후 전체 배포로 안전하게 검증
- 하루 약 15억 건·0.5TB app log를 days transform·bucket·event type partition으로 재설계하고, adaptive execution으로도 안 풀린 hot-key skew는 device ID hash repartition으로 분산해 batch 시간을 최대 90% 단축
- Kafka Connect Iceberg sink로 5분 commit exactly-once 준실시간 적재, worker의 90% 이상을 spot·Graviton으로 운영하며 interruption은 graceful shutdown·fault tolerance로 복구
왜 읽나Trino·Iceberg lakehouse를 실운영 규모로 굴리며 탄력성·안정성·비용을 동시에 잡아야 하는 데이터 플랫폼 엔지니어에게 구체적 튜닝 결정들의 종합판.