Senior Site Reliability Engineer
Design and implement scalable, secure, and highly available distributed systems for ClickHouse Cloud. Establish SLOs/SLAs, manage monitoring and alerting, lead incident response and post-mortems, drive chaos engineering initiatives, and build automation tools to enhance reliability and performance. Collaborate across engineering teams to improve the resilience of cloud infrastructure components including Dataplane, Control Plane, and ClickHouse Core.