Cloudlinux

Senior Database Reliability Engineer (DBRE) & Architect (worldwide remote)

Posted: 1 hours ago

Job Description

CloudLinux is transforming the Linux infrastructure market by ensuring security and stability for over 500,000 servers worldwide. Our products - CloudLinux OS, TuxCare, and Imunify360 - are the de facto standard in the hosting industry and Enterprise segment.We are seeking a visionary engineer to lead the evolution of our data platform. In 2025, we are shifting from classic database administration to an Internal Database-as-a-Service (DBaaS) model. We need a specialist who doesn't just "configure backups," but designs resilient distributed systems, writes code to automate infrastructure, and transforms databases into a reliable service for product teams.If you are tired of endless tickets and want to build platforms capable of processing petabytes of data, this role is for you.Your Challenges & Responsibilities:DBaaS Architecture: Design and implement a self-service platform based on Terraform and Ansible, enabling the deployment of HA clusters (PostgreSQL and ClickHouse, MongoDB, Redis) in a heterogeneous environment (Bare Metal + OpenNebula + Kubernetes + Public Clouds). You will turn infrastructure into a productScaling ClickHouse: Manage exponentially growing analytics clusters (12+ clusters, tens of terabytes of data). You will tackle sharding, table engine optimization (ReplicatedMergeTree), and building reliable S3 backup pipelines under high loadData Platform & Analytics Support: Maintain and scale the infrastructure for Apache Airflow and Redash. You will ensure the reliability of ETL pipelines and visualization tools, bridging the gap between raw infrastructure and the data analytics teamReliability as Code: Implement SRE practices in data management. Replace manual incident response with automated self-healing mechanisms. Define and implement SLO/SLI for all databasesStack Modernization: Lead the migration process from legacy solutions to modern cloud patterns. Participate in decision-making regarding the implementation of Kubernetes operators for stateful workloadsExpertise & Mentorship: Serve as the technical authority for product teams, helping them optimize data schemas and SQL queries for high-load systemsOur Tech Stack:Databases: PostgreSQL 15+ (Patroni, PgBouncer), ClickHouse (Sharded/Replicated), MongoDB, Redis, KafkaData & Analytics: Apache Airflow, Redash (Infrastructure & Integration)Infrastructure: Own 3+DC colocation (OpenNebula, Kubernetes, Bare Metal), AWS, Google Cloud, Azure, DO - Hybrid CloudAutomation & IaC: Terraform, Ansible, Python/Go, GitLab, Jenkins, GerritObservability: VictoriaMetrics, Grafana, LokiWhy CloudLinux?Culture: A Remote-first company with an "Employees First" principle. We value results, not hours in the officeImpact: Your architectural decisions will determine the stability of services used by thousands of companies around the worldGrowth: We support professional development and pay for training and conferencesRequirementsWhat We Expect From You:Deep PostgreSQL Expertise (5+ years): You know MVCC internals, understand locking mechanics, can configure Patroni and PgBouncer "with your eyes closed," and have experience with seamless major version upgrades under loadClickHouse Mastery: Experience operating large clusters, understanding ZooKeeper/ClickHouse Keeper, sharding, replication internals, and the ability to diagnose performance issues at the data-part levelEngineering Mindset (SRE/DevOps): You hate doing the same task twice by hand. Experience writing complex Terraform modules and Ansible roles is mandatory. Programming skills in Python or Go for automation are a huge plusHybrid Environment Experience: You understand the differences between running DBs on Bare Metal vs. Kubernetes vs. Cloud and know how to optimize TCO and disk subsystem performance (NVMe, Network Storage)Systems Approach: You see the big picture - from the network packet to the application business logic. You understand the importance of security (FIPS, Audit logs) and Disaster RecoveryNice to Have:Experience building an Internal Developer Platform (IDP)Experience operating databases in Kubernetes (CloudNativePG, Altinity Operator)Experience working in Cloud and Hosting providers on similar servicesBenefitsWhat's in it for you?A focus on professional developmentInteresting and challenging projectsFully remote work with flexible working hours, which allows you to schedule your day and work from any location worldwidePaid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leavesCompensation for private medical insuranceCo-working and gym/sports reimbursementBudget for educationThe opportunity to receive a reward for the most innovative idea that the company can patentBy applying for this position, you agree with CloudLinux Privacy Policy and give us your consent to maintain and process your personal data with this respect. Please read our Privacy Policy for more information.

Job Application Tips

  • Tailor your resume to highlight relevant experience for this position
  • Write a compelling cover letter that addresses the specific requirements
  • Research the company culture and values before applying
  • Prepare examples of your work that demonstrate your skills
  • Follow up on your application after a reasonable time period

You May Also Be Interested In