← WorkMundi · 1M+ jobs from around the world, liveSign inCreate free account

Senior Principal Site Reliability Engineer

bybit · Hong Kong SAR

📅 20/08/2026
🔔 Alert me about jobs like this
No password, no sign-up. Just the email — and you can leave the list anytime.
🔓 Apply — free →
Opens this job on WorkMundi. The account is free and takes under a minute.

See the other 7,002 jobs in Hong Kong →

🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
About Us Established in 2018, Bybit is one of the world’s leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3 — connecting users to the future of digital finance. Our core values define how we build. We listen, care and improve to create products and experiences that put users first. Backed by a global team of ambitious builders, problem-solvers, and innovators, we foster a high-performance and fast-moving environment where talent is empowered to drive real impact at the global scale. Supported by 24/7 multilingual customer service and a strong commitment to innovation, we are shaping the future of finance through technology, collaboration, and bold execution. Today, Bybit is recognized as one of the most trusted and transparent platforms in the digital asset industry, continuing to expand its global presence while building the infrastructure for the next generation of financial services. Core Responsibilities Chaos Engineering Platform Architecture & Development (50%) Design and build an enterprise-grade chaos engineering platform supporting multi-cluster (K8s + EC2 hybrid), multi-region, and multi-environment (testnet/mainnet) deployments Core capability development: Fault Injection Engine: Pod-level / Node-level / AZ-level fault simulation, network latency / packet loss / partition, dependency timeout / error injection Production Safety Assurance: Blast radius control, one-click Kill Switch, automatic rollback, real-time impact monitoring Traffic Isolation: Experiment traffic tagging and isolation to ensure fault injection does not impact real users Fault Isolation: Precise impact scoping at service / cluster / AZ granularity Design experiment orchestration capabilities supporting complex fault scenario composition (e.g., simultaneous network latency + downstream timeout + cache invalidation) Deep integration with existing monitoring, alerting, and SLO systems to achieve an automated closed loop: inject fault → observe impact → determine pass/fail Production Resilience Validation Framework (30%) Define safety standards and approval workflows for mainnet fault injection Design and drive routine chaos experiments: Daily patrol-level experiments: Low-risk experiments executed automatically on a daily/weekly basis Periodic validation experiments: Monthly/quarterly resilience verification of critical paths Large-scale drills: Cross-AZ / cross-region disaster recovery failover validation Establish a resilience scoring system to quantify system health based on experiment results Deliver improvement recommendations and drive business teams to remediate identified weaknesses 3. Technology Selection & Team Enablement (20%) Evaluate and select the technology foundation (Chaos Mesh / Litmus / custom components — hybrid strategy) Develop chaos engineering best practices and playbooks to enable SRE teams and application developers Mentor and grow the team (2–3 engineers) in chaos engineering capabilities Stay current with industry developments and introduce cutting-edge practices (e.g., AI-driven fault scenario discovery) ────── Requirements Must-Have: 8+ years of backend / infrastructure engineering experience, with 3+ years dedicated to chaos engineering or stability engineering Hands-on experience with large-scale fault injection in production environments (not just test environments), with deep understanding of production safety constraints Expert-level proficiency in Kubernetes fault injection (Chaos Mesh / Litmus / custom solutions), familiar with CRD / Operator development Proficient in at least one backend language (Go preferred), with platform-level system architecture design capability Deep understanding of distributed system failure modes (network partitions, split-brain, cascading failures, data inconsistency, etc.) Familiarity with observability tech stack (Prometheus / Grafana / Thanos / OpenTelemetry) Excellent technical documentation and solution design skills Nice-to-Have: Experience in financial / trading system stability (understanding of transaction consistency and fund safety constraints) Experience building SLO / Error Budget frameworks Experience building automated fault recovery (self-healing) systems Familiarity with AWS infrastructure (EC2 / EKS / Multi-AZ / Multi-Region) Knowledge of Netflix Chaos Engineering / AWS Fault Injection Simulator / Gremlin Open-source community contributions (Chaos Mesh / Litmus or similar projects) Soft Skills: Ability to balance "safety" and "validation depth" — not afraid of production injection, while maintaining strict risk control Strong cross-team collaboration and influence — chaos engineering requires buy-in from business teams; this role demands persuasion skills Self-driven, capable of independently planning and executing in ambiguous situations Why Join Us At Bybit, we are committed to fostering a supportive and enriching work environment. Our benefits include: - Study Growth Fund: We support your professional development and continuous learning. - Internal Events: Participate in regular team-building activities, workshops, and events designed to promote collaboration and innovation. - Global Collaboration: Be part of a diverse, international team, working alongside colleagues from around the world. - Career Advancement: Access opportunities for growth and advancement within a rapidly expanding global company. - Internal Mobility: Grow with us- Your long-term development is important to us. We offer internal job opportunities to help build your career path.
Read the rest of the job →
For people searching Engineer

144,883 engineer jobs are open right now

Here's how to pick the right one and stand out in your application.

144.883Jobs
31.687IN
81%EN

That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.

Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.

Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.

When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.

👁 21 have read this
0 comments
Want to comment?

Leave your e-mail to comment, react and follow the posts for your role. It is free.

Similar jobs

Job on WorkMundi — the world's largest job board. See more jobs from every continent, updated live.

📢
🎁

Before you apply, rehearse this interview.

Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card.

I want my training →