← WorkMundi · 1M+ jobs from around the world, liveSign inCreate free account

Senior Site Reliability Engineer

nextventures · Kuala Lumpur, Wilayah Persekutuan Kuala Lumpur, Malaysia

📅 10/08/2026
🔔 Alert me about jobs like this
No password, no sign-up. Just the email — and you can leave the list anytime.
🔓 Apply — free →
Opens this job on WorkMundi. The account is free and takes under a minute.

See the other 3,787 jobs in Malaysia →

🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Who We Are NEXT Ventures is a global fintech group powering FundedNext — one of the world's fastest-growing proprietary trading platforms — and FNmarkets, a regulated CFD brokerage. Across offices in Bangladesh, Malaysia, Sri Lanka, Cyprus, and Dubai, we build and operate the technology that lets traders access global markets at scale. Our Platform Engineering team owns the infrastructure, reliability, and observability backbone that every product squad depends on. Your Role in Our Mission As our Site Reliability Engineer, you are the dedicated specialist who keeps our services observable, fast, and resilient. You own the centralized logging and alerting backbone, drive service-level optimization across the stack, and perform log analysis across both Linux and Windows environments. Working within the Platform Engineering squad, you execute the reliability initiatives that free the squad lead to focus on architecture — and you are the reason incidents are short, signals are clean, and detection is fast. This role is distinct from our DevSecOps Engineer: DevSecOps builds and secures the platform foundation; you measure, detect, and optimize on top of it. How You'll Make an Impact Centralized Logging & Alerting Own and operate the centralized log management platform — ingestion, parsing, structured logging standards, and retention across all services. Build and tune alerting with tiered thresholds — catching real problems early while minimizing noise and alert fatigue. Perform log analysis across Linux and Windows systems to diagnose incidents and surface root causes. Drive MTTD under 15 minutes through better signals, dashboards, and runbooks. Service-Level Optimization & Reliability Identify, diagnose, and optimize service latency and inefficiency across edge, application, and backend layers — profile before guessing, measure every fix. Define, implement, and own SLOs, SLIs, and error budgets for critical customer-facing services, and drive improvements against them. Build deep observability with Datadog — APM, dashboards, monitors, log management, and SLO tracking. Lead reliability and performance root-cause analysis and drive durable fixes. Support load testing and capacity planning — identify breaking points before traffic growth causes production issues. Operations & Cross-Team Support Participate in a shared on-call rotation with solid runbooks and blameless post-incident reviews. Continuously reduce manual toil through automation and better tooling. Collaborate with product squads to instrument services, define meaningful SLIs, and surface the right signals. Document runbooks, dashboards, and operational procedures so any engineer can respond to incidents with clear guidance. What You Bring 5–7 years of professional engineering experience, with at least 3 years in SRE, Platform Engineering, or strongly reliability-focused DevOps work. Strong hands-on experience with centralized log management platforms — ingestion, parsing, structured logging, and retention using ELK/OpenSearch, Datadog Logs, Loki, or similar. Able to diagnose incidents through log analysis across both Linux and Windows environments, isolating root causes under pressure. Designs alerting systems with tiered thresholds that minimize noise while catching real problems early, with clear escalation paths. Proficient with Datadog across APM, dashboards, monitors, log management, and SLO tracking. Experienced defining and operating SLOs, SLIs, and error budgets for customer-facing services. Can isolate and resolve latency and inefficiency across edge, application, and backend layers — profiles before guessing, measures every fix. Comfortable scripting in Python, Bash, or Go to automate alerting, diagnostics, and toil reduction. Hands-on experience operating services on Kubernetes, EKS preferred. Familiar with Infrastructure-as-Code tooling such as Terraform for collaboration with the DevSecOps team. Evidence-driven — profiles and measures before guessing; validates every optimization against before/after data. Reliability-oriented — treats detection speed and signal quality as first-class engineering problems. Good communicator — works with product squads to define SLIs and explains reliability constraints clearly. X-Factor: AI-Native Engineering You actively use modern AI agentic workflows daily — not limited to Copilot autocomplete. You are proficient with Claude Code, Cursor, Windsurf, or equivalent tools. You are comfortable with project-level AI configuration ( CLAUDE.md , rules files), agentic task delegation, and AI-driven code review. You think in terms of 5–10x productivity through AI-augmented development — and you can demonstrate it. Your Journey After Applying Stage 1 — TA Interview Stage 2 — Screening Questionnaire Stage 3 — Hiring Manager Interview Stage 4 — Head of IT Interview Why Join NEXT Work on reliability challenges at real scale — 100M+ row data stores, multi-region traffic, high-frequency trading infrastructure. A team that treats observability and detection speed as first-class engineering problems, not afterthoughts. Flat structure — your work directly shapes how the platform operates, not filtered through layers of process. Offices across Bangladesh, Malaysia, Sri Lanka, Cyprus, and Dubai — a genuinely global engineering team. Competitive compensation benchmarked to your market, with room to grow as the team scales.
Read the rest of the job →
For people searching Engineer

144,883 engineer jobs are open right now

Here's how to pick the right one and stand out in your application.

144.883Jobs
31.687IN
81%EN

That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.

Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.

Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.

When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.

👁 21 have read this
0 comments
Want to comment?

Leave your e-mail to comment, react and follow the posts for your role. It is free.

Similar jobs

Job on WorkMundi — the world's largest job board. See more jobs from every continent, updated live.

📢
🎁

Before you apply, rehearse this interview.

Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card.

I want my training →