← WorkMundi · 1M+ jobs from around the world, liveSign inCreate free account

System Architect for Observability and Monitoring Platforms

BT Group · Bangalore

📅 06/08/2026
🔔 Alert me about jobs like this
No password, no sign-up. Just the email — and you can leave the list anytime.
🔓 Apply — free →
Opens this job on WorkMundi. The account is free and takes under a minute.

See the other 113,540 jobs in India →

🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Software Engineering Specialist Recruiter: Nishita Jena Hiring Manager: Hari Annamalai Career Grade: D About the role We are looking for a System Architect for Observability & Monitoring Platforms with 15+ years of experience who will own and drive the end-to-end architecture of a large-scale infrastructure monitoring and cloud orchestration platform. This role focuses on system architecture, technical strategy, and long-term platform evolution, enabling the platform to help metrics, logs, distributed tracing, and application performance monitoring at scale. This role combines judicious architectural ownership with hands-on validation through design reviews, code reviews, and targeted proof-of-concepts (POCs). What youll be doing Architecture Ownership & Vision Own the overall system architecture of the observability platform across ingestion, processing, storage, and query layers. Work closely with Enterprise Architect to align system architecture with long-term architectural vision and technical roadmap aligned with business and platform goals. Create and review high-level system architecture, data flows, integration patterns, and core technology choices. Act as the technical authority for complex architectural decisions, trade-offs, and create reviews. Platform & Domain Leadership Architect systems for infrastructure monitoring, metrics, logs, and distributed tracing at scale. Guide the evolution from infrastructure monitoring to a full observability platform including APM. Define architectural patterns for high-throughput telemetry ingestion, real-time processing, and query-at-scale. Ensure architectural consistency for multi-tenant, cloud-native distributed platforms. Standards, Governance & Enablement Establish and evolve architecture standards, design-principles, and best practices across teams. Identify architectural risks and proactively drive mitigation strategies. Enable teams through reference architectures, design-frameworks, and technical guidance. Support modernization initiatives including scalability, performance optimization, resilience, and cost efficiency. Hands-On Architectural Validation Perform code reviews of critical and performance-aware components. Build and guide proof-of-concepts (POCs) to formalize architectural decisions and reduce risk new technologies. Develop reference implementations to demonstrate architectural intent. Collaborate closely with senior engineers to troubleshoot complex system-level issues Delivery Collaboration & Execution Enablement Work closely with Software Engineering Managers to align architectural decisions with delivery and release plans. Assist in breaking down large architectural initiatives into phased, incremental deliverables. Identify architectural and technical risks early and proactively surface them to influence release planning. Support release readiness by formalizing that architecture, scalability, and nonfunctional requirements are addressed ahead of key milestones. Essential Skills / Experience Strong foundation in system design, distributed systems, and application scalability Experience creating microservice based, event driven architectures Ability to make architectural trade offs involving scalability, reliability, performance, and cost Demonstrate experience creating large scale, cloud based distributed platforms Strong backend experience with Java (Spring Bootbased microservices) Working knowledge of Python and/or Go for scripting, automation, and collectors Ability to review and reason about performance critical backend code Strong understanding of infrastructure monitoring and observability concepts with hands on experience on follow: o Metrics, logs, and distributed tracing o Agent based and agentless monitoring models o Push / pull / subscription based data collection Hands on or deep working knowledge of: o Metric data models and storage (e.g., VictoriaMetrics or equivalent) o Log aggregation and search platforms (Elasticsearch / OpenSearch) Familiarity with OpenTelemetry, exporters, and custom collectors Understanding of common monitoring tools and protocols (e.g., SNMP, Prometheus style systems) Strong experience with: o Time series databases for metrics o Relational databases (PostgreSQL / MySQL) for metadata and control plane services o NoSQL / in memory stores (e.g., Redis) for high volume or low latency workloads Working knowledge of search and analytics engines for logs and traces Understanding of data modeling, query patterns, and data lifecycle management Experience with Kafka class distributed messaging systems Designing and operating event driven telemetry ingestion pipelines Understanding of throughput, backpressure, and reliability concerns in streaming systems Strong hands on exposure to Docker and Kubernetes based platforms Infrastructure automation using: o Terraform (IaC) o Ansible (deployment and configuration automation) .
Read the rest of the job →

Similar jobs

Job on WorkMundi — the world's largest job board. See more jobs from every continent, updated live.

📢
🎁

Before you apply, rehearse this interview.

Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card.

I want my training →