Observability Architect- Insurance- 25kMYR- Contract
Argyll Scott ·www.argyllscott.com
Apply directWe're seeking an experienced Observability Architect to define, design, implement, and govern enterprise observability capabilities across our multi-cloud, hybrid, and containerized environments
You will establish consistent practices for monitoring, logging, distributed tracing, event correlation, service health, and operational intelligence across Azure, Alibaba Cloud, and other enterprise platforms.
Dynatrace is the primary enterprise observability platform and is integrated with ServiceNow ITOM to support event management, incident enrichment, service mapping, root-cause analysis, and operational automation. You will also define complementary standards and integration patterns for Elastic, OpenSearch, Prometheus, Grafana, OpenTelemetry, and cloud-native monitoring services.
A key focus of this role is advancing AIOps and self-healing capabilities to reduce alert noise and manual intervention, improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR), and strengthen overall application and platform reliability.
Key Responsibilities
Observability Architecture and Strategy
- Define and maintain enterprise observability reference architectures, standards, design patterns, and governance across Group and its Business Units.
- Design end-to-end observability solutions covering infrastructure monitoring, application performance monitoring, distributed tracing, logging, digital experience monitoring, network observability, and service-health dashboards.
- Establish standards for observability KPIs, Service-Level Objectives (SLOs), Service-Level Indicators (SLIs), error budgets, alerting, and service-reliability reporting.
- Review solution designs and ensure observability requirements are embedded throughout the architecture and delivery lifecycle.
- Promote Site Reliability Engineering (SRE) principles and consistent operational practices across the organisation.
Dynatrace Platform Architecture
- Lead the architecture, adoption, and governance of Dynatrace as ’s enterprise observability platform.
- Define monitoring standards for applications, Kubernetes, virtual machines, databases, middleware, APIs, networks, and cloud-native services.
- Establish Dynatrace standards covering tagging, management zones, dashboards, service mapping, synthetic monitoring, Real User Monitoring, and Davis AI.
- Define scalable platform-onboarding models, configuration standards, reporting requirements, and observability data-retention policies.
- Drive effective platform adoption and ensure Dynatrace delivers measurable operational value.
ServiceNow ITOM Integration
- Architect integrations between Dynatrace and ServiceNow ITOM Event Management and ITOM Health.
- Design event-correlation, alert-enrichment, topology-aware service-impact analysis, and automated incident-creation workflows.
- Establish standards for noise reduction, event deduplication, priority mapping, escalation, and operational ownership.
- Integrate observability insights into ITSM processes, command-centre dashboards, and operational reporting.
AIOps and Self-Healing Automation
- Define and implement AIOps patterns for anomaly detection, predictive alerting, root-cause identification, and closed-loop remediation.
- Design self-healing workflows across Dynatrace, ServiceNow ITOM, automation platforms, and cloud-native services.
- Identify recurring operational issues and convert suitable failure patterns into automated remediation runbooks.
- Reduce alert fatigue, incident recurrence, manual intervention, MTTD, and MTTR through continuous improvement and automation.
Logging and Observability Data Platforms
- Architect centralised logging and analytics solutions using Elastic, OpenSearch, cloud-native logging services, and enterprise logging standards.
- Define standards for log collection, normalisation, enrichment, indexing, retention, access control, and cost management.
- Establish correlation patterns across logs, metrics, traces, events, configuration data, and service topology.
- Ensure logging platforms support operational troubleshooting, security visibility, auditability, and regulatory compliance.
Open-Source and Cloud-Native Observability
- Design monitoring, telemetry, and visualisation solutions using Prometheus, Grafana, OpenTelemetry, and related open-source technologies.
- Define observability patterns for AKS, Alibaba Cloud ACK, containers, microservices, APIs, service meshes, and DevOps pipelines.
- Establish approaches for metrics federation, dashboarding, alerting, and integration with Dynatrace and other enterprise platforms.
- Promote consistent instrumentation and telemetry standards across modern application environments.
Azure and Alibaba Cloud Monitoring
- Define monitoring standards for Azure services, including Azure Monitor, Log Analytics, Application Insights, Network Watcher, Azure Managed Prometheus, Azure Managed Grafana, Azure Advisor, and Azure Resource Health.
- Define monitoring standards for Alibaba Cloud services, including CloudMonitor, Log Service (SLS), ActionTrail, ARMS, Managed Service for Prometheus, and Security Center.
- Integrate cloud-native monitoring capabilities with enterprise observability, ITOM, incident-management, and reporting processes.
- Develop cost-effective telemetry and data-retention strategies across Azure and Alibaba Cloud environments.
Security, Governance and Compliance
- Embed security-by-design, least-privilege access, data protection, and compliance requirements into the observability architecture.
- Support audits, risk assessments, regulatory reviews, and evidence requests relating to monitoring, logging, and service reliability.
- Establish guardrails for platform access, data retention, data classification, dashboard sharing, and operational reporting.
- Maintain architecture documentation, technical standards, design patterns, operational controls, and runbooks.
Technical Leadership and Collaboration
- Serve as the enterprise subject-matter expert for observability, monitoring, logging, AIOps, and self-healing automation.
- Partner with Cloud Architecture, Cloud Engineering, Operations, Security, DevOps, Application, Service Management, and Business Unit teams.
- Mentor engineering and operations teams on observability practices, platform onboarding, dashboard design, and incident-reduction techniques.
- Lead observability transformation initiatives and influence technology and operational roadmaps.
- Work with vendors and delivery partners to maximise platform value, improve supportability, and reduce operational risk.
What You’ll Bring
- Significant experience designing enterprise observability solutions across large, complex, multi-cloud or hybrid environments.
- Strong architecture and hands-on technical knowledge of Dynatrace, including application and infrastructure monitoring, service mapping, dashboards, synthetic monitoring, Real User Monitoring, and Davis AI.
- Experience integrating observability platforms with ServiceNow ITOM, particularly Event Management and ITOM Health.
- Strong knowledge of AIOps, event correlation, anomaly detection, automated remediation, and self-healing operating models.
- Experience with logging and analytics technologies such as Elastic, OpenSearch, or cloud-native logging platforms.
- Practical knowledge of Prometheus, Grafana, OpenTelemetry, Kubernetes, containers, microservices, APIs, and service meshes.
- Strong understanding of Azure and Alibaba Cloud monitoring and observability services.
- Experience establishing SLOs, SLIs, error budgets, reliability reporting, and SRE-aligned operational practices.
- Knowledge of observability security, access control, data governance, audit, compliance, and cost-management requirements.
- Strong stakeholder-management and communication skills, with the ability to influence architecture, engineering, and operational teams.
- Experience working within regulated industries or large financial-services organisations would be advantageous.
Argyll Scott Asia is acting as an Employment Business in relation to this vacancy.