
Real job — pulled straight from AHEAD’s careers page · Verified September 18, 2026 · No reposts.
Job description
AHEAD is hiring a Principal Observability and Reliability Architect — a full-time, remote role. Apply directly on AHEAD's careers page below.
Principal Observability & Reliability Architect
Team: Observability
Location: Chicago, Illinois, New York, New York, Atlanta, Georgia
Commitment: Full Time
Workplace Type: hybrid
Responsibilities
- Architect Dynatrace at enterprise scale: multi-tenant and hybrid designs, ActiveGate topology, OneAgent and OpenTelemetry instrumentation strategy, consumption and licensing governance, tagging and ownership models, and AI-workload observability.
- Design end-to-end observability architectures across monitoring, logging, metrics, tracing, telemetry pipelines, alerting, event correlation, and service visibility in hybrid and multi-cloud environments.
- Establish and mature SRE practices with clients: SLIs and SLOs, error budgets, production readiness, incident response and postmortems, and reliability roadmaps tied to business impact.
- Lead assessment and advisory workshops that define use cases, maturity roadmaps, operating models, and adoption strategies, including AIOps and automation with Davis AI, alerting profiles, and workflows.
- Define standards for telemetry onboarding, naming, tagging, service ownership, access, dashboards, alert governance, runbooks, and operational handoff, and advise on telemetry governance: data quality, retention, sampling, cardinality, and cost.
- Lead modernization initiatives: tool and alert rationalization, telemetry strategy, migration to Dynatrace from legacy APM and monitoring platforms, and integration with ITSM, CMDB, event management, and automation platforms.
- Lead complex programs, own solution design and architectural review, articulate trade-offs, and act as escalation point for delivery teams.
- Advise client executives on platform strategy and value realization, and report on program health and outcomes.
- Provide architecture and quality oversight across engagements, intervening early where outcomes are at risk.
- Support pursuits as the technical expert: scoping, positioning, demonstrations, estimate validation, and client-facing technical narratives.
- Build reusable assets such as Dynatrace reference architectures, governance models, accelerators, and points of view, and contribute thought leadership through content, partner material, and conference speaking.
- Mentor architects and consultants across the practice, and maintain Dynatrace Professional certification plus professional-level certification on at least one additional platform.
Qualifications
- 8 or more years of hands-on experience in observability, APM, SRE, or related disciplines, including architecting enterprise-scale solutions across distributed systems and multi-cloud estates.
- 4 or more years of hands-on enterprise Dynatrace experience, including architecture and governance, OneAgent and Kubernetes deployment, Smartscape and PurePath, Grail and DQL, Davis AI, SLOs, workflows and automation, and ITSM integration; Dynatrace Professional certification held or attainable within six months.
- Applied SRE experience defining SLIs and SLOs, operating error budgets, running production readiness and incident reviews, and leading reliability programs that measurably reduce incidents and time to resolve.
- Working expertise in OpenTelemetry, Prometheus, and the Grafana ecosystem, and in public cloud monitoring services on AWS, Azure, or GCP.
- Strong knowledge of telemetry governance (routing, transformation, enrichment, retention, access, cost) and experience defining enterprise standards for dashboards, alerts, tagging, and service ownership.
- Expert knowledge of platform architecture, API integration patterns, and automation frameworks (Terraform, Ansible, Python, or similar).
- Strong consultative and executive-facing presence, with experience leading workshops and translating business needs into architecture and delivery plans.
- Demonstrated leadership mentoring technical teams; familiarity with ITIL, ITSM, and DevOps principles and with scoping, estimating, and change control in consulting delivery.
- Strong communication skills, attention to detail, and a self-starting work style; able to travel occasionally (0% to 15%).
Preferred Qualifications
- Both Dynatrace Professional certifications, or Dynatrace partner program experience.
- Migration experience from AppDynamics, New Relic, Datadog, Splunk, or legacy monitoring platforms to Dynatrace; LogicMonitor or other infrastructure monitoring experience is a plus.
- Telemetry pipeline tools such as OpenTelemetry Collector, Grafana Alloy, Fluent Bit, Kafka, Cribl, or Vector, plus Kubernetes, CI/CD, and infrastructure as code.
- Integration with , Jira Service Management, PagerDuty, Opsgenie, BigPanda, or xMatters.
- Published thought leadership, conference speaking, or ownership of a named offering or accelerator; relevant cloud, SRE, ITIL, or FinOps certifications are a plus.
Get Principal Observability and Reliability Architect jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs

Staff Software Engineer, Reporting, Data Platform & Observability



Senior Software Engineer, Observability Accelerators & AI (Remote)
Frequently asked questions
Is Principal Observability and Reliability Architect at AHEAD a remote job?
Yes, Principal Observability and Reliability Architect at AHEAD is a remote position. Candidates in Chicago, IL, New York, NY, Atlanta, GA may be preferred.
What skills are required for Principal Observability and Reliability Architect at AHEAD?
The required skills for Principal Observability and Reliability Architect at AHEAD include: OpenTelemetry, Prometheus, Grafana, AWS, Azure, GCP, Terraform, Ansible, Python, ITIL, DevOps, New Relic, Datadog, Splunk, Kafka, Kubernetes, CI/CD.
What is the seniority level for Principal Observability and Reliability Architect at AHEAD?
Principal Observability and Reliability Architect at AHEAD is a Principal / Senior level position.
How do I apply for Principal Observability and Reliability Architect at AHEAD?
You can view the full description and apply for Principal Observability and Reliability Architect at AHEAD on EchoJobs: https://echojobs.io/job/ahead-principal-observability-reliability-architect-anu1i.