Principal Platform Engineer, Observability (CIPE)
Palo Alto Networks · California
📍 Office - USA - CA - Headquartersvia workday
Apply on company site ↗
CareerRiver pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to Palo Alto Networks.
Our Mission
At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.
Who We Are
In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!
We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.
Job Summary
Your Career
We are looking for a Principal Software Engineer to architect, build, and evolve our observability platform across infrastructure, applications, and developer workflows. This role is ideal for a hands-on technical leader with deep experience in open source observability technologies and Chronosphere, who is equally fluent in building AI-enabled systems and developer experiences using modern AI coding tools such as Claude and Codex.
You will serve as a technical architect for the observability stack, working across engineering, platform, SRE, and product teams to define standards for metrics, logs, traces, profiling, synthetics, alerting, dashboards, and incident response. You will also lead the integration of AI agents, copilots, and skill-based automation into observability workflows — making telemetry, debugging, and reliability operations equally consumable by humans and AI agents. You should be comfortable operating at both strategic and implementation levels: designing architecture, writing production-grade code, reviewing systems, mentoring engineers, and driving adoption across teams.
Your Impact
Observability Architecture
Design and lead the evolution of a modern observability platform using OpenTelemetry, Prometheus, Jaeger, Alertmanager, and related CNCF ecosystem tools.
Define architecture standards for telemetry collection, processing, storage, querying, visualization, alerting, retention, and governance.
Build scalable systems for metrics, distributed tracing, continuous profiling, log aggregation, synthetic monitoring, service health monitoring, and reliability analytics.
Establish best practices for instrumentation across services, infrastructure, Kubernetes workloads, CI/CD systems, and developer platforms.
Evaluate trade-offs around data cardinality, sampling, storage cost, retention, query performance, multi-tenancy, reliability, and operational complexity.
Make pragmatic recommendations on open source, self-managed, managed-service, and hybrid observability approaches.
Create paved-road observability patterns that help engineering teams instrument, monitor, debug, and operate services with minimal friction.
OpenTelemetry and Instrumentation
Lead adoption and standardization of OpenTelemetry across applications, services, infrastructure, and platform components.
Design and implement telemetry pipelines using OpenTelemetry Collector, exporters, processors, receivers, connectors, and custom extensions where needed.
Define conventions for traces, metrics, logs, spans, attributes, resources, service names, correlation IDs, and semantic conventions.
Build libraries, SDK wrappers, golden paths, and internal tooling to simplify observability instrumentation for engineering teams.
Metrics, Monitoring, and Alerting
Architect metrics systems using Prometheus-compatible formats, PromQL, remote write, federation, scraping strategies, service discovery, recording rules, and long-term storage backends.
Design alerting frameworks that reduce noise, improve signal quality, and align with SLOs, SLIs, error budgets, and incident response practices.
Create reusable alerting patterns for Kubernetes, infrastructure, applications, APIs, databases, queues, event-driven systems, and distributed services.
Define standards for dashboarding, runbooks, escalation policies, alert ownership, and production readiness.
Partner with SRE and engineering teams to mature monitoring practices and improve service reliability.
Kubernetes and Platform Engineering
Build observability capabilities for Kubernetes environments, including cluster monitoring, workload telemetry, service mesh visibility, ingress and egress monitoring, and node-level insights.
Develop and maintain Helm charts, Kubernetes manifests, operators, sidecars, agents, DaemonSets, and deployment automation for observability components.
Work with platform teams to ensure observability systems are reliable, secure, multi-tenant, highly available, and easy to operate.
Define standards for resource usage, scaling, upgrades, failover, backup, disaster recovery, access control, and tenant isolation for observability infrastructure.
Support observability across multi-cluster, multi-region, and hybrid cloud environments where applicable.
AI-Enabled Observability and Developer Experience
Design and build AI-enabled observability workflows that allow both humans and AI agents to investigate incidents, query telemetry, summarize signals, and propose remediations.
Define and publish reusable AI skills, agents, and tools (e.g., Claude skills, Codex tools, MCP servers, structured prompts) that encode observability best practices and ma
More California jobs
California jobs · Browse all locations