CareerRiver

Manager, Engineering - Dev Ops/SRE (Hybrid)

CrowdStrike · San Francisco Bay Area

📍 USA - Sunnyvale, CAvia workday
Apply on company site ↗
CareerRiver pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to CrowdStrike.
As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed — we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you. About the Role: At CrowdStrike, Site Reliability Engineering (SRE) is at the forefront of ensuring the reliability and scalability of our cloud-native security platform. In this role, you'll manage a team of talented engineers, providing technical leadership on key projects and empowering them to excel in their roles. As an SRE Manager, you will lead a team of SRE engineers ensuring the reliability, scalability, and performance of CrowdStrike's cloud-native security platform. You'll provide technical leadership and mentorship, owning both reliability engineering and software delivery pipelines - driving engineering velocity while maintaining zero tolerance for downtime in security-critical infrastructure. What you will Do Define and enforce SLOs, SLIs, and error budgets across distributed systems processing millions of events per second Drive system reliability by blending software engineering principles with AI-driven automation, moving from reactive firefighting to proactive, automated operations Lead major incident response and facilitate blameless postmortems, driving systemic reliability improvements Own capacity planning, traffic management, and load shedding strategies for high-throughput distributed systems Own the end-to-end software delivery pipeline strategy — designing, building, and maintaining scalable, reliable pipelines using Jenkins, GitLab CI, and Bitbucket Pipelines Build and maintain observability frameworks including metrics, distributed tracing, and log aggregation across the full stack Champion chaos engineering and resilience validation practices for security-critical systems Lead and grow a high-performing SRE team, mentoring engineers and fostering a culture of continuous learning and operational excellence Partner with cross-functional engineering teams to embed reliability practices early in the software development lifecycle What You'll Need Experience & Leadership Proven track record of building, growing, and retaining high-performing SRE/DevOps engineering teams in a fast-paced, high-growth environment 10+ years of software engineering experience with significant focus on reliability engineering, platform infrastructure, and production operations at scale 3+ years of hands-on management experience overseeing SRE/DevOps engineering teams, including incident command and reliability ownership Bachelor's degree in Computer Science or related field, or equivalent work experience Reliability Engineering Deep understanding of SRE principles including SLOs, SLAs, SLI s, and error budgeting strategies applied to large-scale distributed systems Proven experience owning reliability for high-throughput distributed systems processing millions of events per second, including capacity planning, traffic management, and load shedding strategies Strong incident management facilitating blameless postmortems , and driving system reliability improvements Demonstrated ability to build, operationalize, and maintain highly scalable, security-critical microservices-based distributed systems with zero tolerance for data loss or downtime. Advanced observability experience including Prometheus, Grafana, distributed tracing (Jaeger/OpenTelemetry) , and large-scale log aggregation ( ELK/Splunk ) with a focus on building custom SLO dashboards and reliability scorecards. Experience owning disaster recovery strategies including backup automation, failover testing, and business continuity planning for stateful distributed systems Platform and Delivery Engineering Proficiency in Python and/or Golang for automation, tooling, and platform services Hands-on experience designing and managing scalable software delivery pipelines using Jenkins, GitLab CI, Bitbucket Pipelines , or equivalent Strong proficiency in Infrastructure as Code (IaC) - Terraform, Ansible, Pulumi , or equivalent Familiarity with GitOps workflows using ArgoCD or Flux for managing infrastructure deployments at scale Cloud and  Big Data Exposure Proficiency in at least one cloud environment (AWS, Azure, GCP) with emphasis on multi-region architecture, cloud-native reliability patterns, and security-first cloud design Strong experience with Kubernetes at scale - managing large cluster fleets, workload orchestration, and container lifecycle management Familiarity with distributed data systems including relational databases ( PostgreSQL ), NoSQL ( Cassandra ), OLAP ( Pinot ), Indexing( OpenSearch ) and real-time streaming platforms (Kafka, Flink) Exposure to Big Data and analytics technologies like Spark,Storm . #LI-AP1 Benefits of Working at CrowdStrike: Market

More San Francisco Bay Area jobs

San Francisco Bay Area jobs · Browse all locations