Staff Site Reliability Engineer
Job Description
Job Description – Senior Site Reliability Engineer (SRE)
Job Title: Senior Site Reliability Engineer (SRE)
Company: Palo Alto Networks
Employment Type: Full-Time
Experience: 5+ Years
Education: Bachelor’s or Master’s Degree in Computer Science, Information Technology, or a related technical discipline
About the Company
Palo Alto Networks is a global cybersecurity leader dedicated to protecting organizations from evolving cyber threats. The company fosters innovation, collaboration, and continuous learning while building secure, scalable, and resilient digital infrastructure. Employees benefit from comprehensive learning opportunities, wellbeing programs, and a collaborative work environment.
Job Summary
Palo Alto Networks is seeking an experienced Senior Site Reliability Engineer (SRE) to join its Infrastructure & Cloud Operations team. The successful candidate will be responsible for designing, automating, deploying, and maintaining highly available global IT infrastructure supporting customer-facing platforms.
Working closely with Network, Compute, Security, Database, and Application teams, the engineer will build next-generation IT operations through automation, infrastructure as code, observability, analytics, and continuous improvement.
Key Responsibilities
Infrastructure & Cloud Operations
Design, implement, and support Linux-based infrastructure using Infrastructure as Code (IaC).
Provision, configure, and maintain resilient hybrid cloud environments.
Deploy and support highly available production infrastructure.
Manage large-scale Linux server environments.
Build and operate compute platforms supporting thousands of virtual machines and Kubernetes clusters.
Ensure platform scalability, redundancy, resilience, and disaster recovery readiness.
Automation & DevOps
Develop automation frameworks to simplify infrastructure management.
Automate routine operational tasks using Python, Shell, or Bash scripting.
Manage Infrastructure as Code using Terraform, Ansible, Puppet, and Git.
Build and maintain CI/CD pipelines using Jenkins, CircleCI, or similar platforms.
Improve deployment efficiency through automation and continuous integration practices.
Site Reliability Engineering
Maintain high service availability and performance based on business SLAs.
Perform capacity planning and infrastructure optimization.
Design and implement proactive monitoring and alerting systems.
Create observability solutions using MELT principles:
Metrics
Events
Logs
Traces
Perform trend analysis to improve platform reliability.
Kubernetes & Container Platform
Deploy and manage Docker containers.
Administer Kubernetes clusters.
Ensure container platform reliability, scalability, and security.
Automate Kubernetes operations and deployments.
Monitoring & Observability
Design enterprise monitoring solutions.
Implement centralized logging and tracing platforms.
Improve infrastructure visibility and operational intelligence.
Support AIOps initiatives for predictive monitoring and automated incident response.
Security & Compliance
Support security implementations and compliance audits.
Participate in infrastructure hardening.
Optimize API infrastructure security.
Maintain Public Key Infrastructure (PKI) operational processes.
Follow security best practices across cloud infrastructure.
Operations Support
Provide technical support for internal platform users.
Participate in production incident response and root cause analysis (RCA).
Plan maintenance windows.
Prepare and review change requests.
Maintain operational documentation and runbooks.
Participate in on-call support rotations.
Collaboration
Work closely with:
Network Engineering
Security Teams
Cloud Engineering
Database Teams
Product Engineering
Global IT Teams
Collaborate effectively across multiple time zones.
Required Qualifications
Bachelor’s or Master’s Degree in Computer Science, Information Technology, or related field.
Minimum 5+ years of professional experience in Site Reliability Engineering, Cloud Infrastructure, or DevOps.
Required Technical Skills
Operating Systems
Linux Administration
Cloud Platforms
AWS
Google Cloud Platform (GCP)
Hybrid Cloud Infrastructure
Containerization & Orchestration
Docker
Kubernetes
Infrastructure as Code
Terraform
Ansible
Puppet
Git
Programming & Scripting
Python
Shell
Bash
CI/CD Tools
Jenkins
CircleCI
Similar CI/CD platforms
Monitoring & Observability
Metrics
Logs
Events
Traces (MELT)
Infrastructure Monitoring
Alerting Systems
Networking & Security
Networking fundamentals
Security technologies
PKI
API security
Infrastructure security best practices
API Technologies
API development
API optimization
API security
Preferred Skills
Candidates with experience in the following areas will have an advantage:
AIOps
Machine Learning applications in IT Operations
Cloud Observability
Self-healing infrastructure
Big Data technologies
Data Analytics
Enterprise Business Applications
IT Service Management (ITSM)
Disaster Recovery (DR)
Business Continuity Planning (BCP)
Soft Skills
Excellent analytical and troubleshooting abilities
Strong problem-solving mindset
Effective communication skills
Ability to collaborate across global teams
Positive attitude with a customer-first mindset
Adaptability in fast-paced environments
Strong documentation skills
Leadership and ownership mentality
Why Join Palo Alto Networks?
Work with one of the world’s leading cybersecurity companies.
Exposure to large-scale global cloud infrastructure.
Opportunity to work on cutting-edge automation and observability technologies.
Collaborative and innovation-driven culture.
Continuous learning and career development opportunities.
Flexible wellbeing and employee support programs.
Inclusive and diverse workplace.
Opportunity to build next-generation secure cloud platforms.