JobInAWeekJobInAWeek
JobsInternshipsCompaniesSalariesResumePricing
Sign in
Jobs / DevOps engineer jobs in India / DevOps engineer jobs in Chennai /
P

Sr. Site Reliability Engineer

PayPal · Chennai

Posted
21 days ago
Experience
3+ yrs
Pay
Not stated
Role
DevOps engineer
KubernetesAWSGCPTroubleshootingAzureStakeholder managementPythonLinux
Apply on PayPal's siteOpens the company's own careers page.

About the job

The Company

PayPal has been revolutionizing commerce globally for more than 25 years. Creating innovative experiences that make moving money, selling, and shopping simple, personalized, and secure, PayPal empowers consumers and businesses in approximately 200 markets to join and thrive in the global economy.

We operate a global, two-sided network at scale that connects hundreds of millions of merchants and consumers. We help merchants and consumers connect, transact, and complete payments, whether they are online or in person. PayPal is more than a connection to third-party payment networks. We provide proprietary payment solutions accepted by merchants that enable the completion of payments on our platform on behalf of our customers.

We offer our customers the flexibility to use their accounts to purchase and receive payments for goods and services, as well as the ability to transfer and withdraw funds. We enable consumers to exchange funds more safely with merchants using a variety of funding sources, which may include a bank account, a PayPal or Venmo account balance, PayPal and Venmo branded credit products, a credit card, a debit card, certain cryptocurrencies, or other stored value products such as gift cards, and eligible credit card rewards.  Our PayPal, Venmo, and Xoom products also make it safer and simpler for friends and family to transfer funds to each other. We offer merchants an end-to-end payments solution that provides authorization and settlement capabilities, as well as instant access to funds and payouts. We also help merchants connect with their customers, process exchanges and returns, and manage risk. We enable consumers to engage in cross-border shopping and merchants to extend their global reach while reducing the complexity and friction involved in enabling cross-border trade.

Our beliefs are the foundation for how we conduct business every day.  We live each day guided by our core values of Inclusion, Innovation, Collaboration, and Wellness. Together, our values ensure that we work together as one global team with our customers at the center of everything we do – and they push us to ensure we take care of ourselves, each other, and our communities.

Job Summary: What do you need to know about the role

This is an incident command role. You'll direct application and infrastructure teams during incidents making work assignments, prioritizing troubleshooting paths, and authorizing critical actions like rollbacks and regional failovers. You need the technical depth to rapidly read Infrastructure as Code, Kubernetes manifests, and CI/CD configurations to make informed decisions under pressure.

Meet Our Team

The Site Health Engineering (Command Center) team serves as the operational and technical authority during PayPal's most critical incidents. We are a team of experienced infrastructure and reliability professionals who blend deep technical expertise with sound judgment under pressure, directing cross-functional engineering efforts across PayPal's core platforms and family of brands, including Venmo, Xoom, Zettle, and Braintree. Job Description: Essential Responsibilities:

- Delivers complete solutions spanning all phases of the Software Development Lifecycle (SDLC) (design, implementation, testing, delivery and operations), based on definitions from more senior roles.

- Advises immediate management on project-level issues

- Guides junior engineers

- Operates with little day-to-day supervision, making technical decisions based on knowledge of internal conventions and industry best practices

- Applies knowledge of technical best practices in making decisions

Minimum Qualifications:

- 3+ years relevant experience and a Bachelor’s degree OR Any equivalent combination of education and experience.

Additional Responsibilities & Preferred Qualifications :

Your way to Impact

Rather than building or maintaining infrastructure day-to-day, our team is entrusted with a broader mandate: safeguarding the reliability, resiliency, and availability of some of the world's most heavily trafficked financial platforms. We hold final decision-making authority during high-severity incidents, partner closely with executive leadership on post-incident learnings, and drive the tooling and processes that continuously strengthen our incident response capabilities.

You'll also regularly interface with executive leadership during critical incidents and post-mortems, and drive implementation of tooling that advances the Command Center's capabilities.

Site Resiliency & Infrastructure Management

- - Proactively identify and address vulnerabilities in cloud (AWS, GCP, Azure) and on-premises infrastructure

- Review Infrastructure as Code changes for reliability risks as part of change approval process

- Identify architectural anti-patterns in Kubernetes deployments and cloud migrations

- Conduct regular disaster recovery drills and readiness tests before major events (Thanksgiving, Cyber 5, peak shopping seasons)

- Participate in situation room activities for new product rollouts

- Drive site resilience projects to enhance system reliability and uptime

- Proactively identify and address vulnerabilities in cloud (AWS, GCP, Azure) and on-premises infrastructure

- Implement automated monitoring solutions to detect single points of failure

- Lead new datacenter and CDN certification initiatives

- Conduct regular disaster recovery drills and readiness tests before major events (Thanksgiving, Cyber 5, peak shopping seasons)

- Participate in situation room activities for new product rollouts

- Drive site resilience projects to enhance system reliability and uptime

Incident Management & Response

- - Act as incident commander with final decision authority -- directing engineering teams, authorizing rollbacks, and commanding regional failovers

- Direct application and infrastructure teams during incidents by making work assignments and prioritizing troubleshooting paths

- Rapidly assess incidents by reading Infrastructure as Code (Terraform, CloudFormation), Kubernetes manifests, and CI/CD configurations

- Give final authorization for critical actions including production rollbacks, regional failovers, and emergency changes

- Interface with executive leadership during critical incidents and post-mortems to provide technical guidance and impact assessments

- Identify when incidents stem from teams deviating from established cloud-native patterns

- Command cross-functional teams during high-severity incidents affecting PayPal core and brand platforms (Venmo, Xoom, Zettle, Braintree)

- Lead blameless postmortem sessions and contribute to Root Cause Analysis (RCA) processes

- Drive continuous improvement initiatives based on incident learningsServe as the primary technical escalation point during critical incidents

- Accelerate incident response times through standardized playbooks and automated workflows

- Coordinate cross-functional teams during high-severity incidents affecting PayPal core and brand platforms (Venmo, Xoom, Zettle, Braintree)

- Lead blameless postmortem sessions and contribute to Root Cause Analysis (RCA) processes

- Drive continuous improvement initiatives based on incident learnings

- Manage multiple concurrent incidents during peak periods with efficiency and precision

Change Management & Risk Mitigation

- - Serve as final approver for emergency changes and provide expert guidance on all production changes

- Act as advisor and technical authority during change approval processes, identifying potential reliability risks

- Provide training and guidance to engineering teams on change management best practices

- Maintain change audit documentation and compliance requirements

- Review and approve changes to production systems, ensuring comprehensive risk assessment

- Automate change validation and rollback procedures to minimize service disruptions

- Streamline change management processes to reduce manual errors and bottlenecks

- Provide training and guidance to engineering teams on change management best practices

- Maintain change audit documentation and compliance requirements

Cloud Expertise & Technical Leadership

- - Leverage deep expertise in cloud platforms (AWS, GCP, Azure) to drive incident resolution

- Support Braintree and Venmo cloud infrastructure operations

- Guide teams toward solutions by providing architectural direction during incidents

- Stay current with emerging cloud technologies and best practices

- Mentor team members on cloud technologies and incident management techniques

Cloud Expertise & Technical Leadership

- - Implement automation, dashboards, and tooling to enhance the team's incident response capabilities

- Build runbooks and playbooks for cloud-native incident scenarios

- Develop internal tools and scripts to improve TDO operational efficiency

- Drive projects that advance the Command Center's operational capabilities

In your day-to-day role you will be responsible for:

- - 3 days   -4 days   alternating 12-hour shift pattern.

- Take ownership of system performance monitoring, identify inefficiencies, and lead initiatives to improve the overall availability and reliability of digital platforms and applications.

- Lead and manage the response to complex, high-priority incidents, ensuring prompt resolution and a thorough root cause analysis to prevent future occurrences.

- Design and implement advanced automation frameworks to improve operational efficiency, streamline processes, and reduce human error.

- Lead reliability-focused initiatives, ensuring systems are highly available, resilient, and scalable, and promote best practices across engineering teams.

- Enhance the monitoring infrastructure by identifying key metrics, optimizing alerting, and improving system observability to ensure the reliability of large-scale systems.

- Forecast resource requirements and lead capacity planning activities to ensure systems can scale effectively to meet growing user demand.

- Ensure robust disaster recovery strategies are in place and conduct regular testing to ensure systems can recover quickly from failures.

- Partner with engineering and product teams to identify opportunities for improving system architecture, focusing on scalability, reliability, and fault tolerance.

- Provide mentorship and technical guidance to junior site reliability engineers, fostering skill development and knowledge sharing.

Apply now

Do you fit this job?

Create your profile and every open job, this one included, gets a fit score for you.

More jobs like this

  • DevOps engineer jobs in Chennai
  • All jobs in Chennai
  • All jobs at PayPal

Similar jobs

All devops engineer jobs →
N

Site Reliability Engineer II (GCP & Azure)

NCR
Chennai3–6 yrsDevOps engineer
CI/CDAzureGitGCPDockerKubernetes
Posted yesterdayApply
K

DevOps Engineer

KLA
ChennaiDevOps engineer
CI/CDGCPAzurePythonGitDocker
Posted 3 days agoApply
C

Engineer 3 - DevOps

Comcast
Chennai5–7 yrsDevOps engineer
CI/CDKubernetesLinuxGitAWSGCP
Posted 4 days agoApply
S

Platform Engineer & Cloud Ops Engineer

Sutherland
Hyderabad · Remote OKDevOps engineer
KubernetesAWSCI/CDGCPTerraformGit
Posted todayApply
O

Senior Site Reliability Engineer

Okta
BengaluruDevOps engineer
PythonGoSQLKubernetesCI/CDTerraform
Posted yesterdayApply
L

Lead Platform Engineer, Manager

LSEG
HyderabadDevOps engineer
AzureGoCI/CDGitLinuxTerraform
Posted yesterdayApply

Jobs by role

  • Backend developer jobs in India
  • Full-stack developer jobs in India
  • Program / project manager jobs in India
  • Data scientist jobs in India
  • Data analyst jobs in India
  • Product manager jobs in India
  • Consultant jobs in India
  • Strategy & operations jobs in India
  • QA / Test engineer jobs in India
  • Frontend developer jobs in India
  • DevOps engineer jobs in India
  • Tech support jobs in India
  • Research analyst jobs in India
  • Mobile developer jobs in India

Jobs by city

  • Jobs in Bengaluru
  • Jobs in Hyderabad
  • Jobs in Pune
  • Jobs in Gurugram
  • Jobs in Chennai
  • Jobs in Mumbai
  • Jobs in Noida
  • Jobs in Delhi
  • Jobs in Coimbatore
  • Jobs in Ahmedabad
  • Jobs in Kolkata
  • Jobs in Chandigarh

Popular searches

  • Backend developer jobs in Bengaluru
  • Backend developer jobs in Hyderabad
  • Full-stack developer jobs in Bengaluru
  • Program / project manager jobs in Bengaluru
  • Data scientist jobs in Bengaluru
  • Product manager jobs in Bengaluru
  • Data analyst jobs in Bengaluru
  • Consultant jobs in Bengaluru
  • Backend developer jobs in Chennai
  • QA / Test engineer jobs in Bengaluru
  • Backend developer jobs in Pune
  • Frontend developer jobs in Bengaluru
  • Backend developer jobs in Gurugram
  • DevOps engineer jobs in Bengaluru
  • Full-stack developer jobs in Hyderabad
  • Program / project manager jobs in Hyderabad
  • Backend developer jobs in Mumbai
  • Strategy & operations jobs in Bengaluru

JobInAWeek

  • How to get a job in a week
  • Create your profile
  • Internships & fresher jobs
  • Companies hiring
  • Salary check
  • Swipe jobs
  • Resume builder
  • Pricing
  • Privacy
  • Terms
  • Refunds & cancellation
  • Contact
  • About our job crawler
JobInAWeek – Skills today. Job tomorrow.JobInAWeek – Skills today. Job tomorrow.

Free for job seekers. Jobs come from company careers pages and link to the company's own apply page. Updated every two hours.