Job Title:
Lead Site Reliability Engineer
Job Details:
Business Unit: Tech & Digital
Team: Ent Factory-Channels, Mobility, Payments
Reports to: SRE Manager
Location: Mumbai, Chennai, Gurgaon & Bangalore
Role Type: Non-Supervisory
No of direct reportees: 0
Travel Required: No
Job Band Range: D1/D2
JD Created date: 28th Jan 2023
Job Purpose:
Analyzing, troubleshooting, and designing vital services, platforms, and infrastructure on GCP with a focus on reliability, scalability, resilience, security, and performance.
Lead and Mentor a team of SRE engineers.
Job Responsibilities:
Execute reliability initiatives for the team and organization.
Mentor and lead a team of SRE engineers.
Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentation, and code with other engineering teams.
Solid understanding of observability tools and ability to express reliability metrics via observability.
Define KPI in the form of RPO/RTO/SLI/SLO/Error Budget.
Apply automation and software to any manually performed tasks or system parts.
Troubleshoot complicated, cross-platform issues handling OS, Networking, Database in a cloud-based SaaS environment and manage live production incidents.
Monitor application performance, take steps to improve overall application performance and stability, and follow through with implementation.
Conduct system analysis, configuration management, and develop improvements for system software performance, availability, and reliability.
Design, write, ship, and motivate the creation of software and systems to increase observability, product reliability, and organizational efficiency.
Maintain and monitor deployment, orchestration of servers, docker containers, databases, and general backend infrastructure.
Develop Run Books/Standard Operating Procedure for recurring Production issues and work on permanent solutions.
Perform Incident Analysis regularly to prevent and find long-term solutions for Incidents.
Educational Qualifications:
B Tech in Computer Science or related discipline preferred.
Key Skills:
Experience in monitoring and analyzing infrastructure performance using standard performance monitoring tools.
Demonstrable experience in Containerization (Docker) and orchestration (Kubernetes).
Experience with Infrastructure As Code (Terraform, Cloud Formation, Ansible).
Knowledge and proven hands-on experience in large-scale databases and distributed technologies, such as Kafka and Confluent Platform Kafka.
Basic programming and scripting skills.
Solid understanding of at least 2 observability technologies.
Experience Required:
Total Yrs of experience: 11-13
Major Stakeholders:
Internal: Product Manager from Digital Factory, Business Analyst from BTG team, Incident Management team, Development Team.
Refer to the Job Description
Eligibility typically includes the qualifications and experience outlined in the job description above, with around 0 - 1 Years years of relevant experience expected for this role.
The key responsibilities for this role are detailed in the Key Responsibilities section above, covering the core duties expected of a Team Lead-Business-BIU at HDFC Bank.
This role requires around 0 - 1 Years years of relevant experience, as specified in the job listing. Please refer to the Qualifications & Experience section above for full details.
Skills relevant to this position are outlined in the Qualifications & Experience section above. In general, strong communication, domain knowledge, and the ability to meet role-specific targets are valued across similar BFSI positions.
You can apply directly using the Apply Now button on this page, which will take you to HDFC Bank's application process for this role.