Job overview
Complete role details
Location
Chennai
Employment type
Other
Workplace
Hybrid
Seniority
Lead
Skills
KubernetesSplunk
Role details
Job description
Lead - OpenShift Engineer
Install, configure, administer, and manage Red Hat OpenShift (OCP) clusters on Bare Metal and VMware platforms.
Perform OpenShift patching and version upgrades with minimal disruption to production workloads.
Maintain and support existing DevOps infrastructure pipelines integrated with OpenShift.
Create, manage, and troubleshoot persistent storage solutions, including file systems and persistent volumes.
Configure and manage RBAC and security policies across clusters.
Configure and manage Ingress Controllers, Routes, and network load balancers.
Support inter-cluster communication and hybrid/multi-cluster architectures.
Apply a comprehensive understanding of networking and firewall concepts within Kubernetes and OpenShift environments.
Troubleshoot cluster networking, DNS, and connectivity issues.
Monitor cluster, node, and application health using Prometheus, Grafana, and Splunk.
Configure alerts, dashboards, and metrics for proactive monitoring.
Identify performance bottlenecks and recommend optimization strategies.
Troubleshoot complex issues related to cluster performance, storage, networking, and application deployments.
Resolve pod and container runtime issues.
Respond to production issues promptly and professionally, adhering to SLAs.
Lead resolution of high-priority P1/P2 incidents and provide detailed Root Cause Analysis (RCA) reports to clients and stakeholders.
Demonstrate familiarity with enterprise ITSM processes.
Exhibit excellent troubleshooting, documentation, and communication skills.
Perform OpenShift patching and version upgrades with minimal disruption to production workloads.
Maintain and support existing DevOps infrastructure pipelines integrated with OpenShift.
Create, manage, and troubleshoot persistent storage solutions, including file systems and persistent volumes.
Configure and manage RBAC and security policies across clusters.
Configure and manage Ingress Controllers, Routes, and network load balancers.
Support inter-cluster communication and hybrid/multi-cluster architectures.
Apply a comprehensive understanding of networking and firewall concepts within Kubernetes and OpenShift environments.
Troubleshoot cluster networking, DNS, and connectivity issues.
Monitor cluster, node, and application health using Prometheus, Grafana, and Splunk.
Configure alerts, dashboards, and metrics for proactive monitoring.
Identify performance bottlenecks and recommend optimization strategies.
Troubleshoot complex issues related to cluster performance, storage, networking, and application deployments.
Resolve pod and container runtime issues.
Respond to production issues promptly and professionally, adhering to SLAs.
Lead resolution of high-priority P1/P2 incidents and provide detailed Root Cause Analysis (RCA) reports to clients and stakeholders.
Demonstrate familiarity with enterprise ITSM processes.
Exhibit excellent troubleshooting, documentation, and communication skills.
͏
Do
- Provide adequate support in architecture planning, migration & installation for new projects in own tower (platform/dbase/ middleware/ backup)
- Lead the structural/ architectural design of a platform/ middleware/ database/ back up etc. according to various system requirements to ensure a highly scalable and extensible solution
- Conduct technology capacity planning by reviewing the current and future requirements
- Utilize and leverage the new features of all underlying technologies to ensure smooth functioning of the installed databases and applications/ platforms, as applicable
- Strategize & implement disaster recovery plans and create and implement backup and recovery plans
- Manage the day-to-day operations of the tower
- Manage day-to-day operations by troubleshooting any issues, conducting root cause analysis (RCA) and developing fixes to avoid similar issues.
- Plan for and manage upgradations, migration, maintenance, backup, installation and configuration functions for own tower
- Review the technical performance of own tower and deploy ways to improve efficiency, fine tune performance and reduce performance challenges
- Develop shift roster for the team to ensure no disruption in the tower
- Create and update SOPs, Data Responsibility Matrices, operations manuals, daily test plans, data architecture guidance etc.
- Provide weekly status reports to the client leadership team, internal stakeholders on database activities w.r.t. progress, updates, status, and next steps
- Leverage technology to develop Service Improvement Plan (SIP) through automation and other initiatives for higher efficiency and effectiveness
͏
Team Management
- Resourcing
- Forecast talent requirements as per the current and future business needs
- Hire adequate and right resources for the team
- Train direct reportees to make right recruitment and selection decisions
- Talent Management
- Ensure 100% compliance to WiproâÂÂs standards of adequate onboarding and training for team members to enhance capability & effectiveness
- Build an internal talent pool of HiPos and ensure their career progression within the organization
- Promote diversity in leadership positions
- Performance Management
- Set goals for direct reportees, conduct timely performance reviews and appraisals, and give constructive feedback to direct reports.
- Ensure that organizational programs like Performance Nxt are well understood and that the team is taking the opportunities presented by such programs to their and their levels below
- Employee Satisfaction and Engagement
- Lead and drive engagement initiatives for the team
- Track team satisfaction scores and identify initiatives to build engagement within the team
- Proactively challenge the team with larger and enriching projects/ initiatives for the organization or team
- Exercise employee recognition and appreciation
͏
Deliver
| No | Performance Parameter | Measure |
| 1 | Operations of the tower | SLA adherence Knowledge management CSAT/ Customer Experience Identification of risk issues and mitigation plans Knowledge management |
| 2 | New projects | Timely delivery Avoid unauthorised changes No formal escalations |
