Job Description
The Role
Join us as a Mainframe Site Reliability Engineer (SRE) and embark on an exciting journey of ensuring reliability, resiliency, and innovation in our information systems and ecosystems. As an SRE at Kyndryl, youll be at the forefront of driving continuous improvement and delivering exceptional service to our customers. Your role goes beyond traditional engineering, as youll have the opportunity to analyze business needs, tackle complex problems, and provide strategic advice and designs. Youll be involved in every stage of the software lifecycle, from building and testing to deploying changes and maintaining robust systems. Were looking for a true visionary who can think strategically and help shape the future of our services. Your expertise in building trusted relationships with customers and partnering with them for success will be instrumental in driving our growth. As an SRE, youll have the unique opportunity to work on end-to-end services, spanning customer sites and platforms. Collaboration and proactivity are key as you work alongside a talented team of professionals, eager to make a difference. Youll embrace an entrepreneurial mindset, taking ownership of your responsibilities and constantly seeking innovative solutions. With an unwavering focus on quality, robustness, and security, youll be a driving force in implementing cutting-edge tools that enhance our operations, improve reliability, and gather valuable feedback on our platforms. Your ability to identify and mitigate common operational issues will play a crucial role in delivering seamless experiences to our customers.
We are seeking a Mainframe Site Reliability Engineer (SRE) to join our infrastructure team. In this role, the engineer will bridge traditional mainframe systems programming with modern SRE practices and will serve as a key technical lead for the mainframe environment ensuring high availability, performance, and scalability for mission-critical workloads. As a Kyndryl representative, the engineer will act as the lead consultant during major incidents and architectural changes, driving service improvements and modernization initiatives, and strengthening platform stability through rigorous Root Cause Analysis (RCA) and automation.
Act as a consultant for all technical teams, advising on migrations/modernization upgrades, architectural changes, and DR exercises.
Participate in and document major incidents (P1/P2) and represent Kyndryl in technical discussions within customer forums.
Collaborate with internal stakeholders (z/OS, CICS, DB2, Automation, Storage, etc.) to drive technical solutions and fixes.
Conduct Root Cause Analysis (RCA) and blameless post-mortem reviews to prevent recurrence of incidents.
Analytical Mindset: A "detective" approach to system logs (LOGREC, SYSLOG) and dump analysis.
Proactive Ownership: A mindset focused on "Zero Outages" through predictive monitoring and robust system health checks.
Support internal and external audits by, addressing findings, and tracking mitigations to completion.
Identify performance bottlenecks and implement permanent fixes, proactively tracking action items to completion to ensure continuous service enhancement.
Suggest and implement service improvements, best practices, automation, and modernization opportunities, as well as alerting for seamless operations.
Follow SRE tenets applicable to mainframe platform.
Required Technical and Professional Expertise
10+ years of experience in operational management, including incident management and escalations
Expert level knowledge of z/OS system programming, including LPAR management, IPL processes, Parallel Sysplex, and JES2.
Extensive experience in troubleshooting, RCA, and conducting blameless post-mortems.
Good understanding of how to apply SRE principles to mainframe environments.
Good Knowledge of mainframe application development (COBOL, JCL, Rexx) and a strong understanding of application workflows and scheduling.
Fundamental understanding of CICS, IMS, and MQ transaction management workflows.
Fundamental understanding of DB2 databases.
Fundamental understanding of RACF security.
Fundamental understanding of IBM workload scheduler.
Fundamental understanding of storage(DS8K), VTS and other mainframe hardware that comprises the architecture.
Knowledge of host networking (TCP/IP, VTAM, SSH, and other standard mainframe protocols) and enterprise networking (VLAN, DHCP, DNS, firewalls, etc.).
Preferred Technical and Professional Experience
Experience implementing automation solutions on the mainframe platform
Exposure to mainframe modernization initiatives and platform optimization
Experience with predictive monitoring, proactive alerting, and system health checks
Ability to integrate traditional mainframe operations with modern reliability engineering practices
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
