Job Description
Lead Site Reliability Engineer
Why This Role Is Important To Us
As a Lead Site Reliability Engineer, you will take on strategic and technical ownership within a Product Area, providing leadership across cloud infrastructure, service reliability, and platform operations. This role is ideal for a highly experienced engineer ready to guide others, shape our operational standards, and ensure SimCorp s cloud platform delivers excellence to both newly onboarded and long-standing clients.
You ll lead reliability-focused design sessions, drive the adoption of SRE principles, and champion automation and observability improvements. In addition, you will coach and mentor team members to grow technical capability across the Product Area while contributing to platform-wide initiatives.
What You Will Be Responsible For
- Own the reliability, scalability, and performance of Azure-hosted environments
- Lead operational support and onboarding activities for new and running client platforms
- Provide technical leadership in SRE practices, tooling, and incident response
- Guide the transition of manual processes into automated, scalable solutions
- Lead solutions workshops, engage with vendors, and manage senior stakeholders
- Implement and evolve observability frameworks using SLOs, SLIs, and proactive alerting
- Oversee disaster recovery planning, incident postmortems, and root cause analysis
- Ensure high-quality onboarding delivery through reusable automation pipelines
- Mentor and support the growth of junior and senior engineers in your Product Area
- Collaborate with product owners and architects to influence platform design and future roadmaps
- Provide weekend or on-call support as needed
What We Value
- Bachelor s or Master s degree in Computer Science or related field
- 5-8+ years in Site Reliability Engineering or Cloud Infrastructure leadership roles
- Strong expertise in Microsoft Azure, including production-grade design and operation
- Proficiency with IaC tools like Terraform, Bicep, ARM, Ansible
- Deep understanding of cloud-native monitoring and incident management frameworks
- Experience in leading incident response and platform-wide reliability improvements
- Leverage a strong foundation in ITIL practices, including problem, change, and incident management
- Broad technical knowledge: Kubernetes, Docker, APIs, CI/CD, scripting, SQL
- Experience with SimCorp Dimension or financial services platforms is a plus
- Proven ability to mentor and lead engineers, influence architecture decisions, and manage complexity
- Comfort balancing strategic priorities with hands-on execution
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.