Job Description
About the Role
We're looking for a Sr Software Engineer to join the Platform Ops team (Platform Engineering) and enable developer productivity by building and improving developer tooling and environments, and by improving other productivity tools around observability. Our team builds the shared systems and workflows engineers rely on every day internal developer environments and self-serve tooling that makes the recommended way to ship changes the easiest way. This is a platform tooling role with a focus is on reusable capabilities and self-service workflows adopted across teams.
You'll also help develop AI-based workflows for productivity and efficiency, including automation that improves infrastructure operations (for example, faster triage, automated runbook steps, and better self-service for common issues). You'll prioritize the highest-impact problems by looking at where time is being lost today (slow builds, flaky tests, manual release steps, repetitive operational work) and turning that into concrete projects. You'll set clear success metrics up front and use them to guide design decisions and rollouts.
Expect to partner closely with engineering teams to drive adoption, document changes, and iterate based on feedback. You'll drive cross-team work across internal tooling, developer environments, observability, and reliability, focusing on measurable improvements (for example: shorter lead time, fewer incidents, and faster recovery). You'll add automation and enable developer workflows across all three major cloud environments (AWS, Azure, and GCP).
Over time, you'll own medium-sized projects end-to-end: identify the highest-friction workflows with partner teams, build and ship improvements, roll them out safely, and iterate based on adoption and measurable outcomes (for example: faster builds, fewer flaky tests, shorter time to debug, and less operational toil). You'll collaborate closely with Platform Engineering, Release Engineering, Security, and product engineering teams to make improvements that are easy to adopt and maintain.
What You'll Do
- Build internal tools and services in Python and/or Go
- Build tooling for our internal developer platform (self-service workflows, shared patterns, developer-facing interfaces)
- Add automation and enable developer workflows across all three major cloud environments (AWS, Azure, and GCP)
- Develop AI-based workflows that improve developer experience and reduce repetitive manual work, including efficiency gains in infrastructure operations
- Improve developer feedback loops through better observability (metrics/logs/traces), making failures easier to diagnose
- Automate common engineering workflows (scaffolding, environment setup, releases, operational runbooks)
- Improve observability for platform services (metrics, logs, traces) to make issues easier to detect and debug
- Integrate tools into existing workflows (CLI, Slack, GitHub, dashboards) where it reduces time-to-resolution
- Participate in on-call/incident response as needed and drive follow-up automation to prevent repeat issues
Skills We're Looking For
- Professional software engineering experience in Python and/or Go
- Experience building developer productivity tooling (CI systems, build tooling, developer portals, workflow automation)
- Experience with at least one cloud provider (AWS, Azure, or GCP)
- Strong fundamentals: testing, code review, debugging, and operating production services
- Ability to work effectively in a remote, cross-team environment
You'll Have An Edge If You Have
- Kubernetes experience
- Infrastructure-as-code (Terraform/OpenTofu), including reusable modules and safe state management
