Job Description
Key Responsibilities
Observability Leadership: Enhance telemetry collection and processing using OpenTelemetry, prioritizing actionable and cost-efficient metrics and traces.
Reliability Standards: Guide teams in defining and adopting SLIs/SLOs and foster a culture of service ownership.
Incident Management: Lead incident response efforts, facilitate post-incident reviews, and drive implementation of long-term solutions.
Infrastructure Automation: Use tools such as Pulumi, Terraform, or AWS CDK to manage cloud infrastructure and CI/CD pipelines.
Software Development: Create tools and automation in TypeScript (with optional Rust). Contribute to shared libraries and internal platforms.
Mentorship & Collaboration: Support and mentor other engineers, promoting a reliability-focused mindset across teams.
Continuous Improvement: Explore innovative tools and practices in observability and reliability; lead proof-of-concepts and improvement initiatives.
💡 Quick Summary
Seeking a career-building opportunity? The Site Reliability Engineer position is now open for candidates interested in the IT Engineer & Developer Jobs sector. This role in London offers a professional environment and growth potential.
Requirement Snapshot: Candidates should possess basic communication skills, a proactive attitude, and the ability to work in a team. Experience in IT Engineer & Developer Jobs is a plus.
