Job Description
To be successful in this role you should have:
Proficiency in Programming and Scripting - This includes expertise in languages such as Python, Powershell, or Go, which are essential for automating routine tasks and system deployments.
Incident Management and Troubleshooting - The ability to manage incidents effectively, troubleshoot issues swiftly, and perform root cause analysis to prevent future incidents.
Systems Engineering and Automation - A deep understanding of systems engineering, including operating systems, networking, and cloud infrastructure. Proficiency in automation tools is crucial for maintaining system reliability at scale.
Influential Communication Skills - The ability to communicate effectively with team members and stakeholders, ensuring alignment, inspiring and motivating them to embrace new mindsets, cultures, and SRE working practices. This skill is crucial for driving meaningful change and fostering a collaborative environment where innovative ideas can thrive.
Some other highly valued skills include:
Knowledge of Cloud Computing - Familiarity with cloud platforms and services, which is increasingly important as more infrastructure moves to the cloud.
Strong Problem-Solving Abilities - The capability to approach problems methodically and find effective solutions, which is vital for maintaining system reliability.
You may be assessed on the key critical skills relevant for success in role, such as risk and controls, change and transformation, business acumen, strategic thinking and digital and technology, as well as job-specific technical skills.
The successful candidate can either be based in Knutsford (Radbroke Hall) or Glasgow Campus.
Purpose of the role
To apply software engineering techniques, automation, and best practices in incident response, to ensure the reliability, availability, and scalability of the systems, platforms, and technology through them.
Accountabilities
Availability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.
Resolution, analysis and response to system outages and disruptions, and implement measures to prevent similar incidents from recurring.
Development of tools and scripts to automate operational processes, reducing manual workload, increasing efficiency, and improving system resilience.
Monitoring and optimisation of system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.
Collaboration with development teams to integrate best practices for reliability, scalability, and performance into the software development lifecycle, and work closely with other teams to ensure smooth and efficient operations.
Stay informed of industry technology trends and innovations, and actively contribute to the organization's technology communities to foster a culture of technical excellence and growth.
💡 Quick Summary
Seeking a career-building opportunity? The Senior Site Reliability Engineer position is now open for candidates interested in the Bank Jobs sector. This role in Manchester offers a professional environment and growth potential.
Requirement Snapshot: Candidates should possess basic communication skills, a proactive attitude, and the ability to work in a team. Experience in Bank Jobs is a plus.