Site Reliability Engineer III -(AIML SRE)

💰 $8,960 - $14,336 (Est.) 📍 New York City

Job Description

Full job description
JOB DESCRIPTION

Are you looking for an exciting opportunity to join a dynamic and growing team in a fast paced and challenging area? This is a unique opportunity for you to work in our team to partner with the Business to provide a comprehensive view.

As a Senior AI Reliability Engineer at JPMorgan Chase within the Technology and Operations division, you will join our dynamic team of innovators and technologists. Your mission will be to enhance the reliability and resilience of AI systems that revolutionize how the Bank services and advises clients. You will focus on ensuring the robustness and availability of AI models, deepening client engagements, and promoting process transformation. We seek team members passionate about leveraging advanced reliability engineering practices, AI observability, and incident response strategies to solve complex business challenges through high-quality, cloud-centric software delivery.

Job Responsibilities:

Define and refine Service Level Objectives (SLOs) for large language model serving and training systems, using metrics like accuracy, fairness, latency, drift targets, TTFT, and TPOT, while balancing reliability and development velocity.
Design, implement, and continuously improve monitoring systems to track availability, latency, drift, and other key metrics for robust observability and rapid issue detection.
Collaborate in the design and deployment of high-availability language model serving infrastructure that supports high-traffic internal workloads across multiple regions and cloud providers.
Champion site reliability engineering practices, providing technical leadership and fostering a culture of reliability, resilience, and continuous improvement across teams.
Develop and manage automated failover and recovery systems for model serving deployments, ensuring seamless operation and rapid recovery from failures.
Create and lead AI-specific incident response playbooks for issues like model drift or bias spikes, including automated rollbacks, circuit breakers, and systematic post-incident improvements.
Build and maintain cost optimization systems for large-scale AI infrastructure, leveraging load balancing, caching, optimized GPU scheduling, and AI Gateways to ensure efficient, secure, and scalable operations.
Required qualifications, capabilities, and skills:

Formal training or certification on AI reliability concepts and 3+ years applied experience.
Demonstrate a strong sense of curiosity and a passion for continuous learning, especially in the rapidly evolving field of AI reliability.
Show proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices.
Possess deep knowledge and experience in observability, including white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, and Splunk.
Be proficient with continuous integration and delivery tools like Jenkins, GitLab, or Terraform, as well as container and orchestration technologies such as ECS, Kubernetes, and Docker.
Have experience troubleshooting common networking technologies and issues, and understand the unique challenges of operating AI infrastructure, including model serving, batch inference, and training pipelines.
Communicate effectively and bridge the gap between ML engineers and infrastructure teams, with proven experience implementing and maintaining SLO/SLA frameworks for business-critical services, and working with both traditional and AI-specific metrics.
Preferred qualifications, capabilities, and skills

Experience with AI-specific observability tools and platforms, such as OpenTelemetry and OpenInference.
Familiarity with AI incident response strategies, including automated rollbacks and AI circuit breakers.
Knowledge of AI-centric SLOs/SLAs, including metrics like accuracy, fairness, drift targets, TTFT (Time To First Token), and TPOT (Time Per Output Token).
Expertise in engineering for scale and security, including load balancing, caching, optimized GPU scheduling, and AI Gateways.
Experience with continuous evaluation processes, including pre-deployment, pre-release, and post-deployment monitoring for drift and degradation.
Understand ML model deployment strategies and their reliability implications
Have contributed to open-source infrastructure or ML tooling
Have experience with chaos engineering and systematic resilience testing
#LI-ID1

ABOUT US
JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management.
We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.

We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, ****** orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.

JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans




ABOUT THE TEAM

J.P. Morgan Asset & Wealth Management delivers industry-leading investment management and private banking solutions. Asset Management provides individuals, advisors and institutions with strategies and expertise that span the full spectrum of asset classes through our global network of investment professionals. Wealth Management helps individuals, families and foundations take a more intentional approach to their wealth or finances to better define, focus and realize their goals.

💡 Quick Summary

Seeking a career-building opportunity? The Site Reliability Engineer III -(AIML SRE) position is now open for candidates interested in the Software Developer Jobs sector. This role in New York City offers a professional environment and growth potential.

Requirement Snapshot: Candidates should possess basic communication skills, a proactive attitude, and the ability to work in a team. Experience in Software Developer Jobs is a plus.

Sponsored

Job Details

Company Name: JPMorganChase

Frequently Asked Questions

Click the Apply Now button on this page, login or register for free on CallCenterJob.co.in, fill in your name, mobile number, city, and experience, then submit your application. The recruiter will contact you directly.
The expected salary for Site Reliability Engineer III -(AIML SRE) in New York City is $8,960 - $14,336 (Est.) per month. Actual compensation may vary based on experience and negotiation.
No, Site Reliability Engineer III -(AIML SRE) is an on-site position based in New York City. Candidates must be able to commute or relocate to this location.
Basic communication skills, a proactive attitude, and the ability to work in a team are required for Site Reliability Engineer III -(AIML SRE). Previous experience in Software Developer Jobs is a plus. Freshers may also apply depending on the employer's requirements.
Yes, CallCenterJob.co.in is completely free for job seekers. Never pay money to apply for any job. If anyone asks for payment to process your application, report it immediately using the "Report this Job" button.

Similar Openings

  • Software Developer

    Jaipur SALARY: ₹60,000 - ₹80,000 POSTED: 20 hours ago CATEGORY: Education DEADLINE: May 26, 2026 SKILLS: Software Development LANGUAGES: English, Hindi SOFTWARE DEVELOPER Greetings from Jobingo HR Solutions Pvt. Ltd…!!! This is to inform you we curre...

    Full Time / Part Time

    Salary Estimated: 20K to 32K

    Jaipur, Rajasthan

    August 4, 2026


    Apply Now

  • Dot Net full stack developer

    Position : Dot Net Full Stack Developer Experience: 5 to 6 Yr Location: Mumbai, Pune, Bangalore, Chennai Mode: Work from Office Notice: Immediate Joiners / 7 Days Looking for an experienced, product-oriented backend developer to join our software eng...

    Full Time / Part Time

    Salary Estimated: 21K to 26K

    Mumbai, Maharashtra

    August 4, 2026


    Apply Now

  • Backend Engineer | Python |Gurgram |BSE Listed NBFC

    Note • Applicants who fit the below Job Criteria will only be contacted by TeamUpShoot • Rest, we appreciate your interest and you will only be be contacted by TeamUpShoot if any other opportunity suitable is available with us About The Employer • Th...

    Full Time / Part Time

    Salary Estimated: 16K to 29K

    Gurugram , Haryana

    August 4, 2026


    Apply Now

  • Dotnet Developer

    Shift Timings 12:30 to 10:30 pm / 1:30 pm to 11:30 pm Mandatory Skills- .NET and Azure Relevant Experience- 8-10 years Work Location- Bangalore/Hyderabad/Noida (Temporary Remote) • Relevant years of experience required is 8-12 years. 1-2 years of tea...

    Full Time / Part Time

    Salary Estimated: 21K to 23K

    Bangalore, Karnataka

    August 4, 2026


    Apply Now

  • Java Software Engineer

    Key Responsibilities: • Contribute to build key components, middleware of the platform and developing fully multi-tenant systems • Contribute to develop workflow management functions • Develop REST APIs, as well as contribute to the overall API frame...

    Full Time / Part Time

    Salary Estimated: 17K to 25K

    Remote

    August 4, 2026


    Apply Now

  • Software Developer

    Senior Engineer This local Software as a Service company seeks an experienced Senior Engineer to drive innovation and streamline processes. Key Responsibilities Design, develop, and maintain Java-based software solutions Requirements Technical expert...

    Full Time / Part Time

    Salary Estimated: 25K to 33K

    Chicago, Illinois

    August 4, 2026


    Apply Now