OCBC Bank logo

Machine Learning Operations Lead

OCBC Bank

On-site
Singapore
Full-time
Manager
Principal
5+ yrs
Salary not listedPosted 4w ago

Real job — pulled straight from OCBC Bank’s careers page · Verified July 13, 2026 · No reposts.

Job description

OCBC Bank is hiring a Machine Learning Operations Lead — a full-time, based in Singapore role. Apply directly on OCBC Bank's careers page below.

Data Scientist, Platform AI Squad, Group Data Office (AVP/VP)

Location: OCBC Singapore

Remote Type: Onsite

Time Type: Full time

Job Description

WHO WE ARE:

As Singapore’s longest established bank, we have been dedicated to enabling individuals and businesses to achieve their aspirations since 1932. How? By taking the time to truly understand people. From there, we provide support, services, solutions, and career paths that meet their individual needs and desires.

 Today, we’re on a journey of transformation. Leveraging technology and creativity to become a future-ready learning organisation. But for all that change, our strategic ambition is consistently clear and bold, which is to be Asia’s leading financial services partner for a sustainable future.

 We invite you to build the bank of the future. Innovate the way we deliver financial services. Work in friendly, supportive teams. Build lasting value in your community. Help people grow their assets, business, and investments. Take your learning as far as you can. Or simply enjoy a vibrant, future-ready career.

Your Opportunity Starts Here.

Data Scientist, Platform AI Squad, Group Data Office (AVP/VP) 

About the Role

As a Data Scientist, Platform AI Squad, Group Data Office, you will design, optimize, and maintain the enterprise cloud infrastructure and software frameworks powering Enterprise AI across the bank. Operating at the intersection of AI application engineering, high-performance inference, and cloud platform engineering, you will work across Agentic AI frameworks and applications, low-latency LLM inference stacks leveraging hybrid-cloud services, and platform-level MLOps/DevOps pipelines.

 

In this role, you will establish standardized MLOps processes, optimize LLM inference, architect and build end-to-end AI solutions, maintain production pipelines, and drive enterprise AI engineering best practices. 

 

Key Responsibilities

1. Agentic AI Platforms & Frameworks

- Build and scale multi-agent orchestration frameworks (LangGraph, AutoGen, CrewAI) integrated with AWS.
- Design and implement reusable, enterprise-grade agentic platform components and applications.

- Implement function-calling harnesses and modern protocols (Model Context Protocol / MCP, A2A) connecting LLMs to databases (e.g., AWS Aurora, DynamoDB), vector stores, and enterprise APIs.
- Integrate managed cloud services (Amazon Bedrock Agents & Knowledge Bases) alongside custom open-source agent frameworks.
- Deploy automated evaluation pipelines (RAGAS, DeepEval, Bedrock Guardrails) to validate tool-calling precision, hallucination rates, and prompt safety before production release.

 

2. LLM Inference Optimization & Model Lifecycle

- Deploy and maintain low-latency LLM serving engines (vLLM, TensorRT-LLM, TGI, SGLang) across on-premises GPU clusters and cloud infrastructure.
- Establish GPU capacity planning, utilization tracking, monitoring, and budgeting processes.

- Oversee automated model versioning, artifact storage, and metadata tracking using MLflow or Amazon SageMaker Model Registry.
- Integrate foundation models via Amazon Bedrock and Amazon SageMaker AI for serverless scaling and hybrid workload routing.
- Package models into production-ready microservices supporting REST/gRPC APIs, batch processing, and streaming inference.
- Implement PagedAttention, Dynamic Batching, KV Cache offloading, and Speculative Decoding to minimize Time-to-First-Token (TTFT) and maximize token throughput.

 

3. AWS Cloud Architecture, MLOps & Daily DevOps Support

- Provide daily operational support for MLOps platforms, ensuring high availability (99.9%+ SLA), cluster stability, and rapid incident resolution.
- Build and maintain automated CI/CD workflows (GitHub Actions, GitLab CI, Bitbucket, Jenkins, AWS CodePipeline, ArgoCD) for model testing, containerization, feature store syncing, and release management.

- Provision multi-tenant enterprise infrastructure using Terraform, AWS CloudFormation/CDK, and Helm on Kubernetes (EKS).
- Configure telemetry (AWS CloudWatch, Prometheus, Grafana) for GPU tracking, latency SLAs, token costs, and data/model drift monitoring.

 

Experience & Background

-Education: Bachelor’s, Master’s, or Ph.D. in Computer Science, Data Science, Artificial Intelligence, or a quantitative discipline.

 

Work Experience:
- AVP Level (4+ years): Hands-on experience building LLM applications, microservices, containerized AWS deployments (EKS/Docker), and managing daily MLOps operations, have experience to drive a project from 0 -1. 

- VP Level (8+ years): Demonstrated track record architecting production AI platforms, LLM serving stacks, and multi-agent harnesses at enterprise scale, while leading platform engineering practices, have team leading experience.

 

Technical Skills

- AWS Cloud & AI Services: Core experience with Amazon Bedrock, SageMaker AI, EKS, EC2 GPU instances, S3, IAM, CloudWatch, and PrivateLink.
- Languages & Core AI Frameworks: Strong proficiency in Python. Deep hands-on experience with LangChain/LangGraph and LlamaIndex.
- Agentic Frameworks & Protocols: Practical experience with agentic workflows and protocol standards like MCP and Agent-to-Agent (A2A).
- Inference Stack & Compute: Proficiency with vLLM, TensorRT-LLM, Triton, Ray, MLflow, GPU memory management, and quantization techniques.
- DevOps & Infrastructure: Hands-on experience with Docker, Kubernetes (EKS), Terraform, Helm, Git, and GitOps tools (ArgoCD).
- Operational Mindset: Strong diagnostic skills for troubleshooting pipeline bottlenecks, container crashes, drift anomalies, and platform incidents.

What we offer:


Competitive base salary. A suite of holistic, flexible benefits to suit every lifestyle. Community initiatives. Industry-leading learning and professional development opportunities. Your wellbeing, growth and aspirations are every bit as cared for as the needs of our customers.

Get Machine Learning Operations Lead jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Cognyte logo

Data Science Team Leader

Herzliya, IL
✓ From careers page· 18m ago
Cognyte logo

Data Analyst

Bar Lev, IL
✓ From careers page· 18m ago
Altamira logo

Software Development Intern

McLean, VA
✓ From careers page· 20m ago

Frequently asked questions

What skills are required for Machine Learning Operations Lead at OCBC Bank?

The required skills for Machine Learning Operations Lead at OCBC Bank include: Python, Docker, Kubernetes, AWS, EKS, ECS, Lambda, S3, CloudWatch, IAM, CloudFormation, Terraform, CI/CD, MLflow, Jenkins, Bitbucket, Prometheus, Grafana, gRPC, REST.

What is the seniority level for Machine Learning Operations Lead at OCBC Bank?

Machine Learning Operations Lead at OCBC Bank is a Manager / Principal level position.

How do I apply for Machine Learning Operations Lead at OCBC Bank?

You can view the full description and apply for Machine Learning Operations Lead at OCBC Bank on EchoJobs: https://echojobs.io/job/ocbc-bank-machine-learning-ops-lead-vp-b6bd2.