OKX logo

Data Operations Engineer

OKX

On-site
Singapore, Singapore
Full-time
Senior
3+ yrs
Salary not listedPosted 1mo ago

Real job — pulled straight from OKX’s careers page · Verified July 14, 2026 · No reposts.

Job description

OKX is hiring a Data Operations Engineer — a full-time, based in Singapore, Singapore role. Apply directly on OKX's careers page below.

Big Data Engineer, Web3

Location: Singapore, Singapore

Department: Engineering

OKX will be prioritising applicants who have a current right to work in Singapore, and do not require OKX's sponsorship of a visa.

Who We Are

At OKX, we believe that the future will be reshaped by crypto, and ultimately contribute to every individual's freedom.
 
OKX is a leading crypto exchange, and the developer of OKX Wallet, giving millions access to crypto trading and decentralized crypto applications (dApps). OKX is also a trusted brand by hundreds of large institutions seeking access to crypto markets. We are safe and reliable, backed by our Proof of Reserves. 
 
Across our multiple offices globally, we are united by our core principles: We Before Me, Do the Right Thing, and Get Things Done. These shared values drive our culture, shape our processes, and foster a friendly, rewarding, and diverse environment for every OK-er.

OKX is part of OKG, a group that brings the value of Blockchain to users around the world, through our leading products OKX, OKX Wallet, OKLink and more.
 

About The Opportunity

We are building the foundational data infrastructure that powers one of the world's leading crypto exchanges. As a Big Data Platform Engineer, you will design, build, and evolve the core platform services that enable hundreds of data engineers and analysts to move fast and ship reliable pipelines. You will be at the forefront of our AI agent initiative — embedding LLM-driven capabilities directly into the platform layer so that scheduling, cost optimization, and incident response increasingly run autonomously.

What You’ll Be Doing 

  • Platform Core: Design and operate large-scale distributed data systems
  • Own the big data compute and storage infrastructure (MaxCompute/ODPS, Hologres, Spark)
  • Build and maintain multi-site task orchestration that dynamically selects engines and enforces policy
  • Drive reliability and performance improvements across batch and real-time pipelines

 

  • AI Integration: Build the AI-native platform layer
  • Develop and expose MCP (Model Context Protocol) tool interfaces so AI agents can interact with platform APIs
  • Build the scheduling and cost-optimization agents that auto-tune resource allocation and alert severity
  • Instrument platform telemetry to feed AI-driven SLA monitoring and anomaly detection
  • Design context retrieval pipelines (RAG / vector search) for SQL code and config knowledge bases

 

  • Tooling & DX: Evolve the developer experience
  • Own the internal data development platform — IDE integrations, code review automation, deployment tooling
  • Build APIs-first tools (backfill, ingestion automation) designed for future MCP integration
  • Collaborate with data warehouse and service teams to define platform contracts

 

  • Ops & Governance: Drive operational excellence
  • Establish SLA benchmarks, cost metrics, and latency dashboards as AI optimization targets
  • Build automated incident response and root-cause analysis pipelines
  • Define and enforce infrastructure policies across multi-cloud environments
AI Agent Ownership — Platform Tier
  • Scheduling Agent: auto-configure task dependencies, engine selection, cost/performance trade-offs, and alert tiers
  • Operations Agent: detect pipeline latency, performance degradation, and schema drift; trigger remediation
  • Incident Response Agent: trace SLA breaches to root cause, assign accountability, generate post-mortems
  • MCP Tool Layer: design and maintain the cross-platform tool interfaces that all agents call into

What We Look For In You 

  • 5+ years of experience building large-scale data platforms (Hadoop/Spark/Flink or equivalent)
  • Deep expertise in distributed storage and compute systems (MaxCompute, Hologres, ClickHouse, Hive)
  • Strong software engineering skills in Java, Scala, or Python; experience with API-first design
  • Hands-on experience with task scheduling systems (Airflow, DolphinScheduler, or in-house equivalents)
  • Solid understanding of multi-cloud architectures and cost governance
  • Familiarity with LLM integration patterns: tool calling, RAG pipelines, context management
  • Experience with MCP or similar agent-tool frameworks is a strong plus
  • Passion for building systems that make other engineers 10x more productive

Perks & Benefits 

  • Competitive total compensation package
  • L&D programs and education subsidy for employees' growth and development
  • Various team building programs and company events
  • Wellness and meal allowances
  • Comprehensive healthcare schemes for employees and dependants
  • More that we love to tell you along the process!

 

Notice:
All official OKX vacancies are published on this website. While roles may appear on selected third-party platforms from time to time, information on other sites may be inaccurate or outdated. If in doubt, please apply directly through our official careers website.
Information collected and processed as part of the recruitment process of any job application you choose to submit is subject to OKX's Candidate Privacy Notice.

Get Data Operations Engineer jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Mintlify logo

Senior Backend Engineer

$190k–$265kSan Francisco, CA
✓ From careers page· 21m ago
SandboxAQ logo

Senior Platform Engineer, Medical Devices

$122k–$228kRemote · US-eligible
✓ From careers page· 32m ago
Empower Pharmacy logo

Staff Cloud Infrastructure Architect (Remote)

Remote · US-eligible
✓ From careers page· 36m ago
Empower Pharmacy logo

Principal Technical Product Manager (Remote)

Remote · US-eligible
✓ From careers page· 36m ago

Frequently asked questions

What skills are required for Data Operations Engineer at OKX?

The required skills for Data Operations Engineer at OKX include: AWS, Databricks, Spark, Hive, SQL, Shell, Python, Prometheus, Grafana, CloudWatch, Datadog, Elasticsearch, Kubernetes, Terraform, Ansible, GitLab CI, Jenkins.

What is the seniority level for Data Operations Engineer at OKX?

Data Operations Engineer at OKX is a Senior level position.

How do I apply for Data Operations Engineer at OKX?

You can view the full description and apply for Data Operations Engineer at OKX on EchoJobs: https://echojobs.io/job/okx-data-operations-engineer-x-layer-u97dj.