TikTok

Senior Machine Learning Engineer - Platform, Monetization Generative AI

San Jose, CA
Docker AWS Machine Learning PyTorch Spark Deep Learning Kubernetes GCP Azure
Description
TikTok is the leading destination for short-form mobile video. At TikTok, our mission is to inspire creativity and bring joy. TikTok's global headquarters are in Los Angeles and Singapore, and its offices include New York, London, Dublin, Paris, Berlin, Dubai, Jakarta, Seoul, and Tokyo.

Why Join Us
Creation is the core of TikTok's purpose. Our platform is built to help imaginations thrive. This is doubly true of the teams that make TikTok possible.
Together, we inspire creativity and bring joy - a mission we all believe in and aim towards achieving every day.
To us, every challenge, no matter how difficult, is an opportunity; to learn, to innovate, and to grow as one team. Status quo? Never. Courage? Always.
At TikTok, we create together and grow together. That's how we drive impact - for ourselves, our company, and the communities we serve.
Join us.

About the Generative AI Production Team
The Post-Training pod under Generative AI Production Team is at the forefront of refining and enhancing generative AI models for advertising, content creation, and beyond. Our mission is to take pre-trained models and fine-tune them to achieve state-of-the-art (SOTA) performance in vertical ad categories and multi-modal applications. We optimize models through fine-tuning, reinforcement learning, and domain adaptation, ensuring that AI-generated content meets the highest quality and relevance standards.

We work closely with pre-training teams, application teams, and multi-modal model developers (T2V, I2V, T2I) to bridge foundational AI advancements with real-world, high-performance applications. If you are passionate about pushing cognitive boundaries, optimizing AI models, and elevating AI-generated content to new heights, this is the team for you.

As a Machine Learning Platform Engineer, you will drive the development of our AI platform, ensuring scalability, efficiency, and robustness for training and serving large-scale diffusion models and multimodal generative AI systems. You will work closely with model researchers, infrastructure engineers, and data teams to optimize distributed training, inference efficiency, and production reliability.

Responsibilities
1) Architect and develop scalable and efficient AI infrastructure to support large-scale diffusion models and multi-modal generative AI workloads.
2) Optimize large model training and inference using PyTorch, Triton, TensorRT, and distributed training libraries (DeepSpeed, FSDP, vLLM).
3) Implement and optimize model using sequence parallelism, pipeline parallelism, and tensor parallelism etc to improve performance on high-throughput training clusters.
4) Scale and productionize generative AI models, ensuring efficient deployment on heterogeneous hardware environments (H100, A100, etc.).
5) Develop and integrate model distillation techniques to enhance the efficiency and performance of generative models, reducing computation costs while maintaining quality.
6) Design and maintain an automated model production pipeline for training/inference at scale, integrating distributed data processing frameworks (Ray, Spark, or custom solutions).
7) Enhance platform stability and efficiency by refining model orchestration, checkpointing, and retrieval strategies.
8) Collaborate with cross-functional teams (ML researchers, software engineers, infrastructure engineers) to ensure seamless model iteration cycles and deployments. Stay ahead of emerging trends in deep learning architectures, distributed training techniques, and AI infrastructure optimization, integrating best practices from academia and industry.Minimum Qualifications:
1) B.S., M.S., or Ph.D. in Computer Science, Electrical Engineering, or a related field. 5+ years of hands-on experience in large-scale machine learning infrastructure and distributed AI model training.
2) Deep expertise in PyTorch, CUDA optimization, and ML frameworks such as DeepSpeed, FSDP, and vLLM. Proven experience in optimizing diffusion models, sequence parallelism, and large-scale transformer-based architectures.
3) Strong understanding of high-performance computing, low-latency inference, and GPU acceleration techniques.
4) Hands-on experience in scaling AI infrastructure, leveraging Kubernetes, Docker, Ray, and Triton inference servers. Deep understanding of AI model orchestration, scheduling, and optimization across large clusters. Proficiency in profiling and debugging large-scale model training and inference bottlenecks.

Preferred Qualifications:
1) Experience deploying multi-modal generative AI models in production.
2) Expertise in compiler-level optimizations, TensorRT, and hardware-aware model tuning.
3) Familiarity with large-scale AI workloads in cloud environments (AWS, GCP, Azure).
4) Strong software engineering background, with a focus on scalability, efficiency, and reliability.

TikTok is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At TikTok, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

TikTok is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://shorturl.at/cdpT2
TikTok
TikTok

0 applies

0 views

There are more than 50,000 engineering jobs:

Subscribe to membership and unlock all jobs

Engineering Jobs

60,000+ jobs from 4,500+ well-funded companies

Updated Daily

New jobs are added every day as companies post them

Refined Search

Use filters like skill, location, etc to narrow results

Become a member

🥳🥳🥳 452 happy customers and counting...

Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.

To try it out

For active job seekers

For those who are passive looking

Cancel anytime

Frequently Asked Questions

  • We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
  • We've got over 200,000 jobs from 15,000+ vetted companies. No fake or sleazy jobs here!
  • We aggregate jobs from 15,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
  • We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
  • Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
  • Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
  • Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅

What Fellow Engineers Say