Tech Lead Machine Learning Engineer - Platform, Monetization Generative AI
New York, New York, United States
TikTok is the leading destination for short-form mobile video. At TikTok, our mission is to inspire creativity and bring joy. TikTok's global headquarters are in Los Angeles and Singapore, and its offices include New York, London, Dublin, Paris, Berlin, Dubai, Jakarta, Seoul, and Tokyo.
Why Join Us
Creation is the core of TikTok's purpose. Our platform is built to help imaginations thrive. This is doubly true of the teams that make TikTok possible.
Together, we inspire creativity and bring joy - a mission we all believe in and aim towards achieving every day.
To us, every challenge, no matter how difficult, is an opportunity; to learn, to innovate, and to grow as one team. Status quo? Never. Courage? Always.
At TikTok, we create together and grow together. That's how we drive impact - for ourselves, our company, and the communities we serve.
Join us.
About the Generative AI Production Team
The Post-Training pod under Generative AI Production Team is at the forefront of refining and enhancing generative AI models for advertising, content creation, and beyond. Our mission is to take pre-trained models and fine-tune them to achieve state-of-the-art (SOTA) performance in vertical ad categories and multi-modal applications. We optimize models through fine-tuning, reinforcement learning, and domain adaptation, ensuring that AI-generated content meets the highest quality and relevance standards.
We work closely with pre-training teams, application teams, and multi-modal model developers (T2V, I2V, T2I) to bridge foundational AI advancements with real-world, high-performance applications. If you are passionate about pushing cognitive boundaries, optimizing AI models, and elevating AI-generated content to new heights, this is the team for you.
As a Machine Learning Platform Engineer, you will drive the development of our AI platform, ensuring scalability, efficiency, and robustness for training and serving large-scale diffusion models and multimodal generative AI systems. You will work closely with model researchers, infrastructure engineers, and data teams to optimize distributed training, inference efficiency, and production reliability.
Responsibilities
1) Architect and develop scalable and efficient AI infrastructure to support large-scale diffusion models and multi-modal generative AI workloads.
2) Optimize large model training and inference using PyTorch, Triton, TensorRT, and distributed training libraries (DeepSpeed, FSDP, vLLM).
3) Implement and optimize model using sequence parallelism, pipeline parallelism, and tensor parallelism etc to improve performance on high-throughput training clusters.
4) Scale and productionize generative AI models, ensuring efficient deployment on heterogeneous hardware environments (H100, A100, etc.).
5) Develop and integrate model distillation techniques to enhance the efficiency and performance of generative models, reducing computation costs while maintaining quality.
6) Design and maintain an automated model production pipeline for training/inference at scale, integrating distributed data processing frameworks (Ray, Spark, or custom solutions).
7) Enhance platform stability and efficiency by refining model orchestration, checkpointing, and retrieval strategies.
8) Collaborate with cross-functional teams (ML researchers, software engineers, infrastructure engineers) to ensure seamless model iteration cycles and deployments. Stay ahead of emerging trends in deep learning architectures, distributed training techniques, and AI infrastructure optimization, integrating best practices from academia and industry.
Why Join Us
Creation is the core of TikTok's purpose. Our platform is built to help imaginations thrive. This is doubly true of the teams that make TikTok possible.
Together, we inspire creativity and bring joy - a mission we all believe in and aim towards achieving every day.
To us, every challenge, no matter how difficult, is an opportunity; to learn, to innovate, and to grow as one team. Status quo? Never. Courage? Always.
At TikTok, we create together and grow together. That's how we drive impact - for ourselves, our company, and the communities we serve.
Join us.
About the Generative AI Production Team
The Post-Training pod under Generative AI Production Team is at the forefront of refining and enhancing generative AI models for advertising, content creation, and beyond. Our mission is to take pre-trained models and fine-tune them to achieve state-of-the-art (SOTA) performance in vertical ad categories and multi-modal applications. We optimize models through fine-tuning, reinforcement learning, and domain adaptation, ensuring that AI-generated content meets the highest quality and relevance standards.
We work closely with pre-training teams, application teams, and multi-modal model developers (T2V, I2V, T2I) to bridge foundational AI advancements with real-world, high-performance applications. If you are passionate about pushing cognitive boundaries, optimizing AI models, and elevating AI-generated content to new heights, this is the team for you.
As a Machine Learning Platform Engineer, you will drive the development of our AI platform, ensuring scalability, efficiency, and robustness for training and serving large-scale diffusion models and multimodal generative AI systems. You will work closely with model researchers, infrastructure engineers, and data teams to optimize distributed training, inference efficiency, and production reliability.
Responsibilities
1) Architect and develop scalable and efficient AI infrastructure to support large-scale diffusion models and multi-modal generative AI workloads.
2) Optimize large model training and inference using PyTorch, Triton, TensorRT, and distributed training libraries (DeepSpeed, FSDP, vLLM).
3) Implement and optimize model using sequence parallelism, pipeline parallelism, and tensor parallelism etc to improve performance on high-throughput training clusters.
4) Scale and productionize generative AI models, ensuring efficient deployment on heterogeneous hardware environments (H100, A100, etc.).
5) Develop and integrate model distillation techniques to enhance the efficiency and performance of generative models, reducing computation costs while maintaining quality.
6) Design and maintain an automated model production pipeline for training/inference at scale, integrating distributed data processing frameworks (Ray, Spark, or custom solutions).
7) Enhance platform stability and efficiency by refining model orchestration, checkpointing, and retrieval strategies.
8) Collaborate with cross-functional teams (ML researchers, software engineers, infrastructure engineers) to ensure seamless model iteration cycles and deployments. Stay ahead of emerging trends in deep learning architectures, distributed training techniques, and AI infrastructure optimization, integrating best practices from academia and industry.
* Salary range is an estimate based on our AI, ML, Data Science Salary Index 💰
Job stats:
0
0
0
Tags: Architecture Content creation Deep Learning Diffusion models FSDP Generative AI Generative modeling Machine Learning ML infrastructure Model training PyTorch Reinforcement Learning Spark TensorRT vLLM
Perks/benefits: Career development
Region:
North America
Country:
United States
More jobs like this
Explore more career opportunities
Find even more open roles below ordered by popularity of job title or skills/products/technologies used.
BI Developer jobsSr. Data Engineer jobsData Engineer II jobsBusiness Intelligence Analyst jobsPrincipal Data Engineer jobsStaff Data Scientist jobsStaff Machine Learning Engineer jobsData Science Manager jobsData Manager jobsPrincipal Software Engineer jobsData Science Intern jobsBusiness Data Analyst jobsJunior Data Analyst jobsData Analyst Intern jobsSoftware Engineer II jobsData Specialist jobsSr. Data Scientist jobsLead Data Analyst jobsDevOps Engineer jobsResearch Scientist jobsStaff Software Engineer jobsAI/ML Engineer jobsData Engineer III jobsSenior Backend Engineer jobsBI Analyst jobs
Git jobsAirflow jobsOpen Source jobsEconomics jobsLinux jobsKafka jobsComputer Vision jobsJavaScript jobsGoogle Cloud jobsMLOps jobsNoSQL jobsKPIs jobsTerraform jobsData Warehousing jobsPhysics jobsRDBMS jobsPostgreSQL jobsScikit-learn jobsBanking jobsHadoop jobsScala jobsGitHub jobsData warehouse jobsStreaming jobsPandas jobs
R&D jobsClassification jobsBigQuery jobsOracle jobsDistributed Systems jobsCX jobsPySpark jobsdbt jobsScrum jobsReact jobsLooker jobsRAG jobsMicroservices jobsJira jobsRobotics jobsRedshift jobsSAS jobsIndustrial jobsData Mining jobsPrompt engineering jobsNumPy jobsGPT jobsELT jobsMySQL jobsData strategy jobs