Big Data Engineer, Data Lake / Feature Store

Singapore

ByteDance

ByteDance is a technology company operating a range of content platforms that inform, educate, entertain and inspire people across languages, cultures and geographies.

View all jobs at ByteDance

Apply now Apply later

Responsibilities

About ByteDance
Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join Us
Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible.
Together, we inspire creativity and enrich life - a mission we aim towards achieving every day.
To us, every challenge, no matter how ambiguous, is an opportunity; to learn, to innovate, and to grow as one team. Status quo? Never. Courage? Always.
At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve.
Join us.

About The Team
The batch processing team is responsible for the company's offline data processing and distributed training, supporting various business scenarios such as offline ETL and machine learning within the company. The components involved include the offline computing engine Spark, the in-house distributed training framework Primus, feature storage solutions like Iceberg and Hudi, as well as Ray, a next-generation distributed application framework. Faced with massive-scale scenarios, extensive functional and performance optimizations have been carried out in Spark, Primus, Feature Store, and support for the adoption of the new-generation distributed application framework Ray in relevant company scenarios.

What you will be doing:
- Responsible for the development and performance optimisation of the in-house Feature Store functionality based on Iceberg;
- Participant in optimisation of the integration of Iceberg with various upper-level computing engines;
- Involve in platform-related infrastructure development.

Qualifications

Minimum Qualifications
- Bachelor's Degree or above, majoring in Computer Science, or related fields, with 2+ years of relevant development experience in the field with a strong programming ability, and proficiency in Java, Python, C++, with the ability to develop and optimize large-scale distributed systems.
- In-depth research and relevant experience in one or more data lake formats such as Delta, Hudi, or Iceberg.

Preferred Qualifications
- In-depth research or practical experience in open-source big data computing frameworks and scenarios like Hadoop, Spark, Flink, Presto, and more.

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

#LI-CT

Apply now Apply later

* Salary range is an estimate based on our AI, ML, Data Science Salary Index 💰

Job stats:  0  0  0

Tags: Big Data Computer Science Distributed Systems ETL Flink Hadoop Java Machine Learning Open Source Python Research Spark

Perks/benefits: Career development

Region: Asia/Pacific
Country: Singapore

More jobs like this