Job Summary:
ByteDance is dedicated to pioneering new paths toward artificial general intelligence. The Research Scientist role involves developing multimodal foundation models and optimizing world models for reasoning, planning, and interaction, contributing to the advancement of multimodal assistant products.
Responsibilities:
• Develop multimodal foundation models integrating vision, language, audio, and environment signals.
• Design and optimize world models for reasoning, planning, and interaction.
• Build training pipelines including data curation, alignment, and reinforcement learning.
• Improve agent capabilities such as perception, memory, decision-making, and tool use.
• Explore next-generation interaction paradigms between humans and intelligent systems.
Qualifications:
Required:
• Currently pursuing a PhD in computer science, mathematics, engineering, or a related field, with an expected graduation date in 2027 and the ability to commit to an onboarding date by the end of 2027.
• Excellent coding ability, data structures, and fundamental algorithm skills, proficient in C/C++ or Python, etc.
• Experience in multimodal learning, reinforcement learning, or agent systems.
• Familiarity with large-scale model training or simulation environments.
Preferred:
• Strong research track record in relevant areas.
• Strong problem-solving and collaboration skills.
Company:
ByteDance is a technology company that develops content creation platforms and services. Founded in 2012, the company is headquartered in Beijing, CHN, with a team of 10001+ employees. The company is currently Late Stage.