Multimodal Vision Foundation Models Research Scientist

Posted 17 days ago

bytedanceSan Jose (CA)
Computer and Information Research ScientistsResearch and Development in the Physical, Engineering, and Life Sciences (except Nanotechnology and Biotechnology)

SENIORITY

Senior

Apply

About the role

ByteDance in San Jose seeks a PhD candidate for the Seed Vision team focused on visual generation models. Responsibilities include developing and scaling foundation models, optimizing architectures, and exploring real-world applications. Ideal candidates have excellent coding skills in C/C++ or Python, experience in computer vision or multimodal learning, and a strong understanding of deep learning methodologies. The position offers robust benefits including healthcare, a 401(k) plan, and generous paid time off. #J-18808-Ljbffr

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this

Multimodal Vision Foundation Models Research Scientist Jobs at...