Muhammad Ferjad Naeem
Member of Technical Staff, Prometheus · Zürich, Switzerland
I work on building intelligence that can work with any kind of multimodal signal. At Prometheus I am building multimodal for the Artificial General Engineer (AGE).
Before that I spent five years at Google. I was part of Gemini Multimodal with the Astra team, where I worked across data, evaluation and post-training for proactive video understanding. Earlier, with Google Brain and later Google DeepMind, I built Google's strongest open-source multimodal encoders, SigLIP 2 and SILC, unifying self-supervised learning with image–text representation learning.
I completed my PhD at ETH Zürich, supported by a Google PhD Fellowship and advised by Luc Van Gool and Federico Tombari, with foundational work on open-set computer vision with language guidance.
Experience
-
2026 – present
-
2021 – 2026Research Scientist · GoogleZürich
-
2021 – 2024PhD in Machine Learning · Computer Vision Lab, ETH ZürichAdvised by Luc Van Gool and Federico Tombari · Google PhD Fellowship
Thesis: Towards Open-Set Computer Vision with Language Guidance.
-
Earlier
Computer Vision Intern, NVIDIA DriveIX (2020 – 2021) · Visiting Researcher in Zeynep Akata's Explainable ML group, University of Tübingen (2020 – 2021) · AI Research Intern, NAVER CLOVA AI Research (2019) · MSc Biomedical Computing and Research Fellow with Nassir Navab at CAMP, TU Munich (2018 – 2020).
Publications
Also on Google Scholar. * denotes equal contribution or core contributor.