Portrait of Muhammad Ferjad Naeem

Muhammad Ferjad Naeem

Member of Technical Staff, Prometheus · Zürich, Switzerland

I work on building intelligence that can work with any kind of multimodal signal. At Prometheus I am building multimodal for the Artificial General Engineer (AGE).

Before that I spent five years at Google. I was part of Gemini Multimodal with the Astra team, where I worked across data, evaluation and post-training for proactive video understanding. Earlier, with Google Brain and later Google DeepMind, I built Google's strongest open-source multimodal encoders, SigLIP 2 and SILC, unifying self-supervised learning with image–text representation learning.

I completed my PhD at ETH Zürich, supported by a Google PhD Fellowship and advised by Luc Van Gool and Federico Tombari, with foundational work on open-set computer vision with language guidance.

Experience

  1. 2026 – present
    Member of Technical Staff · Prometheus
    Zürich

    Physical AI.

  2. 2021 – 2026
    Research Scientist · Google
    Zürich
    • 2024 – 2026 Gemini Multimodal.
    • 2023 – 2024 Foundational multimodal encoders: SigLIP 2 and SILC, Google's strongest open-source vision–language encoders.
    • 2021 – 2023 Google PhD Fellow / Research Consultant. Zero-shot image classification with long-form text and LLM-generated synthetic data.
  3. 2021 – 2024
    PhD in Machine Learning · Computer Vision Lab, ETH Zürich
    Advised by Luc Van Gool and Federico Tombari · Google PhD Fellowship

    Thesis: Towards Open-Set Computer Vision with Language Guidance.

  4. Earlier

    Computer Vision Intern, NVIDIA DriveIX (2020 – 2021) · Visiting Researcher in Zeynep Akata's Explainable ML group, University of Tübingen (2020 – 2021) · AI Research Intern, NAVER CLOVA AI Research (2019) · MSc Biomedical Computing and Research Fellow with Nassir Navab at CAMP, TU Munich (2018 – 2020).

Publications

Also on Google Scholar. * denotes equal contribution or core contributor.

2026

2025

2024

2023

2022

2021

2020

2019

2018

2017