
Pre-Doctoral Researcher
Atoof Shakir
Video Intelligence
Pre-Doctoral ResearcherAbout
Atoof Shakir works at the intersection of multimodal AI, representation learning, and large-scale human insight extraction. At ETH Zurich and the Agentic Systems Lab, he is developing systems that can interpret vast streams of real-world digital behavior — from short-form video to social media and news — and turn them into structured, actionable understanding. His work combines strong technical depth with an unusually applied orientation: rather than building models in isolation, he is interested in how multimodal agents can operate in noisy, dynamic environments where language, visuals, and human behavior constantly interact. Before joining ASL, Atoof contributed to state-of-the-art embedding models with broad real-world adoption, including models that have reached millions of downloads on Hugging Face. He has also worked on open-source projects spanning multimodal agentic systems, diffusion models, and real-time music generation. This background gives him a distinctive profile: someone equally comfortable with core model development and with designing systems that can produce tangible value at scale.
Publications
Research Areas
Project
Seldon: Advancing multimodal AI
Seldon builds high-quality data and evaluation infrastructure for the next generation of multimodal AI models. This helps to advance AI & robotic labs to test and improve models that need to understand images, long videos, audio, text, and tool use. Seldon is researching new ways to expose where these models fail on tasks that are easy for humans, especially in long videos, cross-modal alignment and real-world visual reasoning. This helps turn model failures into better training data, RL- environments, stronger evaluations, and more reliable multimodal AI systems.
Other team members
Pre-Doctoral Researcher
Interested in collaborating?
We are always looking for talented students, researchers and industry partners.







































