Roland Memisevic
AI researcher · Senior Director, Qualcomm AI Research
Formerly MILA / Université de Montréal · CIFAR Fellow
About
I am an AI researcher and Senior Director at Qualcomm AI Research. I joined Qualcomm in 2021 through the acquisition of my startup Twenty Billion Neurons, where I was co‑founder, Chief Scientist (2016–2018) and CEO (2018–2021), developing real‑time vision‑language models.
Before co‑founding Twenty Billion Neurons, I was a faculty member in deep learning at the Université de Montréal and a founding member of MILA, the Montreal Institute for Learning Algorithms. Before that I was a faculty member in machine learning at the University of Frankfurt, and a postdoctoral fellow at ETH Zürich in Marc Pollefeys' group and at the University of Toronto. I received my PhD from the University of Toronto in 2008, working with Geoffrey Hinton, and was named a Fellow of the Canadian Institute for Advanced Research (CIFAR) in 2015.
My research interest is in understanding how AI systems could perceive and interact with the world with more human‑like common sense — particularly through end‑to‑end learning and situated, real‑time multimodal interaction. I have published over 100 research papers and patents, served on the program committees of the leading AI conferences (NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV), and co‑organized numerous workshops on deep learning since 2010 — including the CRM/CIFAR Deep Learning Summer School in Montréal with Yoshua Bengio and Yann LeCun.
Research interests
Selected publications
A selection of representative work across the years. For the full list, see Google Scholar →
-
Can vision-language models answer face-to-face questions in the real world?International Conference on Learning Representations (ICLR 2026)
-
On the “induction bias” in sequence modelsInternational Conference on Machine Learning (ICML 2026)
-
Revisiting bi-linear state transitions in recurrent neural networksNeural Information Processing Systems (NeurIPS 2025)
-
Look, remember and reason: Grounded reasoning in videos with language modelsInternational Conference on Learning Representations (ICLR 2024)
-
ClevrSkills: Compositional language and visual reasoning in roboticsNeural Information Processing Systems (NeurIPS 2024)
-
AirLetters: An open video dataset of characters drawn in the airEuropean Conference on Computer Vision (ECCV 2024), HANDS Workshop · video dataset for evaluating and learning motion features
-
The Jester dataset: A large-scale video dataset of human gesturesInternational Conference on Computer Vision (ICCV 2019), HANDS Workshop · widely used gesture-recognition benchmark
-
The “something something” video database for learning and evaluating visual common senseInternational Conference on Computer Vision (ICCV 2017)
-
Modeling deep temporal dependencies with recurrent “grammar cells”Neural Information Processing Systems (NIPS 2014)
-
Learning to relate imagesIEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), special issue on Learning Deep Architectures
-
On multi-view feature learningInternational Conference on Machine Learning (ICML 2012) · oral
-
Learning to represent spatial transformations with factored higher-order Boltzmann machinesNeural Computation, 22(6): 1473–1492 · cover article
-
Unsupervised learning of image transformationsComputer Vision and Pattern Recognition (CVPR 2007)
Selected talks
- Expo Talk Panel on embodied AI, NeurIPS 2025 · San Diego
- Invited talk, CVPR Workshop on Computer Vision in Sports (CVsports) · Nashville
- Invited talk, MILA Workshop on NLP in the Era of Generative AI · Montréal
- Qualcomm webinar on multimodal AI
- Panelist on generative AI at the edge, Embedded Vision Summit · Santa Clara
- Invited talk, KLA-Tencor · Chennai, India
- Invited talk, Embedded Vision Summit (virtual)
- Invited tutorial, INIT/AERFAI Summer School on Machine Learning · Benicàssim, Spain
- Invited talk, Research and Applied AI Summit (RAAIS) · London
- Invited talk, ISCAS 2017 · Baltimore
- Invited talk, Re-Work Deep Learning Summit · San Francisco
- Invited talk, MIT Deep Learning Workshop · Boston
- Keynote, ICISP 2016 · Trois-Rivières
- Invited tutorial, HotChips · Cupertino
- Invited tutorial, ICIP (with Yoshua Bengio) · Québec City
- Invited tutorial, Training Workshop on Deep Architectures in Vision and NL · Wrocław, Poland
- Invited lecture, CRM/CIFAR Deep Learning Summer School · Montréal
- Co-organizer, CRM/CIFAR Deep Learning Summer School (with Yoshua Bengio, Yann LeCun) · Montréal
- Invited talks at Google, University of Toronto, KLA-Tencor, ICML Deep Learning workshop
- Keynote, ICLR 2014 · Banff
- Co-organizer, NIPS 2014 workshop on Deep Learning and Representation Learning
- Invited tutorial, CIFAR NCAP Summer School · Toronto
- Tutorial on multi-view feature learning, CVPR 2012 · Providence
- Invited tutorial, IPAM Graduate Summer School on Deep Learning · UCLA
- Invited tutorial, DAGM 2011 · Frankfurt
Tutorials & short courses
-
Deep learningINIT/AERFAI Summer School on Machine Learning · Benicàssim, Spain
-
Deep Learning Summer School 2015 — full lecture archive27 lectures from the summer school I co-organized with Yoshua Bengio and Yann LeCun · Montréal
-
Deep learning in image processing and visionICIP tutorial, with Yoshua Bengio · Québec City
-
Deep architectures in vision and NLTraining workshop · Wrocław, Poland
-
Multi-view feature learningCVPR tutorial · Providence
Teaching
- Machine Learning for Vision (IFT 6268) · Université de Montréal (materials →)
- Machine Learning for Vision (IFT 6268) · Université de Montréal (materials →)
- Databases (IFT 2821) · Université de Montréal (materials →)
- Fundamentals of Machine Learning (IFT 3395/6390) · Université de Montréal (materials →)
- Machine Learning for Vision (IFT 6268) · Université de Montréal (materials →)
- Visual Feature Learning (IFT 6085) · Université de Montréal (materials →)
- Machine Learning · University of Frankfurt (materials →)
- Digital Image Processing (in German) · University of Frankfurt
- Teaching Assistant · CSC 321 (Neural Networks & Machine Learning), CSC 2515 (Machine Learning), CSC 207 (Software Design) · University of Toronto
Grants & awards
- Fellow, Canadian Institute for Advanced Research (CIFAR)
- FQRNT Team Research Project (with co-PIs Y. Bengio, P. Vincent, A. Courville) · ~$150k CAD
- DARPA seedling project on deep learning for time series (PI, with Y. Bengio) · ~$180k USD
- Facebook research gift
- Google Faculty Research Award
- NSERC Discovery Grant · $100k CAD (five years)
- NSERC Strategic Grant, in collaboration with Ubisoft (with Y. Bengio, P. Vincent, A. Courville) · $575k CAD
- Government of Canada Award