Roland Memisevic

AI researcher  ·  Senior Director, Qualcomm AI Research

Formerly MILA / Université de Montréal  ·  CIFAR Fellow

Portrait of Roland Memisevic

About

I am an AI researcher and Senior Director at Qualcomm AI Research. I joined Qualcomm in 2021 through the acquisition of my startup Twenty Billion Neurons, where I was co‑founder, Chief Scientist (2016–2018) and CEO (2018–2021), developing real‑time vision‑language models.

Before co‑founding Twenty Billion Neurons, I was a faculty member in deep learning at the Université de Montréal and a founding member of MILA, the Montreal Institute for Learning Algorithms. Before that I was a faculty member in machine learning at the University of Frankfurt, and a postdoctoral fellow at ETH Zürich in Marc Pollefeys' group and at the University of Toronto. I received my PhD from the University of Toronto in 2008, working with Geoffrey Hinton, and was named a Fellow of the Canadian Institute for Advanced Research (CIFAR) in 2015.

My research interest is in understanding how AI systems could perceive and interact with the world with more human‑like common sense — particularly through end‑to‑end learning and situated, real‑time multimodal interaction. I have published over 100 research papers and patents, served on the program committees of the leading AI conferences (NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV), and co‑organized numerous workshops on deep learning since 2010 — including the CRM/CIFAR Deep Learning Summer School in Montréal with Yoshua Bengio and Yann LeCun.

Research interests

Selected publications

A selection of representative work across the years. For the full list, see Google Scholar →

  1. 2026
    Can vision-language models answer face-to-face questions in the real world?
    Pourreza, R., Dagli, R., Bhattacharyya, A., Panchal, S., Berger, G., Memisevic, R.
    International Conference on Learning Representations (ICLR 2026)
  2. 2026
    On the “induction bias” in sequence models
    Ebrahimi, M. R., Defferrard, M., Panchal, S., Memisevic, R.
    International Conference on Machine Learning (ICML 2026)
  3. 2025
    Revisiting bi-linear state transitions in recurrent neural networks
    Ebrahimi, M. R., Memisevic, R.
    Neural Information Processing Systems (NeurIPS 2025)
  4. 2024
    What to say and when to say it: Live fitness coaching as a testbed for situated interaction
    Panchal, S., Bhattacharyya, A., Berger, G., Mercier, A., Bohm, C., Dietrichkeit, F., Pourreza, R., Li, X., Madan, P., Lee, M., Todorovich, M., Bax, I., Memisevic, R.
    Neural Information Processing Systems (NeurIPS 2024)  ·  Datasets & Benchmarks Track  ·  continuation of work started at Twenty Billion Neurons
  5. 2024
    Look, remember and reason: Grounded reasoning in videos with language models
    Bhattacharyya, A., Panchal, S., Pourreza, R., Lee, M., Madan, P., Memisevic, R.
    International Conference on Learning Representations (ICLR 2024)
  6. 2024
    ClevrSkills: Compositional language and visual reasoning in robotics
    Haresh, S., Dijkman, D., Bhattacharyya, A., Memisevic, R.
    Neural Information Processing Systems (NeurIPS 2024)
  7. 2024
    AirLetters: An open video dataset of characters drawn in the air
    Dagli, R., Berger, G., Materzynska, J., Bax, I., Memisevic, R.
    European Conference on Computer Vision (ECCV 2024), HANDS Workshop  ·  video dataset for evaluating and learning motion features
  8. 2022
    Metaphors we learn by
    Memisevic, R.
    arXiv:2211.06441
  9. 2019
    The Jester dataset: A large-scale video dataset of human gestures
    Materzynska, J., Berger, G., Bax, I., Memisevic, R.
    International Conference on Computer Vision (ICCV 2019), HANDS Workshop  ·  widely used gesture-recognition benchmark
  10. 2017
    The “something something” video database for learning and evaluating visual common sense
    Goyal, R., Kahou, S. E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fründ, I., Yianilos, P., Mueller-Freitag, M., Hoppe, F., Thurau, C., Bax, I., Memisevic, R.
    International Conference on Computer Vision (ICCV 2017)
  11. 2014
    Modeling deep temporal dependencies with recurrent “grammar cells”
    Michalski, V., Memisevic, R., Konda, K.
    Neural Information Processing Systems (NIPS 2014)
  12. 2013
    Learning to relate images
    Memisevic, R.
    IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), special issue on Learning Deep Architectures
  13. 2012
    On multi-view feature learning
    Memisevic, R.
    International Conference on Machine Learning (ICML 2012)  ·  oral
  14. 2010
    Learning to represent spatial transformations with factored higher-order Boltzmann machines
    Memisevic, R., Hinton, G.
    Neural Computation, 22(6): 1473–1492  ·  cover article
  15. 2010
    Gated softmax classification
    Memisevic, R., Zach, C., Hinton, G., Pollefeys, M.
    Neural Information Processing Systems (NIPS 2010)
  16. 2007
    Unsupervised learning of image transformations
    Memisevic, R., Hinton, G. E.
    Computer Vision and Pattern Recognition (CVPR 2007)

See full publication list on Google Scholar →

Selected talks

Tutorials & short courses

Teaching

Grants & awards

Contact

roland.memisevic@gmail.com