Assistant Professor
Department of Computer Science and Artificial Intelligence
Bar-Ilan University
Leading a research group on multimodal learning
Email· CV· Google Scholar· GitHub· X
I work on multimodal learning, mostly generative: aligning modalities such as vision, language, and audio. Much of my work centers on attention, from high-order and factor-graph attention modules for multimodal fusion to attention guidance in diffusion models. I am also interested in what cognitive science can tell us about deep learning, including perception and comprehension in multimodal models.
Before Bar-Ilan, I was a postdoc with Prof. Lior Wolf at Tel Aviv University. I received my PhD in Computer Science from the Technion, co-advised by Prof. Tamir Hazan and Prof. Alexander G. Schwing (UIUC); my thesis was on cognitive models in deep learning (pdf).
In industry, I currently serve as Chief Scientist at Aigency.ai. Previously, I worked on vision and language for eBay’s catalog, meeting insights in Microsoft’s assistant, and cloud workload prediction at Spot.
S. Schiber, O. Lindenbaum, I. Schwartz
Steers cross-attention so concepts appear at the right moment in a generated video, without retraining.

O. Zafar, Y. Cohen, L. Wolf, I. Schwartz
Optimizes a reusable counting token with detector feedback so generated images contain the requested number of objects.

G. Yariv, I. Schwartz, Y. Adi, S. Benaim
Generates multiple images from the prompt and fuses them with a frozen LLM at the last layer, adding visual commonsense without retraining.

G. Yariv, I. Gat, S. Benaim, L. Wolf, I. Schwartz, Y. Adi
Maps audio into the input space of a frozen text-to-video model, generating videos temporally aligned with the sound.

I. Schwartz, V. Snæbjarnarson, S. Benaim, H. Chefer, R. Cotterell, L. Wolf, S. Belongie
Learns a single token from a classifier’s gradient to disambiguate fine-grained classes.

Y. Tewel, Y. Shalev, I. Schwartz, L. Wolf
Combines a frozen LM with CLIP at inference time to caption images, and even does visual arithmetic like image − word + image.

H. Chefer, I. Schwartz, L. Wolf
Fine-tunes ViTs by shaping their relevance maps to focus on the foreground, with large robustness gains under distribution shift.

I. Schwartz, T. Hazan, A. G. Schwing
A general attention mechanism that fuses any number of utilities (image, history, question, …) for visual dialog.
I lead a research group on multimodal generative models and attention at Bar-Ilan University. Students, research highlights, and openings for MSc and PhD students are on the group site.
2025 · watch on YouTube
2024 · watch on YouTube
2019 · watch on YouTube
Department of Computer Science and Artificial Intelligence, Room 213
Building 503, Bar-Ilan University
Ramat Gan, Israel directions