Idan Schwartz

Assistant Professor
Department of Computer Science and Artificial Intelligence
Bar-Ilan University

Leading a research group on multimodal learning

Idan Schwartz

I work on multimodal learning, mostly generative: aligning modalities such as vision, language, and audio. Much of my work centers on attention, from high-order and factor-graph attention modules for multimodal fusion to attention guidance in diffusion models. I am also interested in what cognitive science can tell us about deep learning, including perception and comprehension in multimodal models.

Before Bar-Ilan, I was a postdoc with Prof. Lior Wolf at Tel Aviv University. I received my PhD in Computer Science from the Technion, co-advised by Prof. Tamir Hazan and Prof. Alexander G. Schwing (UIUC); my thesis was on cognitive models in deep learning (pdf).

In industry, I currently serve as Chief Scientist at Aigency.ai. Previously, I worked on vision and language for eBay’s catalog, meeting insights in Microsoft’s assistant, and cloud workload prediction at Spot.

News

Selected work

CVPR 2026
TempoControl: Temporal Attention Guidance for Text-to-Video Models

S. Schiber, O. Lindenbaum, I. Schwartz

Steers cross-attention so concepts appear at the right moment in a generated video, without retraining.

Detection-driven object count teaser
WACV 2026
Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models

O. Zafar, Y. Cohen, L. Wolf, I. Schwartz

Optimizes a reusable counting token with detector feedback so generated images contain the requested number of objects.

LaMI teaser
ACL 2026
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion

G. Yariv, I. Schwartz, Y. Adi, S. Benaim

Generates multiple images from the prompt and fuses them with a frozen LLM at the last layer, adding visual commonsense without retraining.

TempoTokens teaser
AAAI 2024
Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

G. Yariv, I. Gat, S. Benaim, L. Wolf, I. Schwartz, Y. Adi

Maps audio into the input space of a frozen text-to-video model, generating videos temporally aligned with the sound.

Discriminative class tokens teaser
ICCV 2023
Discriminative Class Tokens for Text-to-Image Diffusion Models

I. Schwartz, V. Snæbjarnarson, S. Benaim, H. Chefer, R. Cotterell, L. Wolf, S. Belongie

Learns a single token from a classifier’s gradient to disambiguate fine-grained classes.

ZeroCap teaser
CVPR 2022
ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic

Y. Tewel, Y. Shalev, I. Schwartz, L. Wolf

Combines a frozen LM with CLIP at inference time to caption images, and even does visual arithmetic like image − word + image.

RobustViT teaser
NeurIPS 2022
Optimizing Relevance Maps of Vision Transformers Improves Robustness

H. Chefer, I. Schwartz, L. Wolf

Fine-tunes ViTs by shaping their relevance maps to focus on the foreground, with large robustness gains under distribution shift.

Factor Graph Attention teaser
CVPR 2019
Factor Graph Attention

I. Schwartz, T. Hazan, A. G. Schwing

A general attention mechanism that fuses any number of utilities (image, history, question, …) for visual dialog.

Publications

2026
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion
G. Yariv, I. Schwartz, Y. Adi, S. Benaim
ACL 2026  pdf / project
TempoControl: Temporal Attention Guidance for Text-to-Video Models
S. Schiber, O. Lindenbaum, I. Schwartz
CVPR 2026  arXiv / project
Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models
O. Zafar, Y. Cohen, L. Wolf, I. Schwartz
WACV 2026  pdf / arXiv / project
Single Image Iterative Subject-driven Generation and Editing
Y. Shpitzer, G. Chechik, I. Schwartz
P13N Workshop, CVPR 2026  pdf / project
2024
Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation
G. Yariv, I. Gat, S. Benaim, L. Wolf, I. Schwartz, Y. Adi
AAAI 2024  pdf / project
2023
Discriminative Class Tokens for Text-to-Image Diffusion Models
I. Schwartz, V. Snæbjarnarson, S. Benaim, H. Chefer, R. Cotterell, L. Wolf, S. Belongie
ICCV 2023  pdf / code / project
Zero-Shot Video Captioning with Evolving Pseudo-Tokens
Y. Tewel, Y. Shalev, R. Nadler, I. Schwartz, L. Wolf
BMVC 2023  arXiv / code
AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation
G. Yariv, I. Gat, S. Benaim, L. Wolf, I. Schwartz
INTERSPEECH 2023  pdf / project
2022
Optimizing Relevance Maps of Vision Transformers Improves Robustness
H. Chefer, I. Schwartz, L. Wolf
NeurIPS 2022  pdf / code
ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic
Y. Tewel, Y. Shalev, I. Schwartz, L. Wolf
CVPR 2022  pdf / code
Describing Sets of Images with Textual-PCA
O. Hupert, I. Schwartz, L. Wolf
Findings of EMNLP 2022  pdf / arXiv
Ordered Attention for Coherent Visual Storytelling
T. Braude, I. Schwartz, A. G. Schwing, A. Shamir
ACM Multimedia 2022  pdf
Latent Space Explanation by Intervention
I. Gat, G. Lorberbom, I. Schwartz, T. Hazan
AAAI 2022  pdf / arXiv
Video and Text Matching with Conditioned Embeddings
A. Ali, I. Schwartz, T. Hazan, L. Wolf
WACV 2022  pdf
2021
Perceptual Score: Measuring Perceptiveness of Multi-Modal Classifiers
I. Gat, I. Schwartz, A. G. Schwing
NeurIPS 2021  pdf / code
2020
Removing Bias in Multi-Modal Classifiers: Regularization by Maximizing Functional Entropies
I. Gat, I. Schwartz, A. G. Schwing, T. Hazan
NeurIPS 2020  pdf / code
2019
Factor Graph Attention
I. Schwartz, T. Hazan, A. G. Schwing
CVPR 2019  pdf / arXiv / code
A Simple Baseline for Audio-Visual Scene-Aware Dialog
I. Schwartz, A. G. Schwing, T. Hazan
CVPR 2019  pdf / code
2017
High-Order Attention Models for Visual Question Answering
I. Schwartz, A. G. Schwing, T. Hazan
NeurIPS 2017  pdf / code

Students

I lead a research group on multimodal generative models and attention at Bar-Ilan University. Students, research highlights, and openings for MSc and PhD students are on the group site.

Patents

Talks

Discriminative Models Can Make Generative Models Better (in Hebrew)

2025 · watch on YouTube

Multimodal Attention, Perception, Comprehension

2024 · watch on YouTube

Attention Models for Vision and Language

2019 · watch on YouTube

Contact

idanschwartz@gmail.com

Department of Computer Science and Artificial Intelligence, Room 213
Building 503, Bar-Ilan University
Ramat Gan, Israel  directions