Bar-Ilan University · Computer Science and Artificial Intelligence

The Multimodal Lab

We study how machines perceive, reason, and act across modalities.
Led by Idan Schwartz.

About

We build models that connect vision, language, audio, and video, with a focus on generation. A recurring theme is attention: designing modules that fuse many modalities, and using the attention inside large generative models as a handle on what they produce. We are also drawn to questions from cognitive science, such as how perception, memory, and comprehension arise in machines, and increasingly to models that reason and act in the physical world.

Research

Multimodal attention

Attention modules and pooling that fuse vision, language, audio, and video.

Inference-time optimization

Guidance, search, and learned tokens that steer frozen generative models.

Efficient generative training

Representation learning that lets generative models reach faithfulness to data with less compute.

Self-evolving models

Models that keep improving after deployment as distributions shift, with test-time training and tool use.

Physical AI

Generative and world models for agents acting in the physical world.

Embodied agents

Perception, navigation, and action in interactive environments.

Visual reasoning

Models that answer, explain, and think over images and video.

Memory models

Architectures that store and recall over long contexts and long videos.

Reinforcement learning

Learning from interaction and feedback to align and control models.

Work

CVPR 2026
TempoControl: Temporal Attention Guidance for Text-to-Video Models

S. Schiber, O. Lindenbaum, I. Schwartz

Steers cross-attention so concepts appear at the right moment in a generated video.

Detection-driven object count teaser
WACV 2026
Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models

O. Zafar, Y. Cohen, L. Wolf, I. Schwartz

A reusable counting token, optimized with detector feedback, makes generated images match the requested count.

SISO teaser
P13N Workshop, CVPR 2026
Single Image Iterative Subject-driven Generation and Editing

Y. Shpitzer, G. Chechik, I. Schwartz

Personalizes a diffusion model from a single subject image, iterating between generation and editing.

LaMI teaser
ACL 2026
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion

G. Yariv, I. Schwartz, Y. Adi, S. Benaim

Generates images from the prompt and fuses them with a frozen LLM at the last layer.

TempoTokens teaser
AAAI 2024
Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

G. Yariv, I. Gat, S. Benaim, L. Wolf, I. Schwartz, Y. Adi

Maps audio into the input space of a frozen text-to-video model, keeping video aligned with sound.

Discriminative class tokens teaser
ICCV 2023
Discriminative Class Tokens for Text-to-Image Diffusion Models

I. Schwartz, V. Snæbjarnarson, S. Benaim, H. Chefer, R. Cotterell, L. Wolf, S. Belongie

Learns a single token from a classifier’s gradient to disambiguate fine-grained classes.

ZeroCap teaser
CVPR 2022
ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic

Y. Tewel, Y. Shalev, I. Schwartz, L. Wolf

A frozen LM plus CLIP captions images at inference time, and even does visual arithmetic.

RobustViT teaser
NeurIPS 2022
Optimizing Relevance Maps of Vision Transformers Improves Robustness

H. Chefer, I. Schwartz, L. Wolf

Shaping relevance maps to focus on the foreground brings large robustness gains.

Factor Graph Attention teaser
CVPR 2019
Factor Graph Attention

I. Schwartz, T. Hazan, A. G. Schwing

A general attention mechanism that fuses any number of utilities for visual dialog.

The complete list is on Google Scholar and Idan’s page.

People

Idan Schwartz · Principal Investigator

PhD students

MSc students

Alumni

Contact

idanschwartz@gmail.com

Department of Computer Science and Artificial Intelligence, Room 213
Building 503, Bar-Ilan University
Ramat Gan, Israel  directions