I started from unified multimodal models that represent and generate across vision, language, audio, and video. Now I study how such models become multi-agents that ground language in shared environments to communicate, plan, and act.

I am a Ph.D. student at UC Berkeley (BAIR), advised by Alane Suhr. Before Berkeley, I had wonderful experiences working with Mohit Bansal at UNC-NLP / MURGe-Lab and with Ziyi Yang at Microsoft. I did my undergrad at UNC Chapel Hill.

Publications

Multi-Agent & Embodied Interaction

02

Agents that communicate, cooperate, and act — grounding language in shared environments and social play.

Decoupling Planning and Control for Instructable Agents

Zineng Tang, Kelsey R. Allen, Sjoerd van Steenkiste, Ishita Dasgupta, Alane Suhr

COLM 2026

A VLM planner reasons and issues language subgoals; a learned low-level controller decides how, and for how long, to act on each one — decoupling instruction understanding from motor control.

Multimodal Learning & Generation

12

Models that connect vision, language, audio, and video — from unified representations to any-to-any generation.

TULIP: Towards Unified Language-Image Pretraining

Zineng Tang, Long Lian, Seun Eisape, XuDong Wang, Roei Herzig, Adam Yala, Alane Suhr, Trevor Darrell, David M. Chan

ICCV Workshops 2025

An open-source drop-in replacement for CLIP-style encoders that learns fine-grained visual features while preserving global semantic alignment.

arXiv code project

Others

01

Excursions beyond the two threads above.

Continuous Language Generative Flow

Zineng Tang, Shiyue Zhang, Hyounghun Kim, Mohit Bansal

ACL 2021

Normalizing flows as a continuous latent generative model for language generation and semi-supervised learning.

paper code

Teaching & Service
  • organizingLSEI Workshop, co-located with COLM 2026 (San Francisco, Oct 9)
  • teachingGraduate Student Instructor, CS 288 — Natural Language Processing, UC Berkeley
Awards
Personal

Outside of research, I spend my time on music production and video games.