I started from unified multimodal models that represent and generate across vision, language, audio, and video. Now I study how such models become multi-agents that ground language in shared environments to communicate, plan, and act.

I am a Ph.D. student at UC Berkeley (BAIR), advised by Alane Suhr. Before Berkeley, I had wonderful experiences working with Mohit Bansal at UNC-NLP / MURGe-Lab and with Ziyi Yang at Microsoft. I did my undergrad at UNC Chapel Hill.

Publications

Multi-Agent & Embodied Interaction

02

Agents that communicate, cooperate, and act — grounding language in shared environments and social play.

Decoupling Planning and Control for Instructable Agents

Zineng Tang, Kelsey R. Allen, Sjoerd van Steenkiste, Ishita Dasgupta, Alane Suhr

COLM 2026

A VLM planner reasons and issues language subgoals; a learned low-level controller decides how, and for how long, to act on each one — decoupling instruction understanding from motor control.

Multimodal Learning & Generation

12

Models that connect vision, language, audio, and video — from unified representations to any-to-any generation.

Others

01

Excursions beyond the two threads above.

Teaching & Service
  • organizingLSEI Workshop, co-located with COLM 2026 (San Francisco, Oct 9)
  • teachingGraduate Student Instructor, CS 288 — Natural Language Processing, UC Berkeley
Awards