JAV-CG The 1st International Workshop on Joint Audio-Video Comprehension and Generation
A dedicated forum for the next wave of audio-video intelligence, spanning robust multimodal understanding, synchronized generation, and unified models that bridge perception and creation.
Joint AV intelligence.
Real-world multimedia is naturally audio-visual, yet many multimodal systems still treat sound as a side channel. JAV-CG focuses on models that listen, see, reason, and generate synchronized multimedia in one coherent framework.
The workshop brings together multimedia, audio and speech, computer vision, and multimodal foundation-model communities to sharpen the research agenda for audio-video comprehension and generation.
Timeline and milestones.
Workshop Paper Submission
16 July 2026Official ACM MM 2026 workshop contribution deadline.
Author Notification
06 August 2026Acceptance notification for workshop submissions.
Camera-Ready / Final Metadata
20 August 2026Final accepted material and metadata due under the ACM MM 2026 workshop schedule.
Author Registration
20 August 2026Registration deadline for accepted workshop contributions.
ACM Multimedia 2026
10-14 November 2026Conference venue: Rio de Janeiro, Brazil.
Workshop Day
Coming SoonFinal agenda and exact workshop day will be updated after logistics are confirmed.
Topics and submission.
JAV-CG welcomes archival workshop papers intended for the ACM MM 2026 workshop proceedings, as well as non-archival featured-paper submissions for workshop presentation. Technical, position, and perspective papers may be up to 8 pages plus references.
Audio-Visual Comprehension
- Sound source localization and source separation
- Audio-visual event detection and localization
- Question answering, grounding, and scene reasoning
- Trustworthy and long-form audio-visual understanding
Audio-Video Generation
- Video-to-audio and text-to-audio-video synthesis
- Audio-driven video generation and talking heads
- Foley, spatial audio, multimodal editing, and music
- Controllable synchronized generation across modalities
Unified AV Frameworks
- Any-to-any multimodal generation involving audio and video
- Joint tokenization, alignment, and representation learning
- Unified encoder-decoder or MLLM architectures
- Benchmarks, datasets, metrics, safety, and evaluation
Submission portal is live on OpenReview.
We will present a Best Paper Award to recognize outstanding workshop submissions.
Keynote speakers.
Talk titles and remaining speaker details will be updated as they are confirmed.
His research integrates computer vision, audition, and machine learning for multisensory perception, audio-visual scene understanding, audio-visual scene generation, accessibility, healthcare, and image/video processing.
Yapeng Tian is an Assistant Professor in the Computer Science Department at UT Dallas, where he leads the Computer Vision and Multimodal Computing Lab. His work has been recognized by the AAAI New Faculty Highlights, Cisco Faculty Research Award, and Amazon Research Award.
His research teaches machines to understand dynamic visual scenes through video, sound, and language, spanning computer vision, audio-visual learning, and trustworthy AI.
Chenliang Xu is a tenured Associate Professor of Computer Science at the University of Rochester and affiliated faculty of the Goergen Institute for Data Science and Artificial Intelligence. He received his Ph.D. from the University of Michigan in 2016 and was honored with the 2025 Edmund A. Hajim Outstanding Faculty Award.
More keynote speakers
Additional invited speakers will be announced soon.
Program schedule.
Detailed agenda, timing, and speaker assignments are TBD and will be updated once confirmed.
Workshop day program
Exact workshop day, session order, timings, speaker order, paper presentation slots, and breaks are TBD.
Meet the organizing team.
You Qin
National University of Singapore
Homepage
Kai Liu
Zhejiang University
HomepageShengqiong Wu
University of Oxford
HomepageJuncheng Li
Zhejiang University
Homepage
Wei Ji
Nanjing University
Homepage
Hao Fei
University of Oxford
Homepage
Liang Zheng
Australian National University
Homepage
Roger Zimmermann
National University of Singapore
Homepage
Jiebo Luo
University of Rochester
Homepage
Tat-Seng Chua
National University of Singapore
HomepageWorkshop correspondence
- You Qin qinyou@u.nus.edu
- Kai Liu kail@zju.edu.cn
- Shengqiong Wu shengqiongwu@gmail.com
- Hao Fei haofei7419@gmail.com