Hello there, I'm Dengming Zhang, a Ph.D. student at Shanghai Jiao Tong University, advised by Prof. Mingda Chen. Previously, I received my master's degree from Zhejiang University and worked at Tencent as a Large Language Models (LLMs) Post-Training Engineer.
My current research focuses on World Models and Agentic AI, especially JEPA-based WAM and VLA. My previous work explores Multimodal Large Models under low-data and low-compute constraints. On the low-data side, I study how to merge multiple domain-specialized, fine-tuned expert LLMs into one generalist model using only 1–5 samples, while retaining SOTA-level performance (ICLR 2026). On the low-compute side, I explore how to equip vision foundation models with audio using a single RTX 4090, and improve audio-visual affective understanding to a SOTA-level. I am also interested in Generative AI (Image/Music), Affective Computing, Meta-learning, and HCI.
By the way, I am good at combining scientific research with engineering implementation, and I have rich experience in front-end development, back-end development, and cluster devops. Some of the open source projects that I lead/participate in can be found on my GitHub.