I am Yuxuan Jiang (江宇轩), a PhD student at
Tsinghua University.
My research interests include multimodal learning, diffusion models,
and audio generation.
I study unified representations for speech, environmental sounds,
and music that capture acoustic details, semantic content, and temporal structure. I also explore
cross-modal and interaction-driven models that predict how sound evolves, enabling continuous and
controllable audio generation grounded in physical and structural priors.
I welcome conversations and collaborations on these and related topics, from informal exchanges to in-depth discussions. Please feel free to reach out by email.
© Yuxuan Jiang, 2026