Thanh V. T. Tran

avatar.jpg

I am a first-year Ph.D. student in the School of Electrical and Electronic Engineering at Nanyang Technological University (NTU), advised by Professor Woon-Seng Gan. Previously, I spent three years at FPT Software – AI Center as an AI Research Resident, where I worked under the supervision of Dr. Van Nguyen and Professor Truong-Son Hy.

I’m always open to collaborations, discussions, and new opportunities. Feel free to reach out if you’re interested in my research or would like to discuss potential projects.

Research: My research focuses on spatial audio signal processing and machine learning, particularly in the areas of auditory scene analysis and 3D sound event localization and detection.

News

Aug 10, 2026 Started my Ph.D. at Nanyang Technological University (NTU), Singapore, advised by Prof. Woon-Seng Gan.
Jun 17, 2026 Flowley got accepted at ECCV 2026, wrapping up my journey at FPT Software – AI Center.
Jun 04, 2026 DiFlow-TTS got accepted at Interspeech 2026 (Long Paper track).
May 01, 2026 DiFlowDubber got accepted at CVPR Findings 2026. DiFlowDubber and Flowley also got accepted at Sight and Sound Workshop, CVPR 2026.
Jan 10, 2026 Honored to receive the Best Performance Award 2025, ranking in the top 3 out of 100+ AI engineers and researchers at FPT Software – AI Center.
May 20, 2025 RESOUND got accepted at Interspeech 2025.
Dec 21, 2024 ConxGNN got accepted at ICASSP 2025.
Nov 17, 2024 GROOT got accepted at KDD 2025.

Selected Publications

  1. Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space
    Thanh V. T. Tran, Ngoc-Son Nguyen, Luong Tran, and 4 more authors
    European Conference on Computer Vision, 2026
  2. DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow Matching
    Ngoc-Son Nguyen, Thanh V. T. Tran, Hieu-Nghia Huynh-Nguyen, and 2 more authors
    Interspeech, 2026
  3. Effective Context Modeling Framework for Emotion Recognition in Conversations
    Cuong Tran Van*, Thanh V. T. Tran*, Van Nguyen, and 1 more author
    International Conference on Acoustics, Speech, and Signal Processing, 2025
  4. KDD
    GROOT: Effective Design of Biological Sequences with Limited Experimental Data
    Thanh V. T. Tran*, Nhat Khang Ngo*, Viet Anh Nguyen, and 1 more author
    Conference on Knowledge Discovery and Data Mining, 2025