Bio
I am a second-year Ph.D. student at Language Technologies Institute, Carnegie Mellon University, advised by Prof. Shinji Watanabe. My research interest mainly focuses on speech and language, and recently I have been interested in developing spoken language models.
Previously, I was a research assistant at Speech Processing Lab, National Taiwan University. I was also an R&D engineer at MediaTek Inc., where I designed and trained lightweight networks for super-resolution and frame-rate conversion (MEMC) that run on mobile devices in real time. I received the M.S. degree from National Taiwan University in 2021. During the time, I joined the Speech Processing Laboratory led by Prof. Lin-shan Lee and Prof. Hung-yi Lee.
Publications
See all publications on Google Scholar →
2026
Causal Tracing of Audio-Text Fusion in Large Audio Language Models
Interspeech 2026PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding
Findings of the Association for Computational Linguistics: ACL 2026Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions
Preprint 2026
2025
DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
Preprint 2025Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
The Thirteenth International Conference on Learning Representations 2025SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2025A Preliminary Exploration with GPT-4o Voice Mode
Preprint 2025
2024
Fusion Of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
IEEE Spoken Language Technology Workshop 2024Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024
2023
Prompting and Adapter Tuning for Self-supervised Encoder-Decoder Speech Model
IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) 2023
2022
Toward Degradation-Robust Voice Conversion
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2022
2021
Utilizing Self-supervised Representations for MOS Prediction
Interspeech 2021Investigating on Incorporating Pretrained and Learnable Speaker Representations for Multi-Speaker Multi-Style Text-to-Speech
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2021How Far Are We from Robust Voice Conversion: A Survey
IEEE Spoken Language Technology Workshop 2021Defending Your Voice: Adversarial Attack on Voice Conversion
IEEE Spoken Language Technology Workshop 2021
Honors
- TIGP-X Research Fellowship Academia Sinica, Taiwan 2026
- Graduate Thesis Award ACLCLP, Taiwan 2021
- 2nd Place, M2VoC Challenge IEEE ICASSP 2021
- Advanced Speech Technologies Scholarship (國立臺灣大學前瞻語音科技獎學金) National Taiwan University, Taiwan 2020
- Excellence Achievement in AI CUP Competition Ministry of Education, Taiwan 2020
