USC Long Single-Speaker (LSS) Dataset
The USC Long Single-Speaker (LSS) dataset contains real-time MRI video of the vocal tract dynam- ics and simultaneous audio obtained during speech pro- duction. This unique dataset contains roughly one hour of video and audio data from a single native speaker of American English, making it one of the longer publicly available single-speaker datasets of real-time MRI speech data. Along with the articulatory and acoustic raw data, we release derived representations of the data that are suit- able for a range of downstream tasks. This includes video cropped to the vocal tract region, sentence-level splits of the data, restored and denoised audio, and regions-of-interest timeseries.
A description of the database is given in the following article:
Sean Foley, Jihwan Lee, Kevin Huang, Xuan Shi, Yoonjeong Lee, Louis Goldstein, and Shrikanth Narayanan, “A long- form single-speaker real-time mri speech dataset,” Submitted to ICASSP2026.
To access/download please begin by filling out this form. Please contact Sean Foley at seanfole@usc.edu if you have any issues with data downloading.
