USC 75-Speaker Annot(-16) Database

75-Speaker Annot(-16) is a comprehensive annotation dataset derived from the 75-Speaker vocal tract MRI database. This dataset provides phonetic alignments, articulator contour annotations, and handmade ground-truth articulator contours. Our annotation process integrates automated algorithms with expert verification to ensure accuracy and efficiency. To demonstrate its utility, we establish three benchmark tasks: speech phoneme recognition, articulatory contour segmentation, and articulatory phoneme recognition. Annot-16 can serve as a valuable resource for speech modeling, computer vision, and cross-modal learning, bridging engineering applications, speech science, and linguistic research.

A description of the database is given in the following article:

Xuan Shi*, Yubin Zhang*, Yijing Lu*, Marcus Ma*, Tiantian Feng*, Asterios Toutios, Haley Hsu, Louis Goldstein, Shrikanth S. Narayanan, “75-Speaker Annot-16: A benchmark dataset for speech articulatory rt-MRI annotation with articulator contours and phonetic alignment,” Interspeech 2025. (* equal contribution)


The database is publicly available on zenodo. Please contact Xuan Shi at xuanshi@usc.edu if you have any issues with data downloading.