# 2020-LRRo_data_set LRRo: A Lip Reading Data Set for the Under-resourced Romanian Language https://doi.org/10.5281/zenodo.3753559 If you make use of this collection, please acknowledge the work of the authors by citing the following publication: Andrei Cosmin Jitaru, Şeila Abdulamit, and Bogdan Ionescu. 2020. LRRo: A Lip Reading Data Set for the Under-resourced Romanian Language. In 11th ACM Multimedia Systems Conference (MMSys'20), June 8–11, 2020, Istanbul, Turkey. ACM, New York, NY, USA, 6 pages. https://doi.org/10.1145/3339825.3394932 Both datasets are distributed into three subsets, a training subset intended for training the models, a validation subset for validating and optimizing methods’ parameters, and a final testing subset for the actual evaluation. In Figure 1, the directory tree of both data sets is presented. In order to emphasis the differences between the data sets, we have presented the distribution of instances in each data set, according to Figure 2. Figure 1

Figure 2 ![](https://github.com/ajitaru/2020-LRRo_data_set/blob/master/img/samples_distrib.jpg) The following videos represent the raw recordings used for generating the data set's classes. Video 1 LAB data set

Video 2 WILD data set