eval_aishell-1: data: prompt: Transcribe the Chinese speech into text. data: - AISHELL-1/test/test_000{000..011}.tar.gz key: text repo_id: mispeech/xares_llm_data batch_size: 16 num_workers: 2 metric: iCER eval_clotho: data: data: - Clotho/test/test_0000{00..04}.tar.gz key: caption;captions prompt: Generate a short caption for the audio. repo_id: mispeech/xares_llm_data batch_size: 16 num_workers: 2 metric: FENSE eval_librispeech: data: data: - LibriSpeech/test/test-clean-0000{00..19}.tar.gz key: trans prompt: Transcribe the English speech into text repo_id: mispeech/xares_llm_data batch_size: 32 num_workers: 4 metric: iWER eval_mecat: data: data: - 0M0/test_0000-0000000.tar.gz - 000/test_0000-0000000.tar.gz - SMA/test_0000-0000000.tar.gz - S00/test_0000-0000000.tar.gz - S00/test_0001-0000000.tar.gz - 0MA/test_0000-0000000.tar.gz - S0A/test_0000-0000000.tar.gz - SM0/test_0000-0000000.tar.gz - SM0/test_0001-0000000.tar.gz - 00A/test_0000-0000000.tar.gz key: long prompt: Write a general audiocaption. repo_id: mispeech/MECAT-Caption batch_size: 64 num_workers: 4 metric: DATE eval_songdescriber: data: data: - Songdescriber/test/valid_song_describer_0000{00..01}.tar.gz key: captions prompt: Generate a caption for the music repo_id: mispeech/xares_llm_data batch_size: 16 num_workers: 0 metric: FENSE