# 2nd-Challenge-Workshop-on-Multilingual-Conversational-Speech-Language-Model- The 2nd Multilingual Conversational Speech Language Models Challenge 2026 is now open for registration. This year’s challenge focuses on Speech LLMs for real-world multilingual conversational speech, covering speaker diarization, speech recognition, acoustic understanding, and semantic understanding. Top-performing teams will share a total prize pool of USD 20,000. Registration is free, and the dataset will be provided free of charge to registered participants. ## Motivation Recent advances in Large Language Models (LLMs) have accelerated the development of Speech LLMs, enabling unified modeling of speech recognition and spoken language understanding. However, progress in this area critically depends on the availability of real-world multilingual conversational speech data that reflect the complexity of natural human communication. The first edition of this workshop demonstrated the importance of such data for advancing Speech LLMs in automatic speech recognition and speaker diarization. The challenge results showed that Speech LLMs have achieved strong performance on speech recognition, suggesting that transcription-centric modeling has been largely addressed. In contrast, speaker diarization remains a key open challenge, with performance still limited in complex multilingual and conversational scenarios. These findings indicate that further progress requires moving beyond transcription accuracy toward deeper modeling of conversational structure and speech content. Motivated by these observations, the second edition of this workshop focuses on advancing Speech LLMs in speaker diarization, acoustic understanding, and semantic understanding. To support this goal, the workshop will release a broader and more diverse multilingual conversational speech dataset, expanding language coverage and conversational scenarios. By encouraging research that jointly models who is speaking, how speech is acoustically realized, and what semantic information is conveyed, this workshop aims to promote holistic spoken language understanding and to drive the next stage of progress in multilingual Speech LLM research. ## Task Setting and Evaluation ### Task 1: Multilingual Conversational Speech Diarization and Recognition No prior or oracle information will be provided during evaluation (e.g., no pre-segmented utterances or speaker labels). Objective: Develop a system for both speaker diarization (identifying who is speaking when), and recognition (transcribing speech to text). Both pipeline-based and end-to-end systems are encouraged, providing flexibility in system design and implementation. Performance will be assessed based on the Diarization Error Rate (DER) and the concatenated minimum permutation WER or CER, referred to as tcpWER or tcpCER. The DER is employed to determine the best speaker ID permutation between oracle annotation and diarization results. Then, the recognition results and references belonging to the same speaker within a recording will be concatenated to calculate the tcpWER or tcpCER. All submissions will be ranked according the tcpWER or tcpCER. ### Task 2: Multilingual Conversational Speech Understanding No prior or oracle information will be provided during evaluation (e.g., no pre-segmented utterances or speaker labels). Objective: Develop a system for acoustic and semantic understanding of multilingual conversation. Both pipeline-based and end-to-end systems are encouraged, providing flexibility in system design and implementation. Evaluate the system's ability to understand the entire conversation in the form of choice questions. ## Important Dates (AOE Time) March 30, 2026: Registration opens April 10, 2026: Training data release April 24, 2026: Development set and baseline system release June 15, 2026: Evaluation set release and leaderboard open June 25, 2026: Leaderboard freeze and paper submission portal opens (CMT system) July 10, 2026: Paper submission deadline July 20, 2026: Notification of acceptance October 2, 2026: Workshop date ## Dataset Description ### Training set The training set (Train) comprises approximately 14 languages: English (en), French (fr), German (de), Italian (it), Portuguese (pt), Spanish (es), Japanese (jp), Korean (ko), Russian (ru), Thai (th), Vietnamese (vi), Tagalog (tl), Urdu (ur),Turkish (tr). Each recording consists of two-speaker conversational speech on randomly assigned topics. Conversations are natural and fluent, with speakers engaging in meaningful dialogues on each topic. Recorded in quiet indoor environments using devices such as iPhones. Each recording will provide the oracle segmentation and speaker label for the development of speech recognition and speaker diarization systems. Both Task I and Task II share the same training set. The English dataset comprises approximately 500 hours of recordings from various regions, including British, American, Australian, Indian, and Philippine English. Other languages contribute around 100 hours each, resulting in a total of approximately 2100 hours of multilingual conversational speech data. This dataset is designed to provide a rich resource for training and evaluating multilingual conversational speech language models (MLC-SLM), addressing the challenges of linguistic diversity, speaker variability, and contextual understanding. ### Development set The development set (Dev) has the same setting as the training set but contains approximately 4 hours of recordings for each language. Both Task I and Task II share the same development set. ### Evaluation set Will be released soon. ## Rules All participants must adhere to the following rules to be eligible for the challenge. Use of Large Language Models: For Tasks 1 and 2, all systems must be built based on large language models, speech large language models, or multimodal large language models. Use of External Resource: For Tasks 1 and 2, the use of external datasets and pre-trained models (including speech foundation models and LLMs) is permitted. All external resources utilized must be freely accessible to all research groups and should be clearly indicated in the final system report. Data augmentation: Data augmentation is allowed on the released training set and may include, but is not limited to, the addition of noise or reverberation, speed perturbation, and tone modification. Prohibition of Evaluation Sets Usage: The use of evaluation sets in any form of non-compliance is strictly prohibited. This includes, but is not limited to, using evaluation sets for fine-tuning or training the model. Multi-System Fusion: Participants are NOT allowed to employ system fusion in either Task 1 and Task 2. Submitted results must be derived from a single model or system rather than through result fusion. Submission Requirement: All participations are required to submit their system. The submission may include final results, models and a Docker that can directly perform inference to obtain the final results, etc. Detailed submission instructions will be provided following the release of the baseline implementation. Please note that we will publicly disclose the name of teams and their affiliated institutions that confirmed participation but did not submit any files. Organizer's Interpretation: The organizers reserve the right to make the final interpretation of these rules. In special circumstances, the organizers will coordinate the interpretation as needed. ## Other Topics In addition to challenge system descriptions, participants are encouraged to submit research papers that showcase innovative findings, practical case studies, and forward-looking ideas. Topics of interest include, but are not limited to: Novel Architectures and Algorithms: Development of new architectures and algorithms for training Speech Large Language Models. Audio Data Processing Pipelines: Innovative pipelines for processing raw audio data that facilitate the collection of diverse internet data for training Speech Large Language Models. Natural and Emotionally Rich Speech Generation: Algorithms designed to generate more natural and emotionally expressive conversational speech for dialogue systems. Leveraging Multi-Turn Conversational History: Approaches that utilize multi-turn conversational history to enhance diarization and understanding results. Evaluation Techniques and Benchmarks: Innovative evaluation techniques or benchmarks specifically tailored for assessing Speech Large Language Models. New Datasets: Creation of new datasets, both real and synthetic, for training Speech Large Language Models. ## Data Access and Usage Registered participants will be granted access to the training and test datasets. Registration for the MLC-SLM Challenge constitutes acceptance of the Data Use Agreement (see below), agree to confidentiality and comply with the data protection agreement. The datasets will only be used for the purpose of the workshop challenge, and redistribution or any other use is strictly prohibited. It is the responsibility of the participant to protect the data from unauthorized access. [Data use agreement](https://www.nexdata.ai/nexdata/static/file/doc/Data%20use%20agreement2026.docx) ## Registration Please register your team using the registration form: [Register for MLC-SLM Challenge 2026](https://forms.gle/LfyNjfjBV3eYwzdp8) The challenge registration begins on March 30, 2026. By registering, you agree to the [Data use agreement](https://www.nexdata.ai/nexdata/static/file/doc/Data%20use%20agreement2026.docx). We welcome both academic and industry teams. Individual researchers are also encouraged to participate. For any other information about registration, please send Email to: mlc-slmw@nexdata.ai More details:https://www.nexdata.ai/competition/mlc-slm ## Baseline System Will be released soon. ## Leaderboard Submission Will be released soon. ## Prize Pool Prizes for Top-Ranking Teams in this Competition(each task): 1st Place:$5,000 2nd Place:$3,000 3rd Place:$2,000 ## Organizers Lei Xie, Northwestern Polytechnical University Shuai Wang, Nanjing University Liumeng Xue, Nanjing University Eng Siong Chng, Nanyang Technological University Hung-yi Lee, National Taiwan University Xie Chen, Shanghai Jiaotong University Khalid Choukri, European Language Resources Association Qiangze Feng, Nexdata Daliang Wang, Nexdata Longshuai Xiao, Huawei Technologies Hexin Liu, Nanyang Technological University Bingshen Mu, Northwestern Polytechnical University Zhennan Lin, Northwestern Polytechnical University