Overview
Who this competition fits
This suits researchers and practitioners with NLP experience on Mandarin meeting transcripts, especially those able to model topic boundaries in long spoken-language documents. It is open to academia and industry, but Alibaba employees and Zhejiang University faculty/students are excluded.
Read the original official blurb
本项目为ICASSP2023 信号处理大挑战的通用会议理解及生成挑战赛(MUG challenge)。赛事构建并发布了目前为止规模最大的中文会议数据集,并基于会议人工转写结果进行了多项口语语言处理(SLP)任务的标注;目标是推动SLP在会议文本处理场景的研究并应对其中的多项关键挑战,包括 人人交互场景下多样化的口语现象、会议场景下的长篇章文档建模 等。 Official category: 算法竞赛. Participants: 76. Organizers: 阿里巴巴达摩院, 魔搭modelscope社区, 浙江大学.
Preparation
From registration to a first submission
- 01
Mandarin text processing
- 02
topic segmentation
- 03
long-document modeling
- 04
Transformer fine-tuning
- 05
ModelScope baseline usage
- 06
meeting transcript preprocessing
Before you commit: The core challenge is finding coherent topic boundaries in long, noisy spoken transcripts: the text highlights disfluencies, redundancies, grammar errors, coreference, colloquial language, and fragmented utterances, while sessions may contain several thousand words. The track is also constrained: participants may use only the specified corpus, permitted public pre-trained models, and listed public corpora.
Source
How this page was assembled
Competition information is structured from the official page. Scores are platform estimates for decision support; official rules take precedence.
- Official competition page
- ModelScope
- Last checked
- Dec 2, 2022