Residential College | false |
Status | 已發表Published |
Word Segmentation by Separation Inference for East Asian Languages | |
Tong, Yu; Guo, Jingzhi; Zhou, Jizhe; Chen, Ge; Zhen, Guokai | |
2022-05 | |
Conference Name | ACL 2022 |
Volume | Findings of the Association for Computational Linguistics: ACL 2022 |
Pages | 3924–3934 |
Conference Date | May 2021 |
Conference Place | ACL | Findings, Dublin |
Country | Ireland |
Publisher | ACL Anthology |
Abstract | Chinese Word Segmentation (CWS) intends to divide a raw sentence into words through sequence labeling. Thinking in reverse, CWS can also be viewed as a process of grouping a sequence of characters into a sequence of words. In such a way, CWS is reformed as a separation inference task in every adjacent character pair. Since every character is either connected or not connected to the others, the tagging schema is simplified as two tags “Connection” (C) or “NoConnection” (NC). Therefore, bigram is specially tailored for “C-NC” to model the separation state of every two consecutive characters. Our Separation Inference (SpIn) framework is evaluated on five public datasets, is demonstrated to work for machine learning and deep learning models, and outperforms state-of-the-art performance for CWS in all experiments. Performance boosts on Japanese Word Segmentation (JWS) and Korean Word Segmentation (KWS) further prove the framework is universal and effective for East Asian Languages. |
DOI | 10.18653/v1/2022.findings-acl.309 |
URL | View the original |
Indexed By | EI |
WOS ID | WOS:000828767404002 |
The Source to Article | https://aclanthology.org/2022.findings-acl.309.pdf |
Scopus ID | 2-s2.0-85149131122 |
Fulltext Access | |
Citation statistics | |
Document Type | Conference paper |
Collection | Faculty of Science and Technology DEPARTMENT OF COMPUTER AND INFORMATION SCIENCE |
Affiliation | University of Macau |
First Author Affilication | University of Macau |
Recommended Citation GB/T 7714 | Tong, Yu,Guo, Jingzhi,Zhou, Jizhe,et al. Word Segmentation by Separation Inference for East Asian Languages[C]:ACL Anthology, 2022, 3924–3934. |
APA | Tong, Yu., Guo, Jingzhi., Zhou, Jizhe., Chen, Ge., & Zhen, Guokai (2022). Word Segmentation by Separation Inference for East Asian Languages. , Findings of the Association for Computational Linguistics: ACL 2022, 3924–3934. |
Files in This Item: | There are no files associated with this item. |
Items in the repository are protected by copyright, with all rights reserved, unless otherwise indicated.
Edit Comment