Residential College | false |
Status | 已發表Published |
PDSum: Prototype-driven Continuous Summarization of Evolving Multi-document Sets Stream | |
Yoon,Susik1; Chan,Hou Pong2; Han,Jiawei1 | |
2023-04-30 | |
Conference Name | 2023 World Wide Web Conference, WWW 2023 |
Source Publication | ACM Web Conference 2023 - Proceedings of the World Wide Web Conference, WWW 2023 |
Pages | 1650-1661 |
Conference Date | 2023/04/30-2023/05/04 |
Conference Place | Austin |
Abstract | Summarizing text-rich documents has been long studied in the literature, but most of the existing efforts have been made to summarize a static and predefined multi-document set. With the rapid development of online platforms for generating and distributing text-rich documents, there arises an urgent need for continuously summarizing dynamically evolving multi-document sets where the composition of documents and sets is changing over time. This is especially challenging as the summarization should be not only effective in incorporating relevant, novel, and distinctive information from each concurrent multi-document set, but also efficient in serving online applications. In this work, we propose a new summarization problem, Evolving Multi-Document sets stream Summarization (EMDS), and introduce a novel unsupervised algorithm PDSum with the idea of prototype-driven continuous summarization. PDSum builds a lightweight prototype of each multi-document set and exploits it to adapt to new documents while preserving accumulated knowledge from previous documents. To update new summaries, the most representative sentences for each multi-document set are extracted by measuring their similarities to the prototypes. A thorough evaluation with real multi-document sets streams demonstrates that PDSum outperforms state-of-the-art unsupervised multi-document summarization algorithms in EMDS in terms of relevance, novelty, and distinctiveness and is also robust to various evaluation settings. |
Keyword | Continuous Summarization Evolving Multi-document Sets Unsupervised Text Summarization |
DOI | 10.1145/3543507.3583371 |
URL | View the original |
Language | 英語English |
Scopus ID | 2-s2.0-85159281316 |
Fulltext Access | |
Citation statistics | |
Document Type | Conference paper |
Collection | University of Macau |
Affiliation | 1.University of Illinois at Urbana-Champaign,United States 2.University of Macau,Macao |
Recommended Citation GB/T 7714 | Yoon,Susik,Chan,Hou Pong,Han,Jiawei. PDSum: Prototype-driven Continuous Summarization of Evolving Multi-document Sets Stream[C], 2023, 1650-1661. |
APA | Yoon,Susik., Chan,Hou Pong., & Han,Jiawei (2023). PDSum: Prototype-driven Continuous Summarization of Evolving Multi-document Sets Stream. ACM Web Conference 2023 - Proceedings of the World Wide Web Conference, WWW 2023, 1650-1661. |
Files in This Item: | There are no files associated with this item. |
Items in the repository are protected by copyright, with all rights reserved, unless otherwise indicated.
Edit Comment