REVISTA
Information

TODAS

Inicio / Information / Vol: 11 Par: 2 (2020) / Art�culo

ART�CULO

TITULO

A Framework for Word Embedding Based Automatic Text Summarization and Evaluation

Tulu Tilahun Hailu

Junqing Yu and Tessfu Geteye Fantaye

Resumen

Text summarization is a process of producing a concise version of text (summary) from one or more information sources. If the generated summary preserves meaning of the original text, it will help the users to make fast and effective decision. However, how much meaning of the source text can be preserved is becoming harder to evaluate. The most commonly used automatic evaluation metrics like Recall-Oriented Understudy for Gisting Evaluation (ROUGE) strictly rely on the overlapping n-gram units between reference and candidate summaries, which are not suitable to measure the quality of abstractive summaries. Another major challenge to evaluate text summarization systems is lack of consistent ideal reference summaries. Studies show that human summarizers can produce variable reference summaries of the same source that can significantly affect automatic evaluation metrics scores of summarization systems. Humans are biased to certain situation while producing summary, even the same person perhaps produces substantially different summaries of the same source at different time. This paper proposes a word embedding based automatic text summarization and evaluation framework, which can successfully determine salient top-n sentences of a source text as a reference summary, and evaluate the quality of systems summaries against it. Extensive experimental results demonstrate that the proposed framework is effective and able to outperform several baseline methods with regard to both text summarization systems and automatic evaluation metrics when tested on a publicly available dataset.

Palabras claves

Automatic evaluation metrics - extrinsic evaluation - intrinsic evaluation - natural language processing - text summarization - word embedding

Acceso

P�GINAS

pp. 0 - 0

N�MERO

Volumen: 11 Parte: 2 (2020)

MATERIAS

INGENIER�A Y CONSTRUCCI�N CIVIL
TECNOLOG�A

REVISTAS SIMILARES

Applied System Innovation
Information
Water

DOI

https://doi.org/10.3390/info11020078

Art�culos similares

Water and Environmental Resources: A Multi-Criteria Assessment of Management Approaches

Acceso

Felipe Armas Vargas, Luzma Fabiola Nava, Eugenio G�mez Reyes, Selene Olea-Olea, Claudia Rojas Serna, Samuel Sandoval Sol�s and Demetrio Meza-Rodr�guez

The present study applied a multi-criteria analysis to evaluate the best approach among six theoretical frameworks related to the integrated management of water?environmental resources, analyzing the frequency of multiple management criteria. The literat... ver m�s

Revista: Water

Web Interface of NER and RE with BERT for Biomedical Text Mining

Acceso

Yeon-Ji Park, Min-a Lee, Geun-Je Yang, Soo Jun Park and Chae-Bong Sohn

The BioBERT Named Entity Recognition (NER) model is a high-performance model designed to identify both known and unknown entities. It surpasses previous NER models utilized by text-mining tools, such as tmTool and ezTag, in effectively discovering novel ... ver m�s

Revista: Applied Sciences

Efficient Conformer for Agglutinative Language ASR Model Using Low-Rank Approximation and Balanced Softmax

Acceso

Ting Guo, Nurmemet Yolwas and Wushour Slamu

Recently, the performance of end-to-end speech recognition has been further improved based on the proposed Conformer framework, which has also been widely used in the field of speech recognition. However, the Conformer model is mostly applied to very wid... ver m�s

Revista: Applied Sciences

Fusion of SoftLexicon and RoBERTa for Purpose-Driven Electronic Medical Record Named Entity Recognition

Acceso

Xiaohui Cui, Yu Yang, Dongmei Li, Xiaolong Qu, Lei Yao, Sisi Luo and Chao Song

Recently, researchers have extensively explored various methods for electronic medical record named entity recognition, including character-based, word-based, and hybrid methods. Nonetheless, these methods frequently disregard the semantic context of ent... ver m�s

Revista: Applied Sciences

Uyghur?Kazakh?Kirghiz Text Keyword Extraction Based on Morpheme Segmentation

Acceso

Sardar Parhat, Mutallip Sattar, Askar Hamdulla and Abdurahman Kadir

In this study, based on a morpheme segmentation framework, we researched a text keyword extraction method for Uyghur, Kazakh and Kirghiz languages, which have similar grammatical and lexical structures. In these languages, affixes and a stem are joined t... ver m�s

Revista: Information

Revistas destacadas

Acceso directo a los n�meros publicados en la revista Infrastructures

Infrastructures

Acceso directo a los n�meros publicados en la revista Informed Infraestructure

Informed Infraestructure

Acceso directo a los n�meros publicados en la revista BiT

Acceso directo a los n�meros publicados en la revista Revista de la Construcci�n

Revista de la Construcci�n

Ver todas las revistas