Inicio  /  Algorithms  /  Vol: 15 Par: 8 (2022)  /  Artículo
ARTÍCULO
TITULO

Short Text Classification with Tolerance-Based Soft Computing Method

Vrushang Patel    
Sheela Ramanna    
Ketan Kotecha and Rahee Walambe    

Resumen

Text classification aims to assign labels to textual units such as documents, sentences and paragraphs. Some applications of text classification include sentiment classification and news categorization. In this paper, we present a soft computing technique-based algorithm (TSC) to classify sentiment polarities of tweets as well as news categories from text. The TSC algorithm is a supervised learning method based on tolerance near sets. Near sets theory is a more recent soft computing methodology inspired by rough sets where instead of set approximation operators used by rough sets to induce tolerance classes, the tolerance classes are directly induced from the feature vectors using a tolerance level parameter and a distance function. The proposed TSC algorithm takes advantage of the recent advances in efficient feature extraction and vector generation from pre-trained bidirectional transformer encoders for creating tolerance classes. Experiments were performed on ten well-researched datasets which include both short and long text. Both pre-trained SBERT and TF-IDF vectors were used in the experimental analysis. Results from transformer-based vectors demonstrate that TSC outperforms five well-known machine learning algorithms on four datasets, and it is comparable with all other datasets based on the weighted F1, Precision and Recall scores. The highest AUC-ROC (Area under the Receiver Operating Characteristics) score was obtained in two datasets and comparable in six other datasets. The highest ROC-PRC (Area under the Precision?Recall Curve) score was obtained in one dataset and comparable in four other datasets. Additionally, significant differences were observed in most comparisons when examining the statistical difference between the weighted F1-score of TSC and other classifiers using a Wilcoxon signed-ranks test.

 Artículos similares

       
 
Yehia Ibrahim Alzoubi, Ahmet E. Topcu and Ahmed Enis Erkaya    
The growth in textual data associated with the increased usage of online services and the simplicity of having access to these data has resulted in a rise in the number of text classification research papers. Text classification has a significant influen... ver más
Revista: Applied Sciences

 
Changwon Kwak, Pilsu Jung and Seonah Lee    
Issue reports are valuable resources for the continuous maintenance and improvement of software. Managing issue reports requires a significant effort from developers. To address this problem, many researchers have proposed automated techniques for classi... ver más
Revista: Applied Sciences

 
Naofumi Fujishiro, Yasuhiro Otaki and Shoji Kawachi    
In this study, we developed a similar text retrieval system using Sentence-BERT (SBERT) for our database of closed medical malpractice claims and investigated its retrieval accuracy. We assigned each case in the database a short Japanese summary of the a... ver más
Revista: Applied Sciences

 
Yaser Altameemi and Mohammed Altamimi    
Using advanced algorithms to conduct a thematic analysis reduces the time taken and increases the efficiency of the analysis. Long short-term memory (LSTM) is effective in the field of text classification and natural language processing (NLP). In this st... ver más
Revista: Applied Sciences

 
Jie Long, Zihan Li, Qi Xuan, Chenbo Fu, Songtao Peng and Yong Min    
The opinion recognition for comments in Internet media is a new task in text analysis. It takes comment statements as the research object, by learning the opinion tendency in the original text with annotation, and then performing opinion tendency recogni... ver más
Revista: Applied Sciences