ARTÍCULO
TITULO

Classification of Scientific Documents in the Kazakh Language Using Deep Neural Networks and a Fusion of Images and Text

Andrey Bogdanchikov    
Dauren Ayazbayev and Iraklis Varlamis    

Resumen

The rapid development of natural language processing and deep learning techniques has boosted the performance of related algorithms in several linguistic and text mining tasks. Consequently, applications such as opinion mining, fake news detection or document classification that assign documents to predefined categories have significantly benefited from pre-trained language models, word or sentence embeddings, linguistic corpora, knowledge graphs and other resources that are in abundance for the more popular languages (e.g., English, Chinese, etc.). Less represented languages, such as the Kazakh language, balkan languages, etc., still lack the necessary linguistic resources and thus the performance of the respective methods is still low. In this work, we develop a model that classifies scientific papers written in the Kazakh language using both text and image information and demonstrate that this fusion of information can be beneficial for cases of languages that have limited resources for machine learning models? training. With this fusion, we improve the classification accuracy by 4.4499% compared to the models that use only text or only image information. The successful use of the proposed method in scientific documents? classification paves the way for more complex classification models and more application in other domains such as news classification, sentiment analysis, etc., in the Kazakh language.

 Artículos similares

       
 
Nehad M. Ibrahim, Dalia G. Gabr, Atta Rahman, Dhiaa Musleh, Dania AlKhulaifi and Mariam AlKharraa    
Plant taxonomy is the scientific study of the classification and naming of various plant species. It is a branch of biology that aims to categorize and organize the diverse variety of plant life on earth. Traditionally, plant taxonomy has been performed ... ver más

 
Yao Ding, Yin Wang, Shuming Yang, Xiaolong Zhao, Lili Ouyang and Chengyue Lai    
The current standards used for nitrogen pollution evaluation are lacking, and scientific classification methods are needed for nitrogen pollution to improve water quality management capabilities. This study addresses the important issue of assessing surf... ver más
Revista: Water

 
Pedro João Rodrigues, Walter Gomes and Maria Alice Pinto    
Honey bee classification by wing geometric morphometrics entails the first step of manual annotation of 19 landmarks in the forewing vein junctions. This is a time-consuming and error-prone endeavor, with implications for classification accuracy. Herein,... ver más

 
Maha Gharaibeh, Mothanna Almahmoud, Mostafa Z. Ali, Amer Al-Badarneh, Mwaffaq El-Heis, Laith Abualigah, Maryam Altalhi, Ahmad Alaiad and Amir H. Gandomi    
Neuroimaging refers to the techniques that provide efficient information about the neural structure of the human brain, which is utilized for diagnosis, treatment, and scientific research. The problem of classifying neuroimages is one of the most importa... ver más

 
Chaoxiang Chen, Shiping Ye, Zhican Bai, Juan Wang, Alexander Nedzved and Sergey Ablameyko    
With the acceleration of urbanization, climate problems affecting human health and safe operation of cities have intensified, such as heat island effect, haze, and acid rain. Using high-resolution remote sensing mapping image data to design scientific an... ver más