Redirigiendo al acceso original de articulo en 19 segundos...
Inicio  /  Information  /  Vol: 12 Par: 8 (2021)  /  Artículo
ARTÍCULO
TITULO

A Study of Analogical Density in Various Corpora at Various Granularity

Rashel Fam and Yves Lepage    

Resumen

In this paper, we inspect the theoretical problem of counting the number of analogies between sentences contained in a text. Based on this, we measure the analogical density of the text. We focus on analogy at the sentence level, based on the level of form rather than on the level of semantics. Experiments are carried on two different corpora in six European languages known to have various levels of morphological richness. Corpora are tokenised using several tokenisation schemes: character, sub-word and word. For the sub-word tokenisation scheme, we employ two popular sub-word models: unigram language model and byte-pair-encoding. The results show that the corpus with a higher Type-Token Ratio tends to have higher analogical density. We also observe that masking the tokens based on their frequency helps to increase the analogical density. As for the tokenisation scheme, the results show that analogical density decreases from the character to word. However, this is not true when tokens are masked based on their frequencies. We find that tokenising the sentences using sub-word models and masking the least frequent tokens increase analogical density.

 Artículos similares

       
 
Nejc Co?, Reza Ahmadian and Roger A. Falconer    
Understanding the impact of various hydraulic structures, such as coastal reservoirs and tidal range impoundments, has been one of the key challenges of hydro?environmental engineering in recent years. Over the last half-century, several proposals for ti... ver más
Revista: Water

 
Sri Efrinita Irwan, Triyana Muliawati     Pág. 111 - 118
One of the important areas in mathematics is graph theory. A graph is a mathematical structure used to model pairwise relations between objects. The theory of graph can be applied in various problems. The purpose of this paper is to solve the dormitory r... ver más

 
Cristiane Canan, Fernanda Delaroza, Rúbia Casagrande, Marcela Maria Baracat, Massami Shimokomaki, Elza Iouko Ida (Author)     Pág. 457 - 463
Rice bran is a by-product of rice processing industry, with high levels of phytic acid or phytate. Considering phytic acid antioxidant activity, its various applications and its high concentration in rice bran, this study had the objective of evaluating ... ver más

 
Yunfei Yang, Zhicheng Zhang, Jiapeng Zhao, Bin Zhang, Lei Zhang, Qi Hu and Jianglong Sun    
Resistance serves as a critical performance metric for ships. Swift and accurate resistance prediction can enhance ship design efficiency. Currently, methods for determining ship resistance encompass model tests, estimation techniques, and computational ... ver más

 
Jiwun Yoon, Sang-Yong Lee and Ji-Yong Lee    
Humans share a similar body structure, but each individual possesses unique characteristics, which we define as one?s body type. Various classification methods have been devised to understand and assess these body types. Recent research has applied artif... ver más
Revista: Applied Sciences