An Improved Retrievability-Based Cluster-Resampling Approach for Pseudo Relevance Feedback

Shariq Bashir

Resumen

Cluster-based pseudo-relevance feedback (PRF) is an effective approach for searching relevant documents for relevance feedback. Standard approach constructs clusters for PRF only on the basis of high similarity between retrieved documents. The standard approach works quite well if the retrieval bias of the retrieval model does not create any effect on the retrievability of documents. In our experiments we observed when a collection contains retrieval bias, then high retrievable documents of clusters are frequently retrieved at top positions for most of the queries, and these drift the relevance feedback away from relevant documents. For reducing (retrieval bias) noise, we enhance the standard cluster construction approach by constructing clusters on the basis of high similarity and retrievability. We call this retrievability and cluster-based PRF. This enhanced approach keeps only those documents in the clusters that are not frequently retrieve due to retrieval bias. Although this approach improves the effectiveness, however, it penalizes high retrievable documents even if these documents are most relevant to the clusters. To handle this problem, in a second approach, we extend the basic retrievability concept by mining frequent neighbors of the clusters. The frequent neighbors approach keeps only those documents in the clusters that are frequently retrieved with other neighbors of clusters and infrequently retrieved with those documents that are not part of the clusters. Experimental results show that two proposed extensions are helpful for identifying relevant documents for relevance feedback and increasing the effectiveness of queries.

Palabras claves

document clustering - machine learning - information retrieval - pseudo-relevance feedback - query expansion - retrieval bias - retrievability measure

Acceso

P�GINAS

pp. 0 - 0

N�MERO

Volumen: 5 Parte: 4 (2016)

MATERIAS

INGENIER�A Y CONSTRUCCI�N CIVIL
TECNOLOG�A

REVISTAS SIMILARES

Computers
South African Journal of Science and Technology
Applied Sciences

DOI

https://doi.org/10.3390/computers5040029

Art�culos similares

Research Trends in the Use of Machine Learning Applied in Mobile Networks: A Bibliometric Approach and Research Agenda

Acceso

Vanessa Garc�a-Pineda, Alejandro Valencia-Arias, Juan Camilo Pati�o-Vanegas, Juan Jos� Flores Cueto, Diana Arango-Botero, Angel Marcelo Rojas Coronel and Paula Andrea Rodr�guez-Correa

This article aims to examine the research trends in the development of mobile networks from machine learning. The methodological approach starts from an analysis of 260 academic documents selected from the Scopus and Web of Science databases and is based... ver m�s

Revista: Informatics

A Survey of OCR in Arabic Language: Applications, Techniques, and Challenges

Acceso

Safiullah Faizullah, Muhammad Sohaib Ayub, Sajid Hussain and Muhammad Asad Khan

Optical character recognition (OCR) is the process of extracting handwritten or printed text from a scanned or printed image and converting it to a machine-readable form for further data processing, such as searching or editing. Automatic text extraction... ver m�s

Revista: Applied Sciences

Using Deep-Learned Vector Representations for Page Stream Segmentation by Agglomerative Clustering

Acceso

Lukas Busch, Ruben van Heusden and Maarten Marx

Page stream segmentation (PSS) is the task of retrieving the boundaries that separate source documents given a consecutive stream of documents (for example, sequentially scanned PDF files). The task has recently gained more interest as a result of the di... ver m�s

Revista: Algorithms

Knowledge-Based Intelligent Text Simplification for Biological Relation Extraction

Acceso

Jaskaran Gill, Madhu Chetty, Suryani Lim and Jennifer Hallinan

Relation extraction from biological publications plays a pivotal role in accelerating scientific discovery and advancing medical research. While vast amounts of this knowledge is stored within the published literature, extracting it manually from this co... ver m�s

Revista: Informatics

Relevance of Machine Learning Techniques in Water Infrastructure Integrity and Quality: A Review Powered by Natural Language Processing

Acceso

Jos� Garc�a, Andres Leiva-Araos, Emerson Diaz-Saavedra, Paola Moraga, Hernan Pinto and V�ctor Yepes

Water infrastructure integrity, quality, and distribution are fundamental for public health, environmental sustainability, economic development, and climate change resilience. Ensuring the robustness and quality of water infrastructure is pivotal for sec... ver m�s

Revista: Applied Sciences

Revistas destacadas

Acceso directo a los n�meros publicados en la revista Infrastructures

Infrastructures

Acceso directo a los n�meros publicados en la revista Informed Infraestructure

Informed Infraestructure

Acceso directo a los n�meros publicados en la revista BiT

Acceso directo a los n�meros publicados en la revista Revista de la Construcci�n

Revista de la Construcci�n

Ver todas las revistas disponibles