Journal article

Search Effectiveness in Nonredundant Sequence Databases: Assessments and Solutions.

Qingyu Chen, Xiuzhen Zhang, Yu Wan, Justin Zobel, Karin Verspoor

Journal of Computational Biology | Mary Ann Liebert | Published : 2018

Abstract

Duplicate sequence records-that is, records having similar or identical sequences-are a challenge in search of biological sequence databases. They significantly increase database search time and can lead to uninformative search results containing similar sequences. Sequence clustering methods have been used to address this issue to group similar sequences into clusters. These clusters form a nonredundant database consisting of representatives (one record per cluster) and members (the remaining records in a cluster). In this approach, for nonredundant database search, users search against representatives first and optionally expand search results by exploring member records from matching clus..

View full abstract