HSIT:Distributed Similarity Query Index for Massive Data
YAO Hui
LIU Wen
Abstract:Similarity queries are commonly used in fields such as information retrieval,biology,and network security to ana-lyze associations between data.The traditional method to perform similarity query often requires the query point and each piece of da-ta in the database to be calculated.As the amount of data increases,the amount of computation increases linearly.In order to im-prove the efficiency of distributed similarity query of massive data,a HBase similarity index tree(HSIT)is proposed.In the process of data storage,the algorithm realizes the dynamic establishment of similarity index tree structure.The HSIT index can divide the similarity data according to the similarity threshold and store it in the adjacent area of HBase.When a user performs a similarity que-ry,the query node can quickly retrieve similar regions through HSIT.The index can achieve efficient pruning,so that only similar regions need to be calculated in pairs.The similarity query is performed through the exponential growth of 20 000 data to 1.28 mil-lion data.Compared with the DSCS-LTS algorithm,the experimental results show that the efficiency of the HSIT algorithm has been improved.
Keywords:massive datadistributedsimilarity queryindexHBase
Publication Date:2025-03-20
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:7( 718-724 )
