A High-dimension Index Based on Variable Grid in Big Data Environment
SONG Baoyan
LIU Yu
DING Linlin
Abstract:With the rapid development of the Internet and cloud computing techniques ,the amount of data in the whole sectors of national economy increases sharply ,especially the high‐dimensional big data ,such as the network transactions da‐ta ,the user reviews data and the multimedia data .A proper index structure to support high‐dimension big data can improve the performance of similarity query on high‐dimensional big data .Therefore ,a distributed two‐level index structure is pro‐posed firstly ,in which global index maintains all the information of subspace in the whole data space ,and in which local in‐dex builds M‐tree on each subspace to organize local high‐dimension data .Secondly ,a similarity search algorithm is proposed based on our two‐level index ,including point query and range query .When processing queries from users ,global index can quickly locate and judge which subspaces are relevant to the query and send the query to relevant subspaces .Queries will be processed on local nodes concurrently .This approach can avoid lots of unnecessary retrieves on query‐irrelevant subspaces . Lastly ,massive of experiments also show that the proposed index is much better than existing high‐dimension index struc‐ture ,and has good query performance and scalability .
Keywords:high-dimensional databig datavariable gridM-tree
Publication Date:2015-01-01
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:7( 1717-1722,1728 )
