首页> 外文会议> >Fast approximate similarity search in extremely high-dimensional data sets

【24h】

Fast approximate similarity search in extremely high-dimensional data sets

机译：在超高维数据集中进行快速近似相似性搜索

获取原文

页面导航

摘要
著录项
引文网络
相似文献
相关主题

摘要

This paper introduces a practical index for approximate similarity queries of large multi-dimensional data sets: the spatial approximation sample hierarchy (SASH). A SASH is a multi-level structure of random samples, recursively constructed by building a SASH on a large randomly selected sample of data objects, and then connecting each remaining object to several of their approximate nearest neighbors from within the sample. Queries are processed by first locating approximate neighbors within the sample, and then using the pre-established connections to discover neighbors within the remainder of the data set. The SASH index relies on a pairwise distance measure, but otherwise makes no assumptions regarding the representation of the data. Experimental results are provided for query-by-example operations on protein sequence, image, and text data sets, including one consisting of more than 1 million vectors spanning more than 1.1 million terms - far in excess of what spatial search indices can handle efficiently. For sets of this size, the SASH can return a large proportion of the true neighbors roughly 2 orders of magnitude faster than sequential search.

机译：本文为大型多维数据集的近似相似性查询引入了一种实用的索引：空间近似样本层次结构（SASH）。 SASH是随机样本的多级结构，通过在大量随机选择的数据对象样本上构建SASH，然后将每个剩余的对象连接到样本中与其近似的最近邻居中的几个，来递归构造。通过首先在样本中定位近似邻居，然后使用预先建立的连接来发现数据集其余部分中的邻居，来处理查询。 SASH索引依赖于成对距离度量，但否则不对数据表示做任何假设。提供了针对蛋白质序列，图像和文本数据集的按示例查询操作的实验结果，其中包括一个由超过100万个向量组成的，跨越110万个以上术语的向量，远远超出了空间搜索索引可以有效处理的范围。对于这种大小的集合，与顺序搜索相比，SASH可以返回很大一部分真实邻居，大约快2个数量级。

著录项

来源
《》|2005年|P.619-630|共12页
会议地点
作者
Houle; M.E.; Jun Sakuma;
展开▼
作者单位

展开▼
会议组织
原文格式 PDF
正文语种
中图分类工业技术;
关键词
data structures; distributed databases; query formulation; query processing; very large databases; approximate nearest neighbors; data objects; fast approximate similarity search; large multidimensional data sets; query processing; spatial approximation sample hie;

机译：数据结构;分布式数据库;查询表述;查询处理;非常大的数据库;近似最近的邻居;数据对象;快速近似相似性搜索;大型多维数据集;查询处理;空间近似样本层;

相似文献

外文文献
中文文献
专利

1. A fast and scalable similarity search in high-dimensional image datasets [J] . Youssef Hanyf, Hassan Silkan International Journal of Computer Applications in Technology . 2019,第1期

机译：在高维图像数据集中快速且可扩展的相似性搜索
2. CSVD: clustering and singular value decomposition for approximate similarity search in high-dimensional spaces [J] . Castelli V., Thomasian A., Chung-Sheng Li IEEE Transactions on Knowledge and Data Engineering . 2003,第3期

机译：CSVD：聚类和奇异值分解，用于在高维空间中进行近似相似性搜索
3. Pivot-based approximate k-NN similarity joins for big high-dimensional data [J] . Cech Premysl, Lokoc Jakub, Silva Yasin N. Information Systems . 2020,第Jana期

机译：基于枢轴的近似k-NN相似度连接可处理大型高维数据
4. Fast Approximate Similarity Search in Extremely High-Dimensional Data Sets [C] . Michael E. Houle, Jun Sakuma International Conference on Data Engineering . 2005

机译：在极高维度数据集中快速近似相似性搜索
5. High-Dimensional Similarity Search for Large Datasets. [D] . Dong, Wei. 2011

机译：大数据集的高维相似性搜索。
6. i-ADHoRe 3.0—fast and sensitive detection of genomic homology in extremely large data sets [O] . Sebastian Proost, Jan Fostier, Dieter De Witte, 2012

机译：i-ADHoRe 3.0-在超大型数据集中快速灵敏地检测基因组同源性
7. Fast Similarity Search for High-Dimensional Dataset [O] . 2008

机译：快速相似搜索高维数据集

Fast approximate similarity search in extremely high-dimensional data sets

摘要

著录项

引文网络

相似文献

相关主题

期刊订阅