A Space and Time Efficient Algorithm for Constructing Compressed Suffix Arrays

Wing-Kai Hon; Tak-Wah Lam; Kunihiko Sadakane; Wing-Kin Sung; Siu-Ming Yiu

首页> 外文期刊>Algorithmica >A Space and Time Efficient Algorithm for Constructing Compressed Suffix Arrays

【24h】

A Space and Time Efficient Algorithm for Constructing Compressed Suffix Arrays

机译：一种高效的时空压缩后缀数组构造算法

获取原文

获取原文并翻译 | 示例

开具论文收录证明 >>

页面导航

摘要
著录项
引文网络
相似文献
相关主题

摘要

With the first human DNA being decoded into a sequence of about 2.8 billion characters, much biological research has been centered on analyzing this sequence. Theoretically speaking, it is now feasible to accommodate an index for human DNA in the main memory so that any pattern can be located efficiently. This is due to the recent breakthrough on compressed suffix arrays, which reduces the space requirement from O(n log n) bits to O(n) bits. However, constructing compressed suffix arrays is still not an easy task because we still have to compute suffix arrays first and need a working memory of O(n log n) bits (i.e., more than 13 gigabytes for human DNA). This paper initiates the study of constructing compressed suffix arrays directly from the text. The main contribution is a construction algorithm that uses only O(n) bits of working memory, and the time complexity is O(n log n). Our construction algorithm is also time and space efficient for texts with large alphabets such as Chinese or Japanese. Precisely, when the alphabet size is |Σ|, the working space is O(n log|Σ|) bits, and the time complexity remains O(n log n), which is independent of |Σ|.

机译：随着第一个人类DNA被解码为约28亿个字符的序列，许多生物学研究都集中在分析该序列上。从理论上讲，现在可以在主内存中容纳人类DNA的索引，以便可以有效地定位任何模式。这是由于最近对压缩后缀数组的突破，将空间需求从O（n log n）位减少到O（n）位。但是，构造压缩后缀数组仍然不是一件容易的事，因为我们仍然必须首先计算后缀数组，并且需要O（n log n）位（即，人类DNA超过13 GB）的工作存储器。本文直接从文本开始研究构建压缩后缀数组。主要贡献是一种仅使用工作内存的O（n）位的构造算法，时间复杂度为O（n log n）。对于具有大字母的文本（例如中文或日语），我们的构造算法在时间和空间上也很有效。精确地，当字母大小为|Σ|时，工作空间为O（n log |Σ|）位，时间复杂度保持为O（n log n），与|Σ|无关。

著录项

来源
《Algorithmica》 |2007年第1期|p.23-36|共14页
作者
Wing-Kai Hon; Tak-Wah Lam; Kunihiko Sadakane; Wing-Kin Sung; Siu-Ming Yiu;
展开▼
作者单位

Department of Computer Science, The University of Hong Kong, Pokfulam, Hong Kong;

展开▼
收录信息美国《科学引文索引》(SCI);美国《工程索引》(EI);
原文格式 PDF
正文语种 eng
中图分类计算技术、计算机技术;
关键词
text indexing; pattern matching; compression; construction;

机译：文本索引;模式匹配;压缩;构造;

相似文献

外文文献
中文文献
专利

1. Two Efficient Algorithms for Linear Time Suffix Array Construction [J] . Nong Ge, Zhang Sen, Chan Wai Hong Computers, IEEE Transactions on . 2011,第10期

机译：线性时间后缀数组构造的两种有效算法
2. Alphabet-independent linear-time construction of compressed suffix arrays using o(nlogn)-bit working space [J] . Joong Chae Na, Kunsoo Park Theoretical computer science . 2007,第1a3期

机译：使用o（nlogn）位工作空间的压缩后缀数组的与字母无关的线性时间构造
3. Time-space trade-offs for compressed suffix arrays [J] . S. Srinivasa Rao Information Processing Letters . 2002,第6期

机译：压缩后缀数组的时空权衡
4. A Space and Time Efficient Algorithm for Constructing Compressed Suffix Arrays [C] . Tak-Wah Lam, Kunihiko Sadakane, Wing-Kin Sung, Computing and Combinatorics . 2002

机译：一种高效的时空压缩后缀数组构造算法
5. Efficient Algorithms for a Class of Space-Time Stochastic Models [D] . Reyes, Brandon C. 2018

机译：一类时空随机模型的高效算法
6. SA-SSR: a suffix array-based algorithm for exhaustive and efficient SSR discovery in large genetic sequences [O] . B. D. Pickett, S. M. Karlinsey, C. E. Penrod, -1

机译：SA-SSR：一种基于后缀数组的算法可在大型遗传序列中进行详尽而高效的SSR发现
7. A Space and Time Efficient Algorithm for Constructing Compressed Suffix Arrays [O] . Wing-Kai Hon 2012

机译：一种高效的时空压缩后缀数组构造算法

A Space and Time Efficient Algorithm for Constructing Compressed Suffix Arrays

摘要

著录项

引文网络

相似文献

相关主题

期刊订阅