Longest Common Prefix Arrays for Succinct k-Spectra

Alanko, Jarno N.; Biagi, Elena; Puglisi, Simon J.

doi:10.1007/978-3-031-43980-3_1

Part of the book series: Lecture Notes in Computer Science ((LNCS,volume 14240))

Included in the following conference series:

International Symposium on String Processing and Information Retrieval

425 Accesses

Abstract

The k-spectrum of a string is the set of all distinct substrings of length k occurring in the string. K-spectra have many applications in bioinformatics including pseudoalignment and genome assembly. The Spectral Burrows-Wheeler Transform (SBWT) has been recently introduced as an algorithmic tool to efficiently represent and query these objects. The longest common prefix ($\textit{LCP}$) array for a k-spectrum is an array of length n that stores the length of the longest common prefix of adjacent k-mers as they occur in lexicographical order. The $\textit{LCP}$ array has at least two important applications, namely to accelerate pseudoalignment algorithms using the SBWT and to allow simulation of variable-order de Bruijn graphs within the SBWT framework. In this paper we explore algorithms to compute the $\textit{LCP}$ array efficiently from the SBWT representation of the k-spectrum. Starting with a straightforward O(nk) time algorithm, we describe algorithms that are efficient in both theory and practice. We show that the $\textit{LCP}$ array can be computed in optimal O(n) time, where n is the length of the SBWT of the spectrum. In practical genomics scenarios, we show that this theoretically optimal algorithm is indeed practical, but is often outperformed on smaller values of k by an asymptotically suboptimal algorithm that interacts better with the CPU cache. Our algorithms share some features with both classical Burrows-Wheeler inversion algorithms and LCP array construction algorithms for suffix arrays. Our C++ implementations of these algorithms are available at https://github.com/jnalanko/kmer-lcs.

Supported in part by the Academy of Finland via grants 339070 and 351150.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Subscribe and save

Springer+ Basic

EUR 32.99 /Month

Get 10 units per month
Download Article/Chapter or Ebook
1 Unit = 1 Article or 1 Chapter
Cancel anytime

Subscribe now

Buy Now

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 59.99; Price excludes VAT (USA)

Softcover Book: USD 74.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

gsufsort: constructing suffix arrays, LCP arrays and BWTs for string collections

Article Open access 22 September 2020

The Colored Longest Common Prefix Array Computed via Sequential Scans

A Framework for Space-Efficient String Kernels

Article 07 February 2017

Notes

1.
We remark here that the LCS array of a colexicographically-ordered spectrum is equivalent to the longest common prefix (LCP) array of the lexicographically-ordered spectrum, and the algorithms we describe in this paper to compute the LCS array are trivially adapted to compute the LCP array.
2.
Wheeler graphs are a class of graphs including de Bruijn graphs, that admit a generalization of the Burrows-Wheeler transform. The SBWT can be seen as a special case of the Wheeler graph indexing framework.
3.
A similar but different structure is described in [2].
4.
Assuming the input to the BWT is terminated with a $-symbol, and there is an added $-edge from the last k-mer of the input to the root of the SBWT graph.

References

Alanko, J.N., Biagi, E., Puglisi, S.J., Vuohtoniemi, J.: Subset wavelet trees. In: Proceedings of the 21st International Symposium on Experimental Algorithms (SEA), LIPIcs. Schloss Dagstuhl - Leibniz-Zentrum für Informatik (2023)
Google Scholar
Alanko, J.N., Puglisi, S.J., Vuohtoniemi, J.: Small searchable k-spectra via subset rank queries on the spectral burrows-wheeler transform. In Proceedings of SIAM Conference on Applied and Computational Discrete Algorithms (ACDA), pp. 225–236. Society for Industrial and Applied Mathematics (2023)
Google Scholar
Alanko, J.N., Vuohtoniemi, J., Mäklin, T., Puglisi, S.J.: Themisto: a scalable colored k-mer index for sensitive pseudoalignment against hundreds of thousands of bacterial genomes. Bioinformatics (2023)
Google Scholar
Beller, T., Gog, S., Ohlebusch, E., Schnattinger, T.: Computing the longest common prefix array based on the Burrows-Wheeler transform. J. Discrete Algorithms 18, 22–31 (2013)
Article MathSciNet MATH Google Scholar
Boucher, C., Bowe, A., Gagie, T., Puglisi, S.J., Sadakane, K.: Variable-order de Bruijn graphs. In: Proceedings of the 25th Data Compression Conference (DCC), pp. 383–392. IEEE (2015)
Google Scholar
Compeau, P.E., Pevzner, P.A., Tesler, G.: Why are de Bruijn graphs useful for genome assembly? Nat. Biotechnol. 29(11), 987 (2011)
Article Google Scholar
Conte, A., Cotumaccio, N., Gagie, T., Manzini, G., Prezza, N., Sciortino, M.: Computing matching statistics on Wheeler DFAs. ar**v preprint ar**v:2301.05338 (2023)
Holley, G., Melsted, P.: Bifrost: highly parallel construction and indexing of colored and compacted de Bruijn graphs. Genome Biol. 21(1), 1–20 (2020)
Article Google Scholar
Jeffery, I.B., et al.: Differences in fecal microbiomes and metabolomes of people with vs without irritable bowel syndrome and bile acid malabsorption. Gastroenterology 158(4), 1016–1028 (2020)
Article Google Scholar
Maillet, N., Lemaitre, C., Chikhi, R., Lavenier, D., Peterlongo, P.: Compareads: comparing huge metagenomic experiments. BMC Bioinf. 13(19), 1–10 (2012)
Google Scholar
Marchet, C., Boucher, C., Puglisi, S.J., Medvedev, P., Salson, M., Chikhi, R.: Data structures based on k-mers for querying large collections of sequencing data sets. Genome Res. 31(1), 1–12 (2021)
Article Google Scholar
Ondov, B.D., et al.: Mash: fast genome and metagenome distance estimation using MinHash. Genome Biol. 17(1), 1–14 (2016)
Article Google Scholar
Salikhov, K.: Efficient algorithms and data structures for indexing DNA sequence data. PhD thesis, Université Paris-Est; Université Lomonossov (Moscou) (2017)
Google Scholar

Download references

Author information

Authors and Affiliations

Helsinki Institute for Information Technology (HIIT), Department of Computer Science, University of Helsinki, Helsinki, Finland
Jarno N. Alanko, Elena Biagi & Simon J. Puglisi

Authors

Jarno N. Alanko
View author publications
You can also search for this author in PubMed Google Scholar
Elena Biagi
View author publications
You can also search for this author in PubMed Google Scholar
Simon J. Puglisi
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Elena Biagi .

Editor information

Editors and Affiliations

ISTI-CNR, Pisa, Italy
Franco Maria Nardini
University of Pisa, Pisa, Italy
Nadia Pisanti
University of Pisa, Pisa, Italy
Rossano Venturini

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Alanko, J.N., Biagi, E., Puglisi, S.J. (2023). Longest Common Prefix Arrays for Succinct k-Spectra. In: Nardini, F.M., Pisanti, N., Venturini, R. (eds) String Processing and Information Retrieval. SPIRE 2023. Lecture Notes in Computer Science, vol 14240. Springer, Cham. https://doi.org/10.1007/978-3-031-43980-3_1

Download citation

DOI: https://doi.org/10.1007/978-3-031-43980-3_1
Published: 20 September 2023
Publisher Name: Springer, Cham
Print ISBN: 978-3-031-43979-7
Online ISBN: 978-3-031-43980-3
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics

Longest Common Prefix Arrays for Succinct k-Spectra

Abstract

Access this chapter

Subscribe and save

Buy Now

Similar content being viewed by others

gsufsort: constructing suffix arrays, LCP arrays and BWTs for string collections

The Colored Longest Common Prefix Array Computed via Sequential Scans

A Framework for Space-Efficient String Kernels

Notes

References

Author information

Authors and Affiliations

Corresponding author

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Publish with us

Subscribe and save

Buy Now

Navigation

Longest Common Prefix Arrays for Succinct k-Spectra

Abstract

Access this chapter

Subscribe and save

Buy Now

Similar content being viewed by others

gsufsort: constructing suffix arrays, LCP arrays and BWTs for string collections

The Colored Longest Common Prefix Array Computed via Sequential Scans

A Framework for Space-Efficient String Kernels

Notes

References

Author information

Authors and Affiliations

Corresponding author

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Share this paper

Publish with us

Search

Navigation