Thomas Konstantinovsky

Thomas Konstantinovsky

PhD candidate, Yaari Lab, Faculty of Engineering, Bar-Ilan University

Algorithms, probabilistic models, and machine learning for biological sequences

I am a computer scientist researching the adaptive immune system. My expertise is in probability theory and mathematical modeling, machine learning and deep learning, data science, and the tools of compression and information theory, which I apply to problems in adaptive immune receptor repertoire (AIRR) analysis and in biology more broadly. I design algorithms and data representations for problems at scale, and I hold them to two standards: as efficient as the problem allows, and approachable by anyone, whatever their field. Most of my work is released as open-source software.

Research contributions

GenAIRR

Generates fully annotated immune-receptor sequences for training and evaluating analysis methods, across 23 species.

AlignAIR

Uses deep learning to identify the gene segments and boundaries in antibody sequences, with calibrated confidence for every call.

LZGraphs

Uses compression-based graph representations to model immune-receptor repertoires without aligning sequences.

Selected publications

2025Enhancing sequence alignment of adaptive immune receptors through multi-task deep learningFirst authorDeep-learning aligner for antibody sequences that returns a calibrated confidence for every gene call; released with a web app and a container.Nucleic Acids Research

Thomas Konstantinovsky†, Ayelet Peres†, Ran Eisenberg, Pazit Polak, Ofir Lindenbaum, Gur Yaari

Sequence alignment of immunoglobulin (Ig) sequences is central to the computational analysis of adaptive immune receptor repertoire sequencing (AIRR-seq) data, impacting adaptive immunity research and antibody engineering. Traditional Ig sequence aligners often struggle with the complexities of V(D)J recombination and somatic hypermutation (SHM), resulting in suboptimal allele assignment accuracy and sequence segmentation. We introduce AlignAIR, a deep learning-based aligner that leverages simulation and a multi-task learning framework. AlignAIR improves allele assignment accuracy, productivity assessment, sequence segmentation and speed, and its latent space captures SHM characteristics. It integrates with existing AIRR-seq pipelines and ships with a web interface and a container image for local processing of millions of sequences.

Cite
@article{konstantinovsky2025enhancing,
  title = {{Enhancing sequence alignment of adaptive immune receptors through multi-task deep learning}},
  author = {Thomas Konstantinovsky and Ayelet Peres and Ran Eisenberg and Pazit Polak and Ofir Lindenbaum and Gur Yaari},
  journal = {Nucleic Acids Research},
  volume = {53},
  number = {13},
  pages = {gkaf651},
  year = {2025},
  doi = {10.1093/nar/gkaf651},
  url = {https://doi.org/10.1093/nar/gkaf651}
}

Click the entry to select it, then copy.

2024An unbiased comparison of immunoglobulin sequence alignersFirst authorIntroduced GenAIRR, a ground-truth simulator, and used it for an unbiased comparison of widely used immunoglobulin aligners.Briefings in Bioinformatics

Thomas Konstantinovsky†, Ayelet Peres†, Pazit Polak, Gur Yaari

Reliable analysis of AIRR-seq data depends on accurate alignment of rearranged immunoglobulin (Ig) sequences, yet there is no unified benchmarking standard that represents the complexities of real data. We introduce GenAIRR, a modular simulation framework that generates Ig sequences alongside their ground truth, realistically simulating V(D)J recombination, somatic hypermutation and a range of sequence corruptions. Using GenAIRR we assess prominent Ig sequence aligners across a range of metrics and reveal distinct performance characteristics for each. The datasets and evaluation criteria provide a basis for unbiased benchmarking of immunogenetics tools.

Cite
@article{konstantinovsky2024unbiased,
  title = {{An unbiased comparison of immunoglobulin sequence aligners}},
  author = {Thomas Konstantinovsky and Ayelet Peres and Pazit Polak and Gur Yaari},
  journal = {Briefings in Bioinformatics},
  volume = {25},
  number = {6},
  pages = {bbae556},
  year = {2024},
  doi = {10.1093/bib/bbae556},
  url = {https://doi.org/10.1093/bib/bbae556}
}

Click the entry to select it, then copy.

2023A novel approach to T-cell receptor beta chain (TCRB) repertoire encoding using lossless string compressionFirst authorLempel-Ziv graph encoding of T-cell repertoires that yields generation probabilities and diversity measures without alignment.Bioinformatics

Thomas Konstantinovsky, Gur Yaari

T-cell receptor beta chain (TCRB) repertoires are crucial for understanding immune responses, but their diversity makes them hard to represent and analyze. We introduce an encoding of TCRB repertoires based on the Lempel-Ziv 76 algorithm that yields a graph-like model of a repertoire. The representation supports generation probability inference, feature-vector derivation, sequence generation, a new diversity estimate and a new sequence centrality measure, and requires no alignment or reference genotype. Applied to four large public TCRB datasets, it demonstrates a scalable approach for large sequencing data. Implemented in the LZGraphs Python package.

Cite
@article{konstantinovsky2023novel,
  title = {{A novel approach to T-cell receptor beta chain (TCRB) repertoire encoding using lossless string compression}},
  author = {Thomas Konstantinovsky and Gur Yaari},
  journal = {Bioinformatics},
  volume = {39},
  number = {7},
  pages = {btad426},
  year = {2023},
  doi = {10.1093/bioinformatics/btad426},
  url = {https://doi.org/10.1093/bioinformatics/btad426}
}

Click the entry to select it, then copy.

2026Flashback: a reversible bilateral run-peeling decomposition of stringsFirst authorA linear-time, reversible string decomposition with an exact run-pairing theorem; the basis of FlashbackGraph.arXivPreprint

Thomas Konstantinovsky, Gur Yaari

We introduce Flashback, a reversible string decomposition that repeatedly peels the maximal leading and trailing character runs from a sentinel-wrapped input, recording each pair as one bilateral token. Decomposition and reconstruction both run in linear time and space. The central result is a run-pairing theorem: Flashback pairs the first run of the string with the last, the second with the second-to-last, and so on, giving an exact token count that matches a lower bound for any admissible bilateral run-peeling scheme. Structural properties follow as corollaries, including a two-symbol irreducible kernel, a characterization of palindromes, a finite-state description of the image, and edit locality.

Cite
@misc{konstantinovsky2026flashback,
  title = {{Flashback: a reversible bilateral run-peeling decomposition of strings}},
  author = {Thomas Konstantinovsky and Gur Yaari},
  year = {2026},
  eprint = {2604.26190},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.26190}
}

Click the entry to select it, then copy.

All publications

Patents

2026Scalable string set graph analysis (Phoenix)US provisional patent applicationPending

Inventors: Thomas Konstantinovsky, Gur Yaari

Filed Feb 2026 · Assignee: Yale University · USPTO

2024A novel machine learning model and mathematical framework for online candidate assessmentUS 11,989,552 B2Granted

Inventors: Thomas Konstantinovsky

Granted May 2024 · Assignee: Altooro · USPTO

Fellowships