I study molecular sequences and genomes through (gene, species, coalescent) trees (phylogenetic trees, genealogies, taxonomies) using scalable algorithms, probabilistic data structures, and a bit of probabilistic modeling.

During my PhD, I developed several metagenomic sequence analysis tools for efficient taxonomic and/or phylogenomic analysis of large-scale datasets. For instance, CONSULT-II computes Hamming distances of (pseudo)-homologous k-mers for highly sensitive taxonomic classification and accurate abundance profiling. Another, and perhaps a more notable, tool I developed is krepp. Using a more principled probabilistic model, krepp can estimate distances from reads to many thousands of reference genomes and perform phylogenetic placement.

In the past, I briefly worked in network analysis and developed a community detection method for dynamic gene co-expression networks (MuDCoD). During my undergraduate and master’s studies, I focused mostly on (applied) machine learning, in particular time series analysis in computational ethology (behavioral analysis of sleep in fruit flies), and active learning in natural language processing. I have never worked in the industry.