Chapter Four · failure evidence
What Bioinformatics Homology & Sequence Alignment got wrong, from 71 dissertations
Computational sequence alignment and homology detection methods frequently fail when confronted with extreme evolutionary divergence, complex domain architectures, or massive scale. Across diverse biological datasets, researchers repeatedly encountered failures from nucleotide saturation, spurious k-mer matches, and trimming artifacts, while advanced learned representations often lost to classical homology-guided baselines. These records come from PhD theses at 25 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Extreme evolutionary divergence and low sequence identity prevent reliable sequence alignment and ortholog detection
Standard pairwise search tools and reference-based read mappers failed to detect homologs or align divergent viral and eukaryotic sequences when sequence identity dropped too low. In response, studies rejected relying on sequence alignment alone, turning instead to profile hidden Markov models or structural comparisons.
Tried and failed
reference-based read alignment to consensus genomes applied to highly diverse viral single-cell RNA. Outcome: no signal. Reason: extensive sequence diversity and serotypic variation prevented read mapping to the static reference
Single-Cell Biology of Respiratory Viral Infections in the Nasal Mucosa · Harvard
Tried and failed
standard sequence similarity search applied to divergent protein domain sequence identification. Outcome: no signal. Reason: Pairwise sequence alignment failed to detect highly divergent homologs across distant eukaryotic lineages.
Tried and failed
strict short-sequence alignment matching applied to CRISPR spacer-to-protospacer virus-host linking. Outcome: no signal. Reason: overly strict alignment thresholds failed to match divergent sequences, requiring relaxed search parameters
Novel microbes and viruses with roles in biogeochemical cycling and eukaryogenesis in marine systems · UT Austin
Tried and failed
automated orthology inference software applied to distantly related cross-phylum sequence homology. Outcome: no signal. Reason: failed to detect highly divergent homologous immune receptors across deep evolutionary distances
Considered and rejected
Considered and rejected: Rejected using sequence homology alone for detecting all viral mimics, acknowledging >70% cannot be detected without structural homology.
Evolutionary History of Immunomodulatory Genes of Giant Viruses · Virginia Tech
Considered and rejected
Considered and rejected: Aligning GBS reads of Melampsora paradoxa to the Melampsora americana reference genome R15-033-03 was rejected due to high sequence divergence (37.6% SNP profile similarity, 90.8% ITS identity) causing poor read alignment and reduced SNP quality.
Considered and rejected
Considered and rejected: Rejected k-mer classifiers (e.g., RDP) and phylogenetic placement (e.g., ARB) for ITS1 sequences due to high sequence divergence, non-orthologous lengths, and common indels.
Dynamics of Microeukaryotes and Archaea in the Mammalian Gut Microbiome · Penn
Considered and rejected
Considered and rejected: Decided against relying solely on pairwise BLAST searches for viral ortholog identification in favor of profile HMMs of ViPhOGs to overcome sequence divergence.
On Bioinformatics of the Human Gut Virome · Publikationssystem UB Tuebingen
Considered and rejected
Considered and rejected: Rejected standard BLAST/PSI-BLAST alone because ~60% of target transcriptome dataset was not yet in annotated NCBI databases and full-length alignment over-penalizes variable effector regions.
Bioinformatics Discovery And Functional Characterization Of Lipid-Binding Lov Photoreceptors · Penn
Tried and failed
sequence similarity search for functional annotation applied to candidate divergent loci in non-model genome. Outcome: no signal. Reason: queries failed to produce significant alignment matches in reference databases due to incomplete annotations and sequence divergence
Tried and failed
cross-species reference germline alignment applied to non-model organism antibody repertoire sequencing. Reason: introduced unpredictable sequence alignment biases and annotation errors due to germline divergence
USING SEROLOGY TO UNDERSTAND BAT-VIRUS INTERACTIONS ACROSS SCALES · Cornell
Tried and failed
conventional sequence alignment algorithms applied to highly diverse protein sequences. Outcome: no signal. Reason: extremely low pairwise sequence identity below 15 percent
UNDERSTANDING THE STRUCTRAL DYNAMICS AND EVOLUTIONARY DIVERSIFICATION OF RIBONUCLEOTIDE REDUCTASES · Cornell
Considered and rejected
Considered and rejected: Rejected identifying functional CiGnRH1 enhancers by sequence conservation alone across distant ascidians (Phallusia, Halocynthia), switching to systematic truncation assays due to lack of detectable sequence homology.
Evolution and development of the olfactory neurons in the olfactores clade · Oxford
Single-locus markers and nucleotide saturation fail to resolve phylogenetic relationships
Phylogenetic reconstruction using single marker genes, individual coding sequences, or raw nucleotide sequences suffered from poor resolution, substitution saturation, and horizontal gene transfer artifacts. Analyses were forced to reject these single loci or saturated nucleotide alignments in favor of codon models, concatenated operons, or core genome alignments.
Considered and rejected
Considered and rejected: Aligning and analyzing raw nucleotide sequences for adenovirus phylogenetic analyses due to nucleotide saturation (translated to amino acid alignment instead).
Phylogeography and Signal Evolution in a Widespread Central American Anole · Harvard
Considered and rejected
Considered and rejected: Rejected classifying fragmented RefSeq phage beta-prime-like sequences and eukaryotic RdRp within the standard multi-subunit RNAP phylogeny due to severe long branch attraction artifacts
Unraveling the Eco-Evolutionary Complexity of Uncultivated Bacteriophages in the Biosphere · Virginia Tech
Considered and rejected
Considered and rejected: Rejected amino acid alignment-based phylogeny for HCP-3 histone fold domains due to poor phylogenetic resolution; adopted codon-based nucleotide alignment instead.
The mechanism of the peel-1 zeel-1 toxin-antidote system · ResearchWorks
Considered and rejected
Considered and rejected: Rejected relying on the HERVH reverse transcriptase (RVT) domain for phylogenetic subtyping due to lower resolution compared to LTR sequence phylogenies
RECOMBINATION-MEDIATED REGULATORY EVOLUTION OF HUMAN ENDOGENOUS RETROVIRUS HERVH · Cornell
Considered and rejected
Considered and rejected: Rejected relying purely on species marker trees (rpoB) for GA operon reconstruction because horizontal gene transfer creates incongruence, selecting concatenated core operon alignments instead
Evolution of gene blocks in bacteria using an event-based model · Iowa State
Tried and failed
Bayesian phylogenetic inference using protein sequences applied to highly conserved single-locus protein alignments. Outcome: no signal. Reason: extremely high sequence conservation provided insufficient polymorphic sites to resolve phylogenetic tree topology
Tried and failed
phylogenetic conservation scoring from small multi-species alignments applied to identifying constrained genomic sites. Outcome: no signal. Reason: insufficient phylogenetic depth and taxa number yielded low statistical power to detect conserved positions
Tried and failed
multilocus sequence typing gene concatenation applied to microbial phylogenetic lineage resolution. Outcome: no signal. Reason: few traditional marker genes lacked sufficient resolution to recover whole-genome phylogenetic structures
Expanding The Bioinformatics Toolbox for Diversity and Taxonomic Studies of Microbial Eukaryotic Pathogens · Georgia Tech
Lost to a baseline
ParSNP core genome alignment resulted in poor isolate separation (collapsing tree clusters) compared to Roary with RAxML-NG.
Investigation into different approaches in the fight against antimicrobial resistance: studying the microbiome using human gastric organoids, bioinformatic analysis of a hospital outbreak and predatory bacteria as a therapeutic · University of Nottingham Repository
Considered and rejected
Considered and rejected: Rejected using nucleotide sequences for deep eukaryote MAPK phylogeny due to poor alignment and high substitution rates outside conserved motifs.
STUDIES OF DEEP HOMOLOGY AND PHENOTYPIC EVOLUTION UNDER A PHYLOGENETIC FRAMEWORK · JScholarship
Considered and rejected
Considered and rejected: Rejected binary pangenome presence/absence matrix phylogenies in favor of core-genome nucleotide alignments, because accessory gene monophyly was driven merely by 8 gene losses rather than true evolutionary descent.
Tried and failed
single marker phylogenetic distance applied to inter-phylum horizontal gene transfer prediction. Outcome: no signal. Reason: 16S rRNA sequence distance alone performed no better than random guessing for recent transfer events
Tried and failed
phylogenetic reconstruction using internal coding sequences applied to endogenous retrovirus subfamily classification. Outcome: no signal. Reason: coding sequences provided insufficient granularity to resolve subfamilies compared to non-coding terminal repeats
RECOMBINATION-MEDIATED REGULATORY EVOLUTION OF HUMAN ENDOGENOUS RETROVIRUS HERVH · Cornell
Learned sequence representations and ab initio models underperform traditional homology baselines
Deep learning representations, foundation model embeddings, and ab initio structure predictors struggled with out-of-distribution generalization and lacked critical positional or mechanistic signals. As a result, simpler homology modeling, profile hidden Markov models, and standard sequence similarity searches consistently outperformed or matched these complex architectures.
Tried and failed
pretrained protein language models for host prediction applied to novel viral variant host classification. Outcome: did not generalise. Reason: models relied on ancestral homology rather than tracking cross-species spillover mutations
Discovering Viral Hosts, Mutations, and Diseases using Machine Learning · Virginia Tech
Lost to a baseline
AlphaFold2 ab initio modelling performed worse on structural alignment to ZmVP14 (RMSD 1.58–1.67 Å) compared to homology modelling with Modeller (RMSD 0.28–0.42 Å).
Drought tolerance and its evolution in conifers including UK commercial species · Oxford
Lost to a baseline
Ab initio programs ESMFold and trRosetta were outperformed by homology-guided AlphaFold2 across nearly all GFLV proteins in both estimated TM-scores and pLDDT metrics.
THE HOST RESPONSE TO GRAPEVINE FANLEAF VIRUS AND IMPLICATIONS FOR SYMPTOMATOLOGY AND TRANSMISSION · Cornell
Tried and failed
sliced-Wasserstein multiset embeddings applied to DNA sequence classification. Outcome: worse than baseline. Reason: degraded taxonomic classification accuracy compared to standard k-mer histogram baseline methods
Distance-preserving set embeddings: theory and applications · Harvard
Tried and failed
restricted Boltzmann machine scoring applied to remote protein homology detection. Outcome: did not generalise. Reason: apparent superiority over baselines was an artifact of testing on pre-aligned sequences without insertions
Restricted Boltzmann Machines and Remote Homology Search · Harvard
Tried and failed
HMM-based sequence inference preprocessing applied to phenotype prediction from genomic sequences. Outcome: worse than baseline. Reason: Failed to improve out-of-sample performance over BLOSUM matrix inference and degraded performance versus raw sequences
Lost to a baseline
BLASTp outperformed KA-Search in identifying the closest sequence based on highest BLOSUM62 score (54 vs 27 as top-1; 96 vs 83 within top-100).
Tried and failed
multimodal 3D CNN and transformer sequence embedding applied to remote protein homology retrieval. Outcome: did not generalise. Reason: Learned representation failed to capture cross-family semantic similarity among divergent structural homologs.
Exploring protein biochemistry with deep learning · UT Austin
Tried and failed
increasing sequence alignment depth applied to protein fitness prediction. Outcome: worse than baseline. Reason: distant homologs introduced noise and diverged too far from target sequence characteristics
Learning from pre-pandemic data to design and test future-proof therapeutics · MIT
Tried and failed
entropy weighting on sequence models applied to remote protein homology detection. Outcome: worse than baseline. Reason: flattens emission probabilities and disproportionately inflates decoy scores relative to target sequences
Restricted Boltzmann Machines and Remote Homology Search · Harvard
Lost to a baseline
AlphaFold2 structural predictions for RhGB01 structural proteins exhibited large poorly supported regions and were outperformed in quality scoring parameters by SWISS-MODEL homology modelling.
Tried and failed
marker gene correlation without sequence similarity weighting applied to cross-species single-cell RNA-seq alignment. Outcome: worse than baseline. Reason: omitting sequence homology weighting reduced cell-type mapping accuracy compared to sequence-aware methods
Computational Analysis of Gene Expression in the Teleost Forebrain and the Cellular Basis of a Social Behavior · Georgia Tech
Tried and failed
pretrained sequence foundation model embeddings alone applied to nonsense-mediated decay efficiency prediction. Outcome: worse than baseline. Reason: Embeddings lacked explicit positional and mechanistic features captured by simple rule-based heuristics
Exhaustive pairwise and multiple sequence alignment incur prohibitive computational overhead on large datasets
Computing full all-versus-all alignments, explicit guide trees, or traditional multiple sequence alignments became computationally intractable when scaling to large numbers of reads or pangenomes. Practitioners rejected exhaustive alignment strategies in favor of alignment-free methods, k-mer clustering, iterative matrix merging, or single-sequence embeddings.
Considered and rejected
Considered and rejected: Rejected aligning all family member sequences with ensemble alignment methods due to non-reviewed sequences, excessive compute time, and alignment gaps/noise.
COMBINING HUERISTIC APPROACHES WITH MACHINE LEARNING TO RECOMMEND POST TRANSLATIONAL MODIFICATIONS FOR STUDY · Georgia Tech
Considered and rejected
Considered and rejected: Decided against performing full all-versus-all alignments for every long read, opting instead for RATTLE's two-step bitvector and longest-increasing-subsequence k-mer clustering to avoid prohibitive computational overhead.
Computational Analysis of the Transcriptome Using Long-Read RNA Sequencing · JScholarship
Considered and rejected
Considered and rejected: Rejected Needleman-Wunsch pairwise alignments for all reads in AmpliCI due to high computational cost; adopted an alignment-free conditional probability strategy.
Model-based clustering methods for high-throughput sequencing data · Iowa State
Considered and rejected
Considered and rejected: Rejected traditional consensus-sequence phylogenetic tree building due to inability to handle thousands of long-read genomes and loss of individual genome abundance and direct MRCA-descendant mutation tracking.
HIV evolution during ART failures revealed by using long-read sequencing and bioinformatics tools · ResearchWorks
Considered and rejected
Considered and rejected: Rejected using traditional multiple sequence alignments (MSAs) or whole-genome graph traversals for large pangenome visualization due to prohibitive computational costs and failure under high sequence divergence.
Methods and applications for large-scale pangenomic analysis · JScholarship
Considered and rejected
Considered and rejected: Rejected explicit tree construction (e.g., UPGMA or Neighbor-Joining) for progressive sequence alignment due to high computational burden when aligning many sequences, adopting iterative distance-matrix merging instead.
Developing a Multiday Travel Demand Modelling System · DalSpace
Considered and rejected
Considered and rejected: Rejected multiple sequence alignments (MSAs) in favor of single-sequence PLM embeddings to eliminate high alignment computational overhead
Predicting One-Dimensional Protein Structures by Leveraging Pre-Trained Language Models (PLMs) and Deep Learning · Research Repository UCD
Short k-mers, ambiguity codes, and unvalidated mapping produce spurious matches and chimeric artifacts
Alignment heuristics relying on short k-mer matches, relaxed mapping constraints, or unphased ambiguity codes produced false-positive connections and chimeric distortions. Filtering reads without strict lowest common ancestor validation or failing to screen repetitive motifs caused widespread misclassification and topological errors.
Tried and failed
CRISPR, tRNA, and k-mer sequence matching applied to viral host prediction in metagenomes. Outcome: did not generalise. Reason: Misbinned contigs, absence of Cas proteins near arrays, and broad tRNA sequence sharing caused false-positive host assignments.
Tried and failed
colored de Bruijn graph k-mer filtering applied to bacterial genome evolutionary sequence graph traversal. Reason: Short k-mers caused spurious non-homologous matches, while longer k-mers evaded prefix-suffix screening.
Studies in Bacterial Genome Dynamics · Harvard
Considered and rejected
Considered and rejected: Rejected discard of all host-classified reads; retained reads meeting 25% non-host sequence threshold via kmer_filter.py to maximize bacterial ARG detection from chimeric reads
Tried and failed
reference-based read mapping without taxonomic validation applied to ancient pathogen sequence identification. Outcome: no signal. Reason: mapped candidate reads represented spurious matches that failed strict lowest common ancestor taxonomic validation
Tried and failed
relaxing mapping constraints to at-least-one applied to unsupervised sequence-to-sequence alignment. Outcome: worse than baseline. Reason: allowing multiple correspondences introduces ambiguity compared to enforcing exact one-to-one mapping
Tried and failed
applying fine-tuned sequence model globally applied to whole genome sequencing basecalling. Outcome: worse than baseline. Reason: fine-tuned repeat biases caused systematic miscalling of similar non-target repetitive sequence motifs elsewhere
Delineating genome alterations in cancer with long-read and linked-read sequencing · Harvard
Considered and rejected
Considered and rejected: Rejected using standard unphased IUPAC ambiguity encoding for diploid heterozygous sites because it creates artificial chimeric sequences that distort phylogenetic branch topologies.
Genetic diversity of Candida albicans and Nakaseomyces glabratus across geographic scales and within-host populations · MSpace - University of Manitoba
Variable domain architectures, repeats, and genomic rearrangements disrupt full-length sequence alignment
Aligning full-length sequences failed when homologous proteins possessed conflicting module histories, variable repeat copy numbers, or distinct domain architectures. These structural discrepancies caused alignment gaps and missed remote homologies, prompting the rejection of whole-protein alignments in favor of domain-focused strategies.
Tried and failed
direct multiple sequence alignment across distant taxa applied to divergent multi-domain homologous proteins. Outcome: no signal. Reason: mismatched N-terminal domain architectures between distant homologues prevented accurate alignment
Bioinformatic Sequence Analysis Reveals Evolutionary History of Synaptic Genes · Harvard
Tried and failed
single-round sequence similarity clustering applied to viral genomes with structural rearrangements. Reason: failed to account for genomic rearrangements, producing clusters lacking unique marker material
Considered and rejected
Considered and rejected: Rejected using whole-protein sequence alignment for homology mining because variable repeat copy numbers and large sequence length discrepancies cause alignment failure and miss remote homologues.
Diversity and Evolution of Cyclic Peptides in Flax (Linum usitatissimum L.) · HARVEST
Considered and rejected
Considered and rejected: Rejected standard sequence homology alignment tools (such as BLAST, ClustalW, MUSCLE, MAFFT) for functional prediction because non-homologous proteins (e.g., Peptides A and D with 19.0% similarity) share identical binding functionality.
Signal processing-based bioinfomatics methods for characterisation and identification of bio-functionalities of proteins · De Montfort Open Research Archive (DORA)
Considered and rejected
Considered and rejected: Aligning full-length multimodular NRPS sequences or whole modules for single phylogenetic tree construction was rejected due to large gaps and conflicting phylogenetic histories between A- and C-domains.
Drift, selection, and convergence in the evolution of a nonribosomal peptide · open_UMR Marburg DSpace 10.0
Tried and failed
short subregion sequence alignment for novelty estimation applied to bacterial genome novelty prediction. Outcome: did not generalise. Reason: subregion sequences have higher sequence identity than full-length sequences, systematically underestimating divergence
Alignment trimming and gap filtering remove informative sites without improving phylogenetic inference
Automated multiple sequence alignment trimming and filtering of gap-rich columns frequently eliminated well-aligned, phylogenetically informative positions from protein datasets. In multiple evaluations, trimmed alignments generated identical phylogenetic tree topologies and divergence times compared to untrimmed alignments, leading authors to abandon trimming.
Tried and failed
automated multiple sequence alignment trimming applied to divergent protein sequence datasets. Reason: consistently excluded blocks of well-aligned, informative phylogenetic sites
The evolution and diversity of noncanonical microbial nitrogen metabolisms · MIT
Tried and failed
alignment trimming prior to phylogenetic inference applied to concatenated maximum-likelihood phylogenomics. Outcome: no signal. Reason: trimmed alignments produced identical phylogenetic tree topologies to untrimmed alignments
The evolution of anthraquinones as an adaptive trait in lichen-forming fungi · Imperial
Tried and failed
filtering gap-rich columns in sequence alignments applied to phylogenetic divergence time estimation. Outcome: no signal. Reason: removing gap positions yielded no significant improvement in divergence time estimates
Bayesian methods for source attribution using HIV deep sequence data · Imperial
Considered and rejected
Considered and rejected: Omission of primer binding sites in all rDNA sequence alignments (16S, 12S, 18S) prior to phylogenetic analysis.
Considered and rejected
Considered and rejected: Rejected using trimmed alignments for concatenated whole-genome phylogenetics after finding untrimmed alignments produced identical topologies.
The evolution of anthraquinones as an adaptive trait in lichen-forming fungi · Imperial
Left open by the authors
Problems the authors named and did not get to.
Left open
Investigate mathematical and sequence correlations between the HIV-1 genome and the human host genome. Blocker: Lack of specific mathematical formulation, target human genomic regions, or methodology definition
Left open
Develop computational models to map within-host viral evolutionary dynamics and spatial dissemination across tissue compartments using multi-tissue sequence data. Blocker: None
Examining viral pathogen evolution and spread through genomic data · Harvard
Left open
Develop and test a convolutional neural network that takes sequence alignment matrices as input to predict outbreak characteristics such as infection timing. Blocker: None
Evolution and population dynamics of mycobacterial pathogens with applications for disease control · Cornell
Left open
Test whether HSV-1 viral genomes exhibiting differential chromatin accessibility are differentially packaged into viral capsids. Blocker: Requires wet-lab experimental biology techniques to isolate capsids and sequence encapsidated viral genomes.
Left open
Basecall plasmid Nanopore Fast5 files using methylation-aware and standard modes, then compare error profiles against reference sequences. Blocker: Requires the raw Nanopore Fast5 sequencing data generated from the specific plasmid templates in the thesis.
ASSESING amplified aene aemplate quality with Nanopore sequencing for improved cell free expression · Iowa State
Left open
Test whether ST258 OmpK36 porin mutations confer similar carbapenem resistance phenotypes when introduced into Klebsiella pneumoniae sequence type ST307. Blocker: Requires a wet microbiology lab and genetic engineering protocols to introduce mutations into K. pneumoniae ST307 strains and measure resistance.
OmpK36-mediated carbapenem resistance in Klebsiella pneumoniae · Imperial
Left open
Analyze resistance genes in genome assemblies to identify predominant resistance mechanisms in poorly characterized Pseudomonas species. Blocker: Requires unpublished isolate sequence data or specific isolate strains generated in the thesis
Left open
Determine whether dual organellar localization of the haem biosynthesis pathway occurs across other rhodophytes using experimental or targeting sequence analysis. Blocker: Determining protein localization typically requires wet lab experiments (e.g. GFP tagging, microscopy, cell fractionation) or unvalidated computational predictions.
The convoluted history of haem biosynthesis · Cambridge
Left open
Analyze RNA sequences to discover sequence motifs or features associated with transcript colocalization patterns. Blocker: No specific hypothesis, sequence modeling method, or concrete target dataset provided beyond a broad research direction
Decoding Spatial Transcriptomics: Computational Frameworks for Subcellular Spatial Patterns and Contact Mediated Signaling · Georgia Tech
Left open
Infer macroinvertebrate phylogeny using combined nucleotide and amino acid sequence alignments. Blocker: None
Characterizing freshwater macroinvertebrates of Bangladesh using metagenetic techniques · Imperial
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.