Chapter Four · failure evidence

What Supervised & Unsupervised Learning got wrong, from 44 dissertations

Unsupervised and supervised learning methods encounter recurring failures when applied across tabular, acoustic, text, and imaging datasets. Clustering approaches frequently struggle with cluster discovery, parameter selection, and class alignment, while supervised models can underperform simple heuristics or suffer from corrupted training signals. These records come from PhD theses at 22 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Clustering algorithms suffer from parameter instability and boundary confusion

8 theses · 8 institutions

Clustering models frequently fail to determine appropriate cluster counts or isolate overlapping boundaries across temporal and spatial feature spaces. Automated pipelines also struggle with hyperparameter tuning, opacity, and lower performance compared to simpler clustering variants.

Lost to a baseline

Unsupervised AE+k-means achieved a lower F-statistic (12.4) for heart failure risk separation than k-means applied directly to raw input data (13.2)

Outcome-driven deep clustering for cardiovascular subtyping · OpenBU

Considered and rejected

Considered and rejected: Rejected pure statistical stream clustering (WCDS, FISHDBC) and unsupervised Autoencoders alone due to inability to untangle overlapping temporal boundaries in IMU data.

2022: A Computational Odyssey - Towards a Deeper Understanding of Clustering Streaming Human Activity Recognition Data · Queens University Institutional Repository

Considered and rejected

Considered and rejected: Rejected simple K-means clustering (with purely unsupervised centroids) for algospeak topic categorization due to inability to generate clear, stable, and balanced clusters.

Checking, Moderating, Adapting: Collective Engagement for Digital Integrity · Cornell

Considered and rejected

Considered and rejected: Decided against using unsupervised clustering methods as the primary behavioral classification platform due to difficulty in tuning hyperparameters, lack of ground truth, and opacity of black-box models.

Rage against the machine: advancing aggression ethology through machine learning · ResearchWorks

Considered and rejected

Considered and rejected: Rejected fully automated unsupervised clustering for defining population-level song types due to high-dimensional metric space properties, choosing semi-supervised clustering with manual human validation.

Cultural evolution in the wild: tracking the landscape of learning in bird song · Oxford

Considered and rejected

Considered and rejected: Rejected K-means and Learning Vector Quantization (LVQ) clustering algorithms for unsupervised spectral classification because they required more manual effort and did not automatically select the optimal number of classes.

Methods for quantifying changes to Lesser Prairie-Chicken lek connectivity and habitat · Texas Tech

Tried and failed

cascade k-means clustering for unsupervised classification applied to remote sensing imagery classification. Outcome: worse than baseline. Reason: underperformed compared to x-means in identifying target spectral clusters accurately

Methods for quantifying changes to Lesser Prairie-Chicken lek connectivity and habitat · Texas Tech

Lost to a baseline

Unsupervised BIC-driven cluster selection frequently selected K=1 or K=2, performing worse than fixing cluster count a priori to K=4.

STATISTICAL METHODS FOR IDENTIFYING AGING-RELATED VULNERABILITY STATES · JScholarship

Lost to a baseline

Gaussian Mixture Models (recall 0.239) lost to unsupervised K-means (recall 0.777) on UAV bridge inspection imagery

Enhancing Structural Inspections by Integrating Computer Vision, Machine Learning, and Risk Assessment for Comprehensive Corrosion Evaluation · Georgia Tech

Unsupervised clustering fails to separate target classes and physical signals

8 theses · 5 institutions

Unsupervised clustering and dimensionality reduction fail to produce distinct separations between target biological conditions, acoustic signals, or sensor responses. Shared spectral overlap, background noise, and physical propagation constraints frequently obscure the underlying classes of interest.

Tried and failed

unsupervised acoustic clustering applied to vessel noise identification. Reason: algorithm grouped co-occurring impulsive sounds instead due to shallow-water low-frequency propagation limits

Understanding the Physical Environment of a French Polynesian Coral Reef Ecosystem Using Ocean Modeling, Acoustics, And Machine Learning · Georgia Tech

Tried and failed

unsupervised dimensionality reduction and clustering applied to graph centrality features from electrophysiological signals. Outcome: no signal. Reason: unsupervised projections failed to produce distinct visual clusters separating abnormal from normal nodes

Beyond Brainstorms: Predicting Epilepsy Surgery Outcome with Functional Connectivity Centrality · Harvard

Tried and failed

unsupervised dimensionality reduction for clustering applied to biomarker lipidomics profiles across timepoints. Outcome: no signal. Reason: unsupervised methods failed to clearly separate disease and control classes across combined timepoints

Metabolomics and Machine Learning for Early-Stage Cancer Diagnosis · Georgia Tech

Tried and failed

unsupervised clustering of mass spectrometry profiles applied to meat sensory and aging traits. Outcome: no signal. Reason: clustering failed to separate samples across sensory responses, tenderness metrics, or product age

Evaluating the Ability of Rapid Evaporative Ionization Mass Spectrometry to Predict the Palatability of Long Aged Beef Cuts · Texas Tech

Tried and failed

unsupervised clustering on mass spectrometry spectra applied to bacterial serovar differentiation. Outcome: no signal. Reason: significant spectral overlap and poor separation among closely related classes

Liver Abscess Effects on Rumen Histology, Morphology, and Feeding Behavior in BeefÍDairy Cattle, Surveillance of Salmonella Enterica Throughout the Beef Carcass, and Characterizati · Texas Tech

Considered and rejected

Considered and rejected: Dimensionality reduction combined with unsupervised clustering was rejected for classifying timing responses because it failed to isolate classes reflecting temporal scaling.

The Production of Interval Timing Activity in the Primary Visual Cortex of Mice · JScholarship

Considered and rejected

Considered and rejected: Rejected using the unsupervised SPICE/clustering detector pipeline for sperm whale daily presence due to clustering difficulty separating clicks from low-frequency (<10 kHz) anthropogenic and vessel noise.

Passive Acoustic Monitoring of Whales and Ocean Ambient Sound to Inform Management · Cornell

Considered and rejected

Considered and rejected: Rejected relying solely on unsupervised sample clustering/PCA because it obscured fine-grained microbial sensitivity to specific chemical gradients

MICROBIAL GENES, GENOMES AND TAXA ASSOCIATED WITH KEY ASPECTS OF PATHOGENESIS AND BIOGEOCHEMICAL CYCLES · JScholarship

Discovered unsupervised clusters fail to align with predefined expert taxonomies and concepts

8 theses · 7 institutions

Unsupervised topic models and cluster representations often fail to match predefined taxonomies, expert categories, or target performance metrics. Without reference information or inductive biases, discovered groupings lack analyst control and produce uninterpretable structures that still require manual labeling.

Tried and failed

unsupervised topic modeling and clustering applied to document classification into expert taxonomy. Outcome: worse than baseline. Reason: discovered clusters failed to align with predefined expert taxonomies without reference information

Advancements in Models and Algorithms for Management Science · MIT

Tried and failed

unsupervised clustering for symbolic concept discovery applied to neuro-symbolic perception learning. Outcome: no signal. Reason: inability to map learned clusters to meaningful symbolic concepts and scale to complex perception

Neuro-symbolic learning of answer set programs from raw data · Imperial

Tried and failed

unsupervised Latent Dirichlet Allocation applied to text classification against predefined taxonomy. Reason: unsupervised topic clusters did not correspond one-to-one with predefined target taxonomy categories

Analytics-Enabled Quality and Safety Management Methods for High-Stakes Manufacturing Applications · MIT

Tried and failed

Unsupervised topic modeling applied to public comments and reviews text. Outcome: no signal. Reason: Failed to produce theoretically or socially meaningful clusters for downstream analysis

Natural Language Processing and Deep Learning Approaches for Sustainability and Infrastructure Policy Analyses · Georgia Tech

Considered and rejected

Considered and rejected: Rejected unsupervised clustering techniques for project grouping because they do not reflect the relationship between input features and target performance metrics.

Novel approaches to benchmark capital project performance : an application to healthcare projects · UT Austin

Considered and rejected

Considered and rejected: Rejected purely unsupervised automatic clustering for class extraction because it still necessitated manual cluster labeling.

A hybrid machine learning and text-mining approach for the automated generation of early warnings in construction project management. · Cranfield

Considered and rejected

Considered and rejected: Rejected unsupervised clustering methods due to insufficient analyst control over high thematic detail category definitions

Generation of a Land Cover Atlas of environmental critic zones using unconventional tools · IRIS - POLITO - prod

Considered and rejected

Considered and rejected: Rejected unsupervised disentangled representation learning without inductive bias/speaker identity because learned representations can be uninterpretable or meaningless.

Exploring Knowledge Transfer with Deep Learning · DukeSpace

Supervised models are beaten by simple baselines and unsupervised alternatives

7 theses · 5 institutions

Supervised machine learning classifiers and neural architectures can underperform simple means, handcrafted statistical tests, or unsupervised clustering baselines. Supervised models also struggle to generalize to unseen attack strategies and perform poorly under high label misclassification.

Lost to a baseline

Mean of training data out-performed all supervised ML models (Ridge, Lasso, ExtraTrees, KernelRidge) across most datasets on held-out perturbation log fold change MSE and MAE.

ASSESSING RELIABILITY OF CAUSAL MODELS OF TRANSCRIPTION · JScholarship

Lost to a baseline

Unsupervised clustering had a lower MAE (6.45 min) than supervised classification with misclassification (13.8 min) on short duration basic feature events.

Machine learning framework for end to end implementation of incident duration prediction · Iowa State

Lost to a baseline

Single-best-score cell-type mapping from RCTD decomposed Stereo-Seq bins performed worse than manual annotation of unsupervised Louvain clustering for discerning rare and nuanced CCS populations.

An investigation of the relationship between transcriptomic signals and functional phenotypes across atrial physiology and pharmacology · Oxford

Lost to a baseline

A 1D chi-square test on the handcrafted invariant mass feature outperformed all supervised multivariate ML classifiers (ANN, BDT, SVM) in discovery power for the two-body decay problem.

The Search for Dark Photons at LHCb and Machine Learning in Particle Physics · MIT

Lost to a baseline

Supervised SVM exhibited higher classification error than unsupervised K-means clustering in largely malicious network scenarios where honest nodes comprised less than 40% of the network.

Identification and mitigation of attacks in trust-based distributed communication networks · Oxford

Lost to a baseline

DeepLabV3+ (recall 0.256) and ViT (recall 0.027-0.321) lost to unsupervised K-means (recall 0.777) on the full UAV-collected dataset

Enhancing Structural Inspections by Integrating Computer Vision, Machine Learning, and Risk Assessment for Comprehensive Corrosion Evaluation · Georgia Tech

Considered and rejected

Considered and rejected: Rejected standard supervised classifiers trained on specific cyberattack models because they fail to generalize to unseen attack strategies.

Computational Methods for Fast and Secure Distributed Optimal Power Flow · Georgia Tech

Unsupervised representations lose to supervised learning baselines

6 theses · 5 institutions

Unsupervised feature representations, clustering techniques, and alignment algorithms consistently underperform supervised baselines across pose estimation, audio, and disaggregation tasks. Supervised models leverage available annotations to achieve significantly higher accuracy and normalized mutual information.

Lost to a baseline

On BBCPose supervised baselines, Pfister et al. (88.01% avg accuracy) outperformed the thesis's unsupervised 75.93%

Unsupervised landmark discovery via self-training correspondence · University of Nottingham Repository

Lost to a baseline

Unsupervised alignment methods (unsupervised beta-cdf and registr) failed to outperform unaligned modularity/participation curves in predicting executive function in the PNC cohort.

Statistical and Machine Learning Methods for Neuroimaging and Neurocognitive Data · Penn

Lost to a baseline

K-means unsupervised clustering underperformed supervised Random Forest, achieving only 38.2% to 82.4% accuracy compared to RF's 81.6% to 94.4%.

IDENTIFYING MIGRATION FATE AND FACTORS CONTRIBUTING TO MORTALITY OF ATLANTIC SALMON SMOLTS · DalSpace

Lost to a baseline

DUE unsupervised disaggregation had lower Estimation Accuracy than supervised Factorial Hidden Markov Models (FHMM)

Flexibility for large-scale deployment of PV systems in low-voltage grids · EPFL

Lost to a baseline

On Kinetics-Sound clustering, SeLaVi achieved 50.2% NMI and 43.2% accuracy, losing to the supervised baseline which achieved 81.7% NMI and 75.0% accuracy

Learning deep neural networks: necessity and scope of prior knowledge, raw data, and labels · Oxford

Lost to a baseline

Handcrafted/supervised OmniSource Slow-8x8-R101x2 beat thesis's linear transformer-pooling feature representation on HMDB-51 (83.8% vs 81.3%) and UCF-101 (98.6% vs 98.0%)

Learning and interpreting deep representations from multi-modal data · Oxford

Supervised learning fails due to corrupted labels, human rating variance, or objective conflicts

3 theses · 2 institutions

Supervised training degrades when benchmark dataset labels contain noise or when subjective human evaluation variance overwhelms the target signal. Optimization also fails when excessive auxiliary penalty weights divert the model away from its primary supervised learning objective.

Tried and failed

risk-aware loss penalty weighting applied to clinical decision support recommendations. Outcome: worse than baseline. Reason: excessive penalty weight diverted the optimization focus away from the primary supervised learning signal

Enhancing the Reliability of Real-World Evidence and Clinical Decision-Making: Robust, Calibrated, and Uncertainty-Aware Methods for Observational Healthcare Research · Penn

Tried and failed

supervised classification using unmodified benchmark dataset labels applied to network intrusion detection. Reason: ground-truth labels were noisy because malicious agents exhibited normal message intervals during attacks

RSU-Based Intrusion Detection and Autonomous Intersection Response Systems · Virginia Tech

Tried and failed

direct supervised regression of subjective condition ratings applied to infrastructure asset condition assessment. Outcome: worse than baseline. Reason: models yielded negative R-squared values due to high human inspection variability masking underlying signal

Human Inspection Variability in Infrastructure Asset Management: A Focus on HVAC Systems · Virginia Tech

Left open by the authors

Problems the authors named and did not get to.

Left open

Develop an automated bioacoustic vocalization detection system for unsegmented, in-the-wild recordings using pre-trained audio self-supervised learning representations. Blocker: None

Transferability of Learnt Speech Representations for Decoding Non-Human Vocal Communication · EPFL

Left open

Correlate discrete acoustic tokens from audio self-supervised learning models with acoustic features like pitch and resonances to build species-specific vocalization inventories. Blocker: None

Transferability of Learnt Speech Representations for Decoding Non-Human Vocal Communication · EPFL

Left open

Develop an unsupervised attribute discovery method to learn articulatory speech features directly from audio without predefined phonetic mappings. Blocker: None

Modeling of Language-Universal Speech Attributes for Multilingual Speech Recognition and Processing · Georgia Tech

Left open

Develop an unsupervised algorithm to induce mode hierarchy and rāga specifications from unannotated performance audio or symbolic music data. Blocker: None

The Structure of Free Polyphony · EPFL

Left open

Implement and evaluate supervised machine learning models like random forests to improve the robustness of bioprocess dynamic phase detection. Blocker: Requires dynamic bioprocess time-series run data from the unit operations.

Statistical approaches supporting QbD milestones via bioprocess digital twins · DSpace-CRIS at TU Wien

Left open

Synthesize network classification rules for mixed traffic types using unsupervised clustering or weak rule ensembles. Blocker: None

Practical Network Programming Automation · Penn

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.