Chapter Four · failure evidence
What Supervised & Unsupervised Learning got wrong, from 44 dissertations
Unsupervised and supervised learning methods encounter recurring failures when applied across tabular, acoustic, text, and imaging datasets. Clustering approaches frequently struggle with cluster discovery, parameter selection, and class alignment, while supervised models can underperform simple heuristics or suffer from corrupted training signals. These records come from PhD theses at 22 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Clustering algorithms suffer from parameter instability and boundary confusion
Clustering models frequently fail to determine appropriate cluster counts or isolate overlapping boundaries across temporal and spatial feature spaces. Automated pipelines also struggle with hyperparameter tuning, opacity, and lower performance compared to simpler clustering variants.
Lost to a baseline
Unsupervised AE+k-means achieved a lower F-statistic (12.4) for heart failure risk separation than k-means applied directly to raw input data (13.2)
Outcome-driven deep clustering for cardiovascular subtyping · OpenBU
Considered and rejected
Considered and rejected: Rejected pure statistical stream clustering (WCDS, FISHDBC) and unsupervised Autoencoders alone due to inability to untangle overlapping temporal boundaries in IMU data.
2022: A Computational Odyssey - Towards a Deeper Understanding of Clustering Streaming Human Activity Recognition Data · Queens University Institutional Repository
Considered and rejected
Considered and rejected: Rejected simple K-means clustering (with purely unsupervised centroids) for algospeak topic categorization due to inability to generate clear, stable, and balanced clusters.
Checking, Moderating, Adapting: Collective Engagement for Digital Integrity · Cornell
Considered and rejected
Considered and rejected: Decided against using unsupervised clustering methods as the primary behavioral classification platform due to difficulty in tuning hyperparameters, lack of ground truth, and opacity of black-box models.
Rage against the machine: advancing aggression ethology through machine learning · ResearchWorks
Considered and rejected
Considered and rejected: Rejected fully automated unsupervised clustering for defining population-level song types due to high-dimensional metric space properties, choosing semi-supervised clustering with manual human validation.
Cultural evolution in the wild: tracking the landscape of learning in bird song · Oxford
Considered and rejected
Considered and rejected: Rejected K-means and Learning Vector Quantization (LVQ) clustering algorithms for unsupervised spectral classification because they required more manual effort and did not automatically select the optimal number of classes.
Methods for quantifying changes to Lesser Prairie-Chicken lek connectivity and habitat · Texas Tech
Tried and failed
cascade k-means clustering for unsupervised classification applied to remote sensing imagery classification. Outcome: worse than baseline. Reason: underperformed compared to x-means in identifying target spectral clusters accurately
Methods for quantifying changes to Lesser Prairie-Chicken lek connectivity and habitat · Texas Tech
Lost to a baseline
Unsupervised BIC-driven cluster selection frequently selected K=1 or K=2, performing worse than fixing cluster count a priori to K=4.
STATISTICAL METHODS FOR IDENTIFYING AGING-RELATED VULNERABILITY STATES · JScholarship
Lost to a baseline
Gaussian Mixture Models (recall 0.239) lost to unsupervised K-means (recall 0.777) on UAV bridge inspection imagery
Unsupervised clustering fails to separate target classes and physical signals
Unsupervised clustering and dimensionality reduction fail to produce distinct separations between target biological conditions, acoustic signals, or sensor responses. Shared spectral overlap, background noise, and physical propagation constraints frequently obscure the underlying classes of interest.
Tried and failed
unsupervised acoustic clustering applied to vessel noise identification. Reason: algorithm grouped co-occurring impulsive sounds instead due to shallow-water low-frequency propagation limits
Tried and failed
unsupervised dimensionality reduction and clustering applied to graph centrality features from electrophysiological signals. Outcome: no signal. Reason: unsupervised projections failed to produce distinct visual clusters separating abnormal from normal nodes
Beyond Brainstorms: Predicting Epilepsy Surgery Outcome with Functional Connectivity Centrality · Harvard
Tried and failed
unsupervised dimensionality reduction for clustering applied to biomarker lipidomics profiles across timepoints. Outcome: no signal. Reason: unsupervised methods failed to clearly separate disease and control classes across combined timepoints
Metabolomics and Machine Learning for Early-Stage Cancer Diagnosis · Georgia Tech
Tried and failed
unsupervised clustering of mass spectrometry profiles applied to meat sensory and aging traits. Outcome: no signal. Reason: clustering failed to separate samples across sensory responses, tenderness metrics, or product age
Tried and failed
unsupervised clustering on mass spectrometry spectra applied to bacterial serovar differentiation. Outcome: no signal. Reason: significant spectral overlap and poor separation among closely related classes
Considered and rejected
Considered and rejected: Dimensionality reduction combined with unsupervised clustering was rejected for classifying timing responses because it failed to isolate classes reflecting temporal scaling.
The Production of Interval Timing Activity in the Primary Visual Cortex of Mice · JScholarship
Considered and rejected
Considered and rejected: Rejected using the unsupervised SPICE/clustering detector pipeline for sperm whale daily presence due to clustering difficulty separating clicks from low-frequency (<10 kHz) anthropogenic and vessel noise.
Passive Acoustic Monitoring of Whales and Ocean Ambient Sound to Inform Management · Cornell
Considered and rejected
Considered and rejected: Rejected relying solely on unsupervised sample clustering/PCA because it obscured fine-grained microbial sensitivity to specific chemical gradients
MICROBIAL GENES, GENOMES AND TAXA ASSOCIATED WITH KEY ASPECTS OF PATHOGENESIS AND BIOGEOCHEMICAL CYCLES · JScholarship
Discovered unsupervised clusters fail to align with predefined expert taxonomies and concepts
Unsupervised topic models and cluster representations often fail to match predefined taxonomies, expert categories, or target performance metrics. Without reference information or inductive biases, discovered groupings lack analyst control and produce uninterpretable structures that still require manual labeling.
Tried and failed
unsupervised topic modeling and clustering applied to document classification into expert taxonomy. Outcome: worse than baseline. Reason: discovered clusters failed to align with predefined expert taxonomies without reference information
Advancements in Models and Algorithms for Management Science · MIT
Tried and failed
unsupervised clustering for symbolic concept discovery applied to neuro-symbolic perception learning. Outcome: no signal. Reason: inability to map learned clusters to meaningful symbolic concepts and scale to complex perception
Neuro-symbolic learning of answer set programs from raw data · Imperial
Tried and failed
unsupervised Latent Dirichlet Allocation applied to text classification against predefined taxonomy. Reason: unsupervised topic clusters did not correspond one-to-one with predefined target taxonomy categories
Analytics-Enabled Quality and Safety Management Methods for High-Stakes Manufacturing Applications · MIT
Tried and failed
Unsupervised topic modeling applied to public comments and reviews text. Outcome: no signal. Reason: Failed to produce theoretically or socially meaningful clusters for downstream analysis
Natural Language Processing and Deep Learning Approaches for Sustainability and Infrastructure Policy Analyses · Georgia Tech
Considered and rejected
Considered and rejected: Rejected unsupervised clustering techniques for project grouping because they do not reflect the relationship between input features and target performance metrics.
Novel approaches to benchmark capital project performance : an application to healthcare projects · UT Austin
Considered and rejected
Considered and rejected: Rejected purely unsupervised automatic clustering for class extraction because it still necessitated manual cluster labeling.
Considered and rejected
Considered and rejected: Rejected unsupervised clustering methods due to insufficient analyst control over high thematic detail category definitions
Generation of a Land Cover Atlas of environmental critic zones using unconventional tools · IRIS - POLITO - prod
Considered and rejected
Considered and rejected: Rejected unsupervised disentangled representation learning without inductive bias/speaker identity because learned representations can be uninterpretable or meaningless.
Exploring Knowledge Transfer with Deep Learning · DukeSpace
Supervised models are beaten by simple baselines and unsupervised alternatives
Supervised machine learning classifiers and neural architectures can underperform simple means, handcrafted statistical tests, or unsupervised clustering baselines. Supervised models also struggle to generalize to unseen attack strategies and perform poorly under high label misclassification.
Lost to a baseline
Mean of training data out-performed all supervised ML models (Ridge, Lasso, ExtraTrees, KernelRidge) across most datasets on held-out perturbation log fold change MSE and MAE.
ASSESSING RELIABILITY OF CAUSAL MODELS OF TRANSCRIPTION · JScholarship
Lost to a baseline
Unsupervised clustering had a lower MAE (6.45 min) than supervised classification with misclassification (13.8 min) on short duration basic feature events.
Machine learning framework for end to end implementation of incident duration prediction · Iowa State
Lost to a baseline
Single-best-score cell-type mapping from RCTD decomposed Stereo-Seq bins performed worse than manual annotation of unsupervised Louvain clustering for discerning rare and nuanced CCS populations.
Lost to a baseline
A 1D chi-square test on the handcrafted invariant mass feature outperformed all supervised multivariate ML classifiers (ANN, BDT, SVM) in discovery power for the two-body decay problem.
The Search for Dark Photons at LHCb and Machine Learning in Particle Physics · MIT
Lost to a baseline
Supervised SVM exhibited higher classification error than unsupervised K-means clustering in largely malicious network scenarios where honest nodes comprised less than 40% of the network.
Identification and mitigation of attacks in trust-based distributed communication networks · Oxford
Lost to a baseline
DeepLabV3+ (recall 0.256) and ViT (recall 0.027-0.321) lost to unsupervised K-means (recall 0.777) on the full UAV-collected dataset
Considered and rejected
Considered and rejected: Rejected standard supervised classifiers trained on specific cyberattack models because they fail to generalize to unseen attack strategies.
Computational Methods for Fast and Secure Distributed Optimal Power Flow · Georgia Tech
Unsupervised representations lose to supervised learning baselines
Unsupervised feature representations, clustering techniques, and alignment algorithms consistently underperform supervised baselines across pose estimation, audio, and disaggregation tasks. Supervised models leverage available annotations to achieve significantly higher accuracy and normalized mutual information.
Lost to a baseline
On BBCPose supervised baselines, Pfister et al. (88.01% avg accuracy) outperformed the thesis's unsupervised 75.93%
Unsupervised landmark discovery via self-training correspondence · University of Nottingham Repository
Lost to a baseline
Unsupervised alignment methods (unsupervised beta-cdf and registr) failed to outperform unaligned modularity/participation curves in predicting executive function in the PNC cohort.
Statistical and Machine Learning Methods for Neuroimaging and Neurocognitive Data · Penn
Lost to a baseline
K-means unsupervised clustering underperformed supervised Random Forest, achieving only 38.2% to 82.4% accuracy compared to RF's 81.6% to 94.4%.
IDENTIFYING MIGRATION FATE AND FACTORS CONTRIBUTING TO MORTALITY OF ATLANTIC SALMON SMOLTS · DalSpace
Lost to a baseline
DUE unsupervised disaggregation had lower Estimation Accuracy than supervised Factorial Hidden Markov Models (FHMM)
Flexibility for large-scale deployment of PV systems in low-voltage grids · EPFL
Lost to a baseline
On Kinetics-Sound clustering, SeLaVi achieved 50.2% NMI and 43.2% accuracy, losing to the supervised baseline which achieved 81.7% NMI and 75.0% accuracy
Learning deep neural networks: necessity and scope of prior knowledge, raw data, and labels · Oxford
Lost to a baseline
Handcrafted/supervised OmniSource Slow-8x8-R101x2 beat thesis's linear transformer-pooling feature representation on HMDB-51 (83.8% vs 81.3%) and UCF-101 (98.6% vs 98.0%)
Learning and interpreting deep representations from multi-modal data · Oxford
Supervised learning fails due to corrupted labels, human rating variance, or objective conflicts
Supervised training degrades when benchmark dataset labels contain noise or when subjective human evaluation variance overwhelms the target signal. Optimization also fails when excessive auxiliary penalty weights divert the model away from its primary supervised learning objective.
Tried and failed
risk-aware loss penalty weighting applied to clinical decision support recommendations. Outcome: worse than baseline. Reason: excessive penalty weight diverted the optimization focus away from the primary supervised learning signal
Tried and failed
supervised classification using unmodified benchmark dataset labels applied to network intrusion detection. Reason: ground-truth labels were noisy because malicious agents exhibited normal message intervals during attacks
RSU-Based Intrusion Detection and Autonomous Intersection Response Systems · Virginia Tech
Tried and failed
direct supervised regression of subjective condition ratings applied to infrastructure asset condition assessment. Outcome: worse than baseline. Reason: models yielded negative R-squared values due to high human inspection variability masking underlying signal
Human Inspection Variability in Infrastructure Asset Management: A Focus on HVAC Systems · Virginia Tech
Left open by the authors
Problems the authors named and did not get to.
Left open
Develop an automated bioacoustic vocalization detection system for unsegmented, in-the-wild recordings using pre-trained audio self-supervised learning representations. Blocker: None
Transferability of Learnt Speech Representations for Decoding Non-Human Vocal Communication · EPFL
Left open
Correlate discrete acoustic tokens from audio self-supervised learning models with acoustic features like pitch and resonances to build species-specific vocalization inventories. Blocker: None
Transferability of Learnt Speech Representations for Decoding Non-Human Vocal Communication · EPFL
Left open
Develop an unsupervised attribute discovery method to learn articulatory speech features directly from audio without predefined phonetic mappings. Blocker: None
Modeling of Language-Universal Speech Attributes for Multilingual Speech Recognition and Processing · Georgia Tech
Left open
Develop an unsupervised algorithm to induce mode hierarchy and rāga specifications from unannotated performance audio or symbolic music data. Blocker: None
Left open
Implement and evaluate supervised machine learning models like random forests to improve the robustness of bioprocess dynamic phase detection. Blocker: Requires dynamic bioprocess time-series run data from the unit operations.
Statistical approaches supporting QbD milestones via bioprocess digital twins · DSpace-CRIS at TU Wien
Left open
Synthesize network classification rules for mixed traffic types using unsupervised clustering or weak rule ensembles. Blocker: None
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.