Chapter Four · failure evidence

What Random Forests & Tree Ensembles got wrong, from 91 dissertations

Tree ensembles and random forests frequently underperform simpler linear baselines or alternative architectures such as gradient boosting and neural networks across diverse tasks. Researchers also encounter severe failures related to overfitting on high-dimensional data, poor calibration on extreme values, class imbalance degradation, and uninterpretable model complexity. These records come from PhD theses at 29 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Tree ensembles underperformed simpler linear and parametric baselines

22 theses · 18 institutions

Across multiple studies, random forests and decision trees achieved lower accuracy or higher error compared to multiple linear regression, logistic regression, and generalized linear models. These models struggled when linear baselines captured the underlying structure better or when tree ensembles failed to generalize on continuous outcomes and spatial surveys.

Tried and failed

generalized random forest heterogeneous treatment effect estimation applied to continuous outcome treatment effect estimation. Outcome: worse than baseline. Reason: generalized random forest performed poorly estimating heterogeneous treatment effects on continuous variables compared to binary outcomes

THREE ESSAYS ON E-COMMERCE DEVELOPMENT AND INEQUALITY IN CHINA · Cornell

Lost to a baseline

Random Forest models exhibited poorer performance and greater subgroup disparities compared to simpler logistic regression models.

The application of machine learning and causal inference to improve suicide outcomes among U.S. veterans: a focus on clinic characteristics · OpenBU

Lost to a baseline

Decision tree predictive R2 (0.436 with DRS, 0.260 without for FPC1) and Random Forest (0.479 / 0.370) were lower than Elastic Net (0.550 / 0.450).

Identifying Risk Factors for Cognitive Decline Using Statistical Learning Techniques and Functional Data Analysis · HARVEST

Lost to a baseline

Random Forest (ensemble method) performed substantially worse than the simpler single Classification Tree, achieving only 42.43% agreement with BorealDB versus CT's 84.88%.

Exploiting Overlapping Landsat Scene Classifications and Focal Context to Identify Boreal Disturbance Mapping Uncertainty · YorkSpace

Lost to a baseline

Random Forest classifier (macro F1-score 0.87, accuracy 0.88) lost to standard Decision Tree classifier (macro F1-score 0.94, accuracy 0.94) on EMG gesture recognition.

EXPLORING THE POTENTIAL OF COMPUTER VISION AND MACHINE LEARNING IN ENHANCING THE FUNCTIONALITY OF AN EMG-CONTROLLED PROSTHETIC HAND · HARVEST

Lost to a baseline

Random Forest classifier (AUC=0.707) was beaten by standard logistic regression (AUC=0.752)

Prediction of second line treatment using comorbidities and patient characteristics in newly diagnosed multiple myeloma patients : a machine learning approach · UT Austin

Lost to a baseline

Random forest variable selection underperformed simpler stepwise AIC and LASSO regression in variable recovery and discrimination/calibration metrics

Enhancing Survival Prediction Models: Insights on Biomarker Inclusion and Model Updating · JScholarship

Lost to a baseline

Random forest regression models did not significantly outperform multiple linear regression, leading to favoring simpler MLR models.

ANALYZING THE IMPACT OF U.S. CENTRAL COMMAND OPERATIONS AND INVESTMENTS IN ITS AREA OF RESPONSIBILITY: AN OPEN SOURCE DATA APPROACH · Calhoun

Lost to a baseline

Random Forest trained on RAVDESS+MESD had lower external validity on TESS than simple Linear Regression and Lasso.

Creating Links: Building an Educational Platform to Ask Relevant Questions in Education · MIT

Tried and failed

random forest classification applied to imbalanced tabular student outcome prediction. Outcome: worse than baseline. Reason: lower AUC and specificity compared to regularised linear models and gradient boosting across imbalanced panels

Using Machine Learning to Advance High School Dropout Prediction and Prevention · Penn

Tried and failed

random forest classification and survival modeling applied to clinical tabular outcome prediction. Outcome: worse than baseline. Reason: complex machine learning models failed to outperform classical logistic regression and Cox proportional hazards baselines

Prediction of second line treatment using comorbidities and patient characteristics in newly diagnosed multiple myeloma patients : a machine learning approach · UT Austin

Tried and failed

Random forests and k-nearest neighbors classification applied to optimization oracle indicator function approximation. Outcome: worse than baseline. Reason: Non-linear classifiers yielded poor generalization accuracy compared to linear baselines like logistic regression.

Consistency in integer programming · Iowa State

Tried and failed

decision tree classifiers applied to sparse bag-of-words text classification. Outcome: worse than baseline. Reason: decision trees produced poor and inconsistent classification performance compared to logistic regression

Andromeda in Education: Studies on Student Collaboration and Insight Generation with Interactive Dimensionality Reduction · Virginia Tech

Tried and failed

random forest regression applied to predicting physical parameters from relaxometry data. Outcome: worse than baseline. Reason: non-linear ensemble failed to outperform a simpler quadratic general linear model baseline

Quantitative Microstructural Imaging for Clinical Use · EPFL

Tried and failed

random forest regression replacing linear mixed-effects models applied to spatial land use regression modeling. Outcome: worse than baseline. Reason: provided no improvement in predictive accuracy over linear mixed-effects regression

Spatiotemporal patterns and socioeconomic inequalities of noise and sound sources in Accra, Ghana · Imperial

Lost to a baseline

Random Forest (0.519 accuracy) and Bagging Classifier (0.618 accuracy) underperformed compared to simpler Decision Tree (0.994) and Logistic Regression (0.984).

Measuring and Analyzing Community Resilience During COVID-19 Using Social Media · Virginia Tech

Lost to a baseline

Random forest (AUC = 0.6345) was outperformed by full logistic regression (AUC = 0.65991) on event status prediction.

ENHANCING NAVAL RETENTION: A STRATEGIC APPROACH TO ALLOCATING SELECTIVE REENLISTMENT BONUSES · Calhoun

Lost to a baseline

Random forest on JDTC data achieved 75.00% test accuracy, beaten slightly by standard logistic regression at 77.08%.

Risk, Need, and Racial Inequality: A Machine Learning Analysis of Rearrest in Juvenile Drug Treatment Courts and Traditional Juvenile Courts · unevada

Considered and rejected

Considered and rejected: Rejected standard/complex classifiers (SVMs, Random Forests) for Rim Thickness Curves in glaucoma diagnosis because logistic regression performed equally or better

Leveraging Uncertainties in Medical Prediction Systems · Publikationssystem UB Tuebingen

Tried and failed

gradient boosted trees with TF-IDF features applied to short text classification. Outcome: worse than baseline. Reason: linear model handles high-dimensional sparse TF-IDF text features better than decision trees

Urban Housing in the Digital Age: Three Applications of Natural Language Processing · Cornell

Tried and failed

random forest regression on functional connectivity features applied to predicting behavioral retrieval accuracy. Outcome: worse than baseline. Reason: connectivity features in the baseline condition lacked predictive signal, yielding negative explained variance

Musical Context Facilitates Event Segmentation and Sequential Learning Through Interconnected Neural Networks and Strengthened Hippocampal Encoding · Georgia Tech

Tried and failed

random forest for small area estimation applied to hierarchical zero-inflated domain data. Outcome: worse than baseline. Reason: predictors lacked domain-level random intercept information, causing massive bias when domains varied independently of features

A Forest for the Trees: Using Random Forests for Small Area Estimation on US Forest Inventory Data · Harvard

Tried and failed

random forests for small area estimation applied to zero-inflated spatial survey data. Outcome: worse than baseline. Reason: failed to outperform linear and mixed-effects models across varied sample sizes, correlations, and target variable ranges

A Forest for the Trees: Using Random Forests for Small Area Estimation on US Forest Inventory Data · Harvard

Lost to a baseline

Random forest models including full factor sets failed to outperform simpler multiple linear regression models in explaining variance in evolutionary rate or purifying selection.

The evolution of the angiosperm plastid genome · Oxford

Random forests suffered from severe overfitting on sparse or high dimensional data

22 theses · 13 institutions

Models fitted training noise and created discontinuous decision boundaries when trained on high-dimensional features or small sample sizes. Consequently, performance dropped sharply on holdout cohorts, validation environments, and pruned feature sets.

Tried and failed

random forest classification applied to high-dimensional sparse genomic variant features. Outcome: overfit. Reason: severe overfitting during training on high-dimensional genomic features compared to logistic regression baseline

Decoding Germline Genetic Influence on Cancer Somatic Mutation Acquisition · Harvard

Tried and failed

random forest classification on low-dimensional biomarker data applied to disease diagnosis prediction. Outcome: overfit. Reason: models formed discontinuous islands and narrow decision bands fitting training noise

Machine-Learning Aided Diagnosis Of Alzheimer's Disease · UT Austin

Tried and failed

random forest regression applied to oceanic tracer distribution modeling. Outcome: overfit. Reason: prone to overfitting on sparse data and underperformed gradient boosting

Environmental pollution as a marine tracer: measuring and modelling anthropogenic lead in the ocean · Imperial

Tried and failed

Random forest regression on dynamic response features applied to crack depth prediction in concrete. Outcome: overfit. Reason: High-dimensional feature space on concrete dynamic response data led to significant overfitting.

Machine learning (ml) approaches to model interdependencies between dynamic loads and crack propagation · Cranfield

Tried and failed

random forest for chemical property prediction applied to mass spectral chromatographic data. Outcome: did not generalise. Reason: pretrained random forest model suffered significant accuracy drop when applied to new experimental chromatographic dataset

A Statistical Methods-Based Novel Approach for Fully Automated Analysis of Chromatographic Data · Virginia Tech

Tried and failed

random forest package implementation for probability estimation applied to discrete choice consideration set formation. Outcome: overfit. Reason: Default implementation overfitted severely, assigning nearly perfect probability to observed alternatives unlike Ranger

Endogeneity and consideration-set issues in residential location choice models · Imperial

Tried and failed

random forest regression applied to tabular spatial feature regression. Outcome: overfit. Reason: Severe overfitting to training data, high computational training time, and poor spatial variance compared to gradient boosting.

Spatial Modeling for Building Design Evaluation: from Visual Landscape Quality Assessment to Devaluation Risk Estimation · EPFL

Tried and failed

random forest classification without feature selection applied to high-dimensional molecular descriptor dataset. Outcome: overfit. Reason: training on large number of molecular descriptors without selection led to severe overfitting and poor AUC

Enabling Data-Driven Experimentation for High-Performance Polymer Thin Film Formulations · Georgia Tech

Lost to a baseline

Random forest cross-validation R2 (0.658) lost to simpler multiple linear regression cross-validation R2 (0.733) due to overfitting.

Spatiotemporal patterns of urbanization during the last four decades in Switzerland and their impacts on urban heat islands · EPFL

Lost to a baseline

Random Forest classifier test accuracy dropped more severely (from 0.78 to 0.69) compared to Support Vector Machine (stayed at 0.71) when pruned to top four LIME-selected features due to overfitting on the full feature set.

Multiscale Integration of Cross-Modal Subsurface Data for Reservoir Characterization under Label-Constrained Environments · Georgia Tech

Considered and rejected

Considered and rejected: Rejected complex non-linear regressors (Random Forest, Gradient Boosting) for speech emotion prediction because they overfit the training dataset and sacrificed cross-dataset generalizability.

Creating Links: Building an Educational Platform to Ask Relevant Questions in Education · MIT

Considered and rejected

Considered and rejected: Rejected Random Forest as the final deployed estimator due to lack of interpretability, high memory footprint, and risk of overfitting on small datasets.

Computational Intelligence for Computer-Aided Design Machine Learning Techniques for Microcontrollers Performance Screenings · IRIS - POLITO - prod

Considered and rejected

Considered and rejected: Rejected Random Forest and KNN regression for destination factor mapping in favor of Multiple Linear Regression due to severe overfitting and poor interpretability.

Leveraging the subtle : hidden factors in recommender systems · DSpace-CRIS at TU Wien

Considered and rejected

Considered and rejected: Decided against random forest and exploratory factor analysis for indicator reduction due to overfitting risks on high-dimensional data.

Suspended sediment transport in rivers: new indicators of transport dynamics for analysis of catchment and climate controls · Cranfield

Considered and rejected

Considered and rejected: Decision tree learning for generalized model development (due to producing variable behavioral thresholds that were overfitted to individual runways)

Incorporation of Causal Factors Affecting Pilot Motivation for Improvement of Airport Runway and Exit Design Modeling · Virginia Tech

Tried and failed

direct execution time prediction with random forest applied to database query execution time modeling. Outcome: overfit. Reason: over-optimized slow queries at the expense of fast ones; predicting decomposed weight parameters was required

Building Instance Aware Systems using Explicit Performance Modeling · MIT

Lost to a baseline

In alternate environment validation, simpler CART and Logistic Regression models (balanced accuracy 79.7% and 78.9%) outperformed Random Forest and AdaBoost models which suffered severe overfitting.

COUNTERING SMALL UNMANNED AIRCRAFT SYSTEMS WITH ADVANCED DATA ANALYSIS AND MACHINE LEARNING · Calhoun

Tried and failed

gradient boosted trees with default hyperparameters applied to weakly supervised anomaly detection. Outcome: overfit. Reason: default tree hyperparameters were overly aggressive for weakly supervised learning labels

LOOKING FOR NEW PHYSICS: FROM DARK MATTER TO MACHINE LEARNING · Cornell

Tried and failed

gradient boosted decision tree classification applied to high-dimensional metabolomics profiles. Outcome: did not generalise. Reason: model failed to perform better than chance on an independent hold-out cohort

DEFINING THE METABOLIC LANDSCAPE OF FIBROLAMELLAR CARCINOMA · Cornell

Tried and failed

random forest on raw radar backscatter features applied to surface water extent mapping. Outcome: did not generalise. Reason: raw backscatter polarizations and incidence angle lacked sufficient discriminative power across diverse environmental conditions

Automated quantification of depressional water storage in prairie pothole landscapes using synthetic aperture radar and random forest classification · Iowa State

Considered and rejected

Considered and rejected: Rejected Tree Parzen Estimators and Random Forests as surrogate models in BO under tight computational budgets due to over-partitioning/overfitting with few evaluations

Exploration and Exploitation Techniques for High-Dimensional Simulation-Based Optimization Problems in Urban Transportation · MIT

Tried and failed

random forest hyperparameter tuning via grid search applied to time series error correction. Outcome: overfit. Reason: overemphasized lag-1 features, degrading downstream simulation accuracy and inflating variance

HYDRO-METEOROLOGICAL UNCERTAINTY QUANTIFICATION FOR WATER RESOURCES PLANNING AND MANAGEMENT: ADVANCES IN SYNTHETIC FORECASTING AND STOCHASTIC WATERSHED MODELS · Cornell

Tree ensembles were outperformed by gradient boosting frameworks and neural networks

14 theses · 10 institutions

Random forests and standard decision trees often lost to algorithms like XGBoost, LightGBM, neural networks, and support vector machines on complex prediction tasks. These competing architectures better captured continuous spatial gradients, complex temporal horizons, and sequential error corrections.

Considered and rejected

Considered and rejected: Rejected using Random Forest models in favor of XGBoost due to extreme computational overhead and boundary overfitting.

Spatial Modeling for Building Design Evaluation: from Visual Landscape Quality Assessment to Devaluation Risk Estimation · EPFL

Tried and failed

random forest regression applied to satellite retrieval bias correction. Outcome: worse than baseline. Reason: significantly underperformed gradient boosted decision tree algorithms like LightGBM and XGBoost

Correcting Biases in Satellite Methane Observations: Applications to Landfill Emissions in the United States and Continental-Scale Emissions in Africa · Harvard

Tried and failed

gradient boosted decision trees regression applied to multimodal extracted image and process features. Outcome: did not generalise. Reason: poorly captured non-linear relationships compared to neural network representations

ML-accelerated pipeline for understanding atomistic hardening of MPEAs and predicting hardness in additive manufacturing · Virginia Tech

Tried and failed

gradient boosted trees for long-horizon regression applied to renewable resource time-series forecasting. Outcome: worse than baseline. Reason: Tree ensembles struggled with complex temporal patterns over long-term forecasting horizons compared to other models.

Machine learning applications for the optimization of renewable energy systems · Iowa State

Tried and failed

gradient boosted decision trees applied to cluster membership classification. Outcome: worse than baseline. Reason: cascading weak learner errors caused poor classification performance

A network-level statistical model for friction estimation in Texas · UT Austin

Tried and failed

random forest regression for field variable surrogate applied to turbulent flow field prediction. Outcome: worse than baseline. Reason: random forest failed to capture complex spatial gradients compared to deep learning framework

Multifidelity machine learning methods for flow field prediction and aerodynamic shape optimization · Iowa State

Lost to a baseline

Random Forest on 5-stage PU UV degradation achieved only 59% accuracy with O-PTIR and 60% with ATR-FTIR, underperforming PLS-DA (80% and 51%) and SVM (82% on O-PTIR).

An investigation on the applications of advanced Infrared Spectroscopy, Spectral Imaging and Machine Learning for Polymer Characterization, including microplastics · Research Repository UCD

Lost to a baseline

Smooth Random Forest underperformed DPGBDT 1x100 on the HELOC dataset.

A consumer centred investigation of differentially private risk assessment models in consumer credit · University of Nottingham Repository

Lost to a baseline

Support Vector Machine (SVM), Random Forests, and Gaussian Processes achieved lower test dataset R2 compared to the 2-hidden-layer ANN model for predicting HFRC tensile stress-strain behavior.

Condition Assessment of Civil Infrastructure and Materials Using Deep Learning · Virginia Tech

Lost to a baseline

Single Decision Tree for HEA phase classification achieved only 67.16% mean CV accuracy, trailing Gradient Boosting (72.99%).

Accelerated Data-Driven Design of Multi-Component Chemistries for Enhanced Mechanical Performance · DSpace at SUNY Buffalo

Lost to a baseline

Gradient boosting (Model 4 R2 0.46–0.92) slightly outperformed random forest (Model 3 R2 0.33–0.98).

Small Effect, Big Picture: Understanding the Role of Transposable Elements, Pleiotropy, and Single-Genes on Maize Genetics and Yield · Cornell

Considered and rejected

Considered and rejected: Rejected Decision Trees, Random Forests, Support Vector Machines, Regularized Linear Regression, and Multi-Layer Perceptrons in favor of XGBoost for Hurricast prediction due to lower performance and higher training time.

Multimodal Machine Learning for Climate Adaptation · MIT

Considered and rejected

Considered and rejected: Random forest classifier rejected in favour of XGBoost due to lower test accuracy and inferior prediction calibration across SES and gender subgroups

Understanding Student Outcomes: The Role of Background, Gender, and Schools in Irish Education · Research Repository UCD

Considered and rejected

Considered and rejected: Rejected Random Forest classification in favor of Gradient Boosted Decision Trees (XGBoost) due to XGBoost's sequential error correction and higher expected performance

Using Machine Learning to Identify Risk Factors Associated with Causes of Fetal Deaths in the United States, 2021 · UT Austin

Models were rejected due to opacity and computational overhead compared to simpler models

10 theses · 8 institutions

Researchers rejected random forest models because their high complexity and lack of interpretability obscured individual variable effects. In addition, tree ensembles required substantially greater runtime and memory footprint without offering accuracy advantages over simpler decision trees or linear alternatives.

Lost to a baseline

Random Forest classifier consumed approximately 8x more runtime than Decision Trees despite achieving comparable accuracy, precision, and recall.

Dynamic Decomposition and Deployment of the Virtual Network Functions with Microservices using Deep Reinforcement Learning · Research Repository UCD

Considered and rejected

Considered and rejected: Rejected Random Forest classification in mlZ in favor of Decision Trees due to lack of accuracy improvement and increased complexity

Characterizing VNTRs in human populations · OpenBU

Considered and rejected

Considered and rejected: Random forest regression rejected in favor of linear regression due to loss of interpretability, poor extrapolation, and inability to handle memory-bound workloads without full target-state counter inputs.

Compute Overlap Stall (COS): Predicting Performance of Power Management for Shared Memory Codes When Throttling Processors, Memory, and Thread Concurrency · Virginia Tech

Considered and rejected

Considered and rejected: Rejected Random Forest as the final model due to its high complexity and lack of interpretability compared to Decision Trees.

Exploring Phishing Detection Using Search Engine Optimization and Uniform Resource Locator based Information · DalSpace

Considered and rejected

Considered and rejected: Rejected Random Forest in favor of Logistic Regression due to logistic regression's superior enlisted AUC (0.836 vs 0.695), interpretability/explainability, and reduced computational complexity.

U.S. ARMY RESERVE RETENTION MODELING FOR MID-LEVEL LEADERS · Calhoun

Considered and rejected

Considered and rejected: Rejected random forest and classification tree models in favor of logistic regression because logistic regression provides clearer interpretability for individual variable effects.

PREDICTING MIDSHIPMEN'S OUTCOMES AT THE UNITED STATES NAVAL ACADEMY · Calhoun

Considered and rejected

Considered and rejected: Random forest regression in favor of multiple linear regression due to similar performance and greater simplicity/interpretability of MLR.

ANALYZING THE IMPACT OF U.S. CENTRAL COMMAND OPERATIONS AND INVESTMENTS IN ITS AREA OF RESPONSIBILITY: AN OPEN SOURCE DATA APPROACH · Calhoun

Considered and rejected

Considered and rejected: Rejected using non-linear and random forest models to focus solely on linear logistic regression of binary outcomes.

Modeling Delays in Foreign-Backed Infrastructure Projects: A BRI Case Study on Inequality, Corruption, and Governance · Harvard

Considered and rejected

Considered and rejected: Tree-based models (such as XGBoost and Random Forests) and Support Vector Machines were rejected in favor of LASSO regression due to multicollinearity issues, extensive hyperparameter tuning requirements, and lack of automatic feature elimination

Tumor Immunogenicity Unlocked: Multi-omics Models Predict Immunotherapy Response and Survival · Publikationssystem UB Tuebingen

Considered and rejected

Considered and rejected: Rejected machine learning (neural networks/random forests) due to high false-positive rates and uninterpretable feature explanations.

Development of computational tools for variant calling in single-cell RNAseq · Oxford

Tree models suffered from poor recall and erratic behavior under class imbalance

7 theses · 6 institutions

When applied to imbalanced target distributions, tree ensembles exhibited poor sensitivity and scattered, noisy predictions. Automated feature elimination and resampling strategies often discarded weak signals essential for detecting rare minority classes.

Lost to a baseline

Random forest demographic-only model yielded an artifactually high sensitivity of 0.308 due to rare outcome imbalances, performing erratically compared to logistic regression baselines.

PREDICTING SUICIDE: UTILITY OF SOCIAL DETERMINANTS OF HEALTH AND ACCESS DATA · JScholarship

Lost to a baseline

Non-parametric models (Random Forest and XGBoost) were outperformed by simpler parametric models (Logistic Regression and LDA) in sensitivity and balanced accuracy when using synthetic resampling (e.g., LR achieved 0.60 sensitivity vs RF 0.32 and XGBoost 0.36 under SMOTE).

Evaluating Factors Contributing to Crash Severity Among Older Drivers: Statistical Modeling and Machine Learning Approaches · Virginia Tech

Lost to a baseline

Logit model achieved higher specificity (0.876) on unbalanced rCSI test data than Random Forest (0.348) and XGBoost (0.627)

Essays in Applied Economics Linking Policy, Obesity, and Health Economics · Texas Tech

Tried and failed

gradient boosted decision trees classification applied to imbalanced fine-grained time-series state prediction. Reason: severe class imbalance produced noisy, scattered predictions with low precision and recall

Understanding and Predicting Sit-Stand Desk Usage Patterns and Willingness among Knowledge Workers: A Data-Driven Approach · Virginia Tech

Tried and failed

recursive feature elimination with cross validation applied to gradient boosted trees on imbalanced data. Outcome: worse than baseline. Reason: feature selection discarded weakly predictive features crucial for detecting the rare minority class

Optimizing Suicide Prevention in Adolescents: A Longitudinal Approach to Risk Modeling and Resource Allocation · Harvard

Lost to a baseline

Traditional Random Forest and Gradient-Boosted Trees suffered from lower recall (0.304 and 0.607, respectively) on HCM classification with CMR features compared to SVM-RBF (0.857) due to class imbalance.

Interpretable machine learning models for investigating and pre-screening cardiomyopathies with high-dimensional genomic and imaging data · Imperial

Considered and rejected

Considered and rejected: Rejected machine learning approaches (decision trees and random forest) for sensor-drone assignment due to poor constraint handling, class imbalance, and lack of generalizability

Development of a Natural Disaster Information System for RPAS Mission Planning and Monitoring · Queens University Institutional Repository

Feature importance metrics and hyperparameter tuning methods failed or introduced bias

7 theses · 6 institutions

Impurity-based feature importance exhibited systematic bias toward continuous variables and features with many categories. Furthermore, hyperparameter tuning via grid search or genetic algorithms often failed to beat untuned base models or removed predictive feature shortcuts.

Tried and failed

random forest mean decrease in impurity applied to feature importance estimation. Reason: biased toward continuous variables and features with many categories

Automated methods to find high-redshift quasars · Imperial

Tried and failed

hyperparameter grid search for random forest applied to tabular demographic and electoral feature classification. Outcome: worse than baseline. Reason: regularization removed single-feature shortcuts that had driven unconstrained performance

Gerrymandering or “Gloria”-mandering? An examination of redistricting effects on women candidates for the U.S. House of Representatives · Harvard

Considered and rejected

Considered and rejected: Rejected evaluating genetic algorithm fitness using Random Forest instead of Logistic Regression because Random Forest's inherent feature importance mechanisms mitigate redundancy and obscure individual feature contributions during GA optimization.

Non-Invasive cancer detection: computational applications in liquid biopsy and radiomics · IRIS - UNITN - prod

Tried and failed

untuned random forest with few estimators applied to blood biomarker disease classification. Outcome: worse than baseline. Reason: insufficient ensemble size and lack of hyperparameter tuning caused extremely poor classification performance

Machine-Learning Aided Diagnosis of Alzheimer's Disease · UT Austin

Lost to a baseline

Feature selection on gas-phase adiabatic ionization potential using random forest failed to improve performance and slightly worsened ionization potential predictions relative to full feature sets.

Using Data-Driven Models to Understand Transition Metal Catalyst Energy Landscapes and Metal-Organic Framework Stability · MIT

Tried and failed

random forest Bayesian optimization for configuration tuning applied to runtime configuration optimization. Outcome: worse than baseline. Reason: direct flag-to-runtime mapping struggled to navigate high-dimensional configuration spaces effectively compared to random search

Accelerating regression testing through test environment tuning · UT Austin

Lost to a baseline

In Experiment 2, Grid Search and Genetic Algorithm tuning of Random Forest (MSE 0.0061, R2 0.8086) lost to the base untuned Random Forest (MSE 0.0042, R2 0.8158).

Using Data Analytics and Machine Learning in Sustainable Forest Management from Remote Sensing Data · YorkSpace

Output variance compression caused inaccurate probability and extreme value predictions

5 theses · 5 institutions

Averaging mechanisms across trees compressed prediction variance, leading models to systematically over-predict low values and under-predict high values. As a result, probability estimates degraded near distribution boundaries and failed to capture extreme tail exceedances.

Tried and failed

random forest regression on continuous clinical scores applied to blood biomarker disease severity prediction. Reason: compressed output variance caused regression to fail at extreme values, resulting in high false positives

Machine-Learning Aided Diagnosis Of Alzheimer's Disease · UT Austin

Tried and failed

indirect mapping predicting multi-item components via random forests applied to health-related utility score estimation. Reason: systematically under-predicted low scores and over-predicted high scores, inflating the overall predicted mean

DEVELOPING BETTER TREATMENT FOR ALCOHOL USE DISORDERS WITH THREE TOOLS: DECISION THEORY, BIG DATA, AND MACHINE LEARNING · Cornell

Tried and failed

random forests with tail oversampling and residual fitting applied to extreme value exceedance prediction. Reason: techniques were ineffective at resolving tail smoothing in predictions

Data-Driven Methods for Modeling Emissions and Atmospheric Composition · Harvard

Considered and rejected

Considered and rejected: Decided against Random Forests for surrogate modeling because of large prediction errors on sampled points and lack of reliable predictive variance estimates.

Accelerating HLS Autotuning of Large, Highly-parameterized Reconfigurable SoC Mappings · Penn

Tried and failed

random forest probability estimation applied to molecular synthetic accessibility prediction. Reason: uncalibrated models yielded poor probability estimates near extreme thresholds

Discovery of synthesisable organic materials · Imperial

Left open by the authors

Problems the authors named and did not get to.

Left open

Perform a grid search over hyperparameter space instead of random search for the Random Forest and XGBoost Alzheimer's diagnosis models. Blocker: Access to the underlying clinical biomarker dataset used in the thesis.

Machine-Learning Aided Diagnosis Of Alzheimer's Disease · UT Austin

Left open

Evaluate varying decision tree depths and random forest models for embedding subspace mapping in neural network generalization prediction. Blocker: None

DEPENDABLE NEURAL NETWORKS FOR SAFETY CRITICAL TASKS · JScholarship

Left open

Prospectively test decision tree and random forest models across diverse healthcare systems to evaluate clinical efficacy in childhood genetic epilepsy. Blocker: Requires prospective clinical deployment and access to private electronic medical record data across multiple healthcare systems.

Quantitative Informatics Approaches to Characterize and Predict Childhood Genetic Epilepsies · Penn

Left open

Evaluate statistical and systematic uncertainty behaviors under Random Forest, SVM, and deep neural network models on student response coding datasets. Blocker: Access to private student physics response datasets used in the thesis.

Evaluating language models applied to student thinking about experiments · Cornell

Left open

Benchmark healthcare capital project performance by evaluating Random Forests, KNN, and DEA grouping algorithms on larger datasets. Blocker: Requires larger proprietary or restricted healthcare capital project BIM and performance datasets

Novel approaches to benchmark capital project performance : an application to healthcare projects · UT Austin

Left open

Develop online prediction for the random forest framework in aerodynamic shape optimization and test alternative architectures like GNN, RNN, and GAN. Blocker: Lack of specific evaluation metrics, clear baseline integration details, or designated datasets for the proposed alternative ML architectures.

Multifidelity machine learning methods for flow field prediction and aerodynamic shape optimization · Iowa State

Left open

Implement decision trees or random forests to combine covert channel detection statistical tests instead of logistic regression. Blocker: None

Real-Time Detection of Storage Covert Channels · Carleton University Institutional Repository

Left open

Train and compare alternative machine learning regression models against random forest for estimating the truncation parameter of truncated exponential distributions. Blocker: None

Estimation of Parameters for Truncated Exponential Distribution · unevada

Left open

Evaluate the reduced feature subsets produced by the fuzzy feature selection methods using classifiers like SVM and Random Forest. Blocker: None

An investigation of fuzzy methods and meta learning for feature selection · University of Nottingham Repository

Left open

Evaluate alternative surrogate models and acquisition functions beyond Random Forest and Expected Improvement for tuning LSM trees in Onix. Blocker: None

Dynamically tuning LSM tree based databases · OpenBU

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.