Chapter Four · failure evidence
What Transformers & Attention Models got wrong, from 100 dissertations
The evaluated doctoral records demonstrate that transformers and attention mechanisms frequently underperform simpler baselines across diverse domains including time series, computer vision, and tabular prediction. Across these experiments, researchers repeatedly encountered severe overfitting, prohibitive computational costs, and degraded accuracy when modifying core attention dynamics or deploying large architectures without sufficient inductive bias. These records come from PhD theses at 29 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Transformers frequently underperform recurrent neural networks and simpler statistical baselines on time series and sequential forecasting tasks
Practitioners observed that pure transformers suffer from high computational scaling, lack local temporal inductive biases, and frequently overfit compared to recurrent models like LSTMs and GRUs or simpler linear baselines. In addition, exposure bias during autoregressive rollouts and difficulties handling raw sensor or trajectory signals led authors to reject or replace standalone transformer architectures.
Tried and failed
time series transformer applied to asset price forecasting. Outcome: worse than baseline. Reason: underperformed simpler autoregressive baseline on predictive accuracy metrics
An Economic Evaluation of NFTs: A Systematic Framework for Prediction and Analysis with Other Comparable Assets · Texas Tech
Tried and failed
transformer architecture for sequence prediction applied to vehicle trajectory forecasting. Outcome: overfit. Reason: transformer backbone suffered severe overfitting compared to recurrent network baseline
Prediction of Traffic Agents' Trajectories in Urban Scenarios Using Deep Learning · Cranfield
Tried and failed
fine-tuning pre-trained transformers with limited data applied to multivariate time series forecasting. Outcome: worse than baseline. Reason: overfitting to seasonal and occupancy patterns degraded performance below zero-shot baseline
Toward Transformer-based Large Energy Models for Smart Energy Management · Virginia Tech
Tried and failed
spatio-temporal transformer with self-attention applied to continuous multi-sensor time series prediction. Outcome: worse than baseline. Reason: performed worse than recurrent baseline while increasing computation without accuracy gains
A Deep-learning based Approach for Foot Placement Prediction · Virginia Tech
Tried and failed
Transformer architecture alone for anomaly detection applied to time-series anomaly detection. Outcome: worse than baseline. Reason: Standalone Transformer encoder underperformed recurrent neural network architectures in sequential state anomaly detection
An assessment of deep cyber-physical situational awareness of power system using real-time testbed · Texas Tech
Tried and failed
transformer models on univariate inputs applied to univariate time series forecasting. Outcome: worse than baseline. Reason: single-dimension embeddings provide insufficient representation capacity compared to recurrent models
A Framework for Generalizing Uncertainty in Mobile Network Traffic Prediction · Virginia Tech
Tried and failed
clustering time series before training transformer forecasters applied to mobile network traffic prediction. Outcome: worse than baseline. Reason: partitioning data into clusters degraded transformer prediction accuracy compared to unclustered baseline
A Framework for Generalizing Uncertainty in Mobile Network Traffic Prediction · Virginia Tech
Tried and failed
physics-informed transformer predicting temporal differences applied to indoor temperature forecasting. Outcome: worse than baseline. Reason: poor percentage error performance during periods when target delta values are near zero
Machine learning on building energy modeling : city and building scale · UT Austin
Lost to a baseline
Transformer trajectory predictor exhibited higher Euclidean error than LSTM on longer prediction horizons (>180s) due to exposure bias between teacher forcing and autoregressive inference.
A Data-driven Methodology for Aircraft Trajectory Analysis to Improve Mid-air Conflict Detection in Terminal Airspace · Georgia Tech
Lost to a baseline
Transformer predictor model achieved higher test RMSE (0.583 vs 0.532) and lower Pearson R (0.818 vs 0.846) compared to the simpler bidirectional LSTM on MMP1.
Amplifying signals in the tumor microenvironment for drug development and diagnostics · MIT
Lost to a baseline
Informer Transformer (0.6° COG MAE, 0.506 kt SOG MAE) was beaten by the simpler 2-layer FNN baseline (0.0093° COG MAE, 0.0167 kt SOG MAE).
USING SPATIAL-TEMPORAL MACHINE LEARNING TO ANALYZE AND PREDICT SHIP MOVEMENTS · Calhoun
Lost to a baseline
Transformer model (MAPE 0.205, IA 0.483) lost to the baseline ARIMAX Top 3 model (MAPE 0.179, IA 0.893).
An Economic Evaluation of NFTs: A Systematic Framework for Prediction and Analysis with Other Comparable Assets · Texas Tech
Lost to a baseline
Transformer model lost significantly to GLM with tCCA (Transformer: RMSE ~0.5410, Corr ~0.3028 vs. GLM with tCCA: RMSE ~0.1117, Corr ~0.9708).
Lost to a baseline
Standard Transformer backbone performed worse than GRU and LSTM on scanpath reconstruction and fixation pre-tasks.
Low-Resource Neural Adaptation: A Unified Data Adaptation Framework for Neural Networks · ResearchWorks
Considered and rejected
Considered and rejected: Rejected pure Transformer sequence modeling in Oh, opting for hybrid attention + LSTM because pure Transformers perform inconsistently on time-series.
From Coarse Single Theoretical to Fine Multiple Practical: Machine Learning-Empowered Real-Time Communications · IRIS - POLITO - prod
Considered and rejected
Considered and rejected: Rejected pure transformer without convolutional layers for raw sensor time series because multi-head attention quadratic scaling makes 10,080-length minute windows infeasible on commodity GPUs
Multimodal Models of Time Series and Text · ResearchWorks
Considered and rejected
Considered and rejected: Standard Transformer ANN architecture was rejected for multivariate PV power forecasting because it cannot natively take multivariate input features without architectural modification.
Investigation of innovative IoT technologies for complex manufacturing process modelling and optimization · IRIS - POLITO - prod
Considered and rejected
Considered and rejected: Rejected using only a single unified transformer for social trajectory forecasting because mixing individual feature extraction with inter-agent attention degraded accuracy.
Deep Generative Models for Autonomous Driving: from Motion Forecasting to Realistic Image Synthesis · EPFL
Considered and rejected
Considered and rejected: Rejected full transformers for long-context meta-RL due to O(t^2) computational cost and empirical failure to learn on T-LS.
Representations in zero-shot meta-reinforcement learning · Oxford
Considered and rejected
Considered and rejected: Transformers and S4 sequence models were considered to model FC temporal dynamics but rejected because vanilla LSTMs performed sufficiently well and preliminary transformer tests showed no huge impact.
Mapping dynamic brain networks with MEG data using machine learning · Oxford
Considered and rejected
Considered and rejected: Rejected Transformer-based spatio-temporal architectures for deforestation forecasting in favor of ConvLSTM due to worse performance on low-dimensional, locally structured spatial data.
Spatio-temporal deep learning for understanding and forecasting tropical deforestation · EPFL
Considered and rejected
Considered and rejected: Rejected TD-Transformer (Crossformer) for building-level downstream evaluations due to poorer effectiveness and high compute/training time (20-27h)
Toward Transformer-based Large Energy Models for Smart Energy Management · Virginia Tech
Tried and failed
pure self-attention without convolutional feature extraction applied to multichannel time-series prediction. Outcome: worse than baseline. Reason: lacked local receptive fields and temporal inductive bias needed for raw signal representation
Investigation of Future Voluntary Movement Prediction for Pathological Tremor-Alleviating Exoskeletons · Virginia Tech
Considered and rejected
Considered and rejected: Self-attention and Transformer architectures were rejected in favor of vanilla GRU, which inherently weights recent timesteps higher and achieves faster inference.
A Deep-learning based Approach for Foot Placement Prediction · Virginia Tech
Tried and failed
increasing Transformer layer depth applied to sensor time series calibration. Outcome: overfit. Reason: deeper architectures led to overfitting or diminishing returns without performance improvements
A fine-scale spatiotemporal air quality modeling framework by combining big but noisy data · Imperial
Tried and failed
sinusoidal positional encoding in transformers applied to small tabular time-series datasets. Outcome: worse than baseline. Reason: outperformed by explicit temporal feature concatenation on limited training data
Developing Transferable Deep Models for Mobile Health · Georgia Tech
Lost to a baseline
Spatio-Temporal Transformer (STT) performed worse than simple Feedforward Neural Networks (FNN) across simpler disaggregation tasks (e.g., PUMA to NTA).
In Pursuit of High-Quality Urban Data · ResearchWorks
Complex multi-transformer architectures and stacked attention layers underperform simpler designs or lightweight baselines
Adding extra self-attention layers, stacking dual transformers, or introducing complex hybrid modules regularly underperformed single transformers, standard decoders, or simple feedforward networks. These elaborate configurations introduced training instability, excessive runtime latency, and error accumulation without delivering predictive gains.
Tried and failed
zero future masking shift in transformer event modeling applied to irregular event sequence prediction. Outcome: worse than baseline. Reason: setting shift to zero reduced predictive performance compared to positive shift windows
Tried and failed
Transformer on execution history sequences applied to database query performance prediction. Outcome: worse than baseline. Reason: representation leakage and query runtime variability when omitting system logs
Machine Learning for Out of Distribution Database Workloads · MIT
Lost to a baseline
The Dual-Transformer + LSTM architecture (macro F1 0.718–0.721) was beaten in Precision, macro F1, and Kappa by the simpler Single-Transformer + LSTM architecture (macro F1 0.728).
Automatic Analysis of Epistemic Stance-Taking in Academic English Writing: A Systemic Functional Approach · Scholars' Bank
Tried and failed
adding self-attention layer to recurrent neural network applied to sequential event classification. Outcome: worse than baseline. Reason: did not improve predictive performance over standard recurrent model despite increasing interpretability
Generalised Modeling of Inquiry Behaviour: From Learning to Understanding · EPFL
Lost to a baseline
Double Self-Attention model (AUROC 80.50% / 87.10%) was outperformed by standard black-box BiLSTM (AUROC 81.24% / 88.02%) on MIMIC-III and eICU delirium prediction
Machine learning applications in Intensive Care Unit · IRIS - UNITN - prod
Lost to a baseline
For User-1 overall accuracy, Self-Attention (0.72) and BiRNN (0.69) were beaten by individual baseline class predictions, or BiRNN (0.69) was beaten by Self-Attention (0.72)
Bridging Non-Intrusive Load Monitoring and Human Activity Recognition: A Multimodal Deep Learning Approach · Carleton University Institutional Repository
Considered and rejected
Considered and rejected: Rejected running self-attention before ON-LSTM (SA-ON-LSTM) to capture sentence context; CEON-LSTM performed better (71.08 vs 70.13 F1).
Structure-based Models for Neural Information Extraction · Scholars' Bank
Considered and rejected
Considered and rejected: Rejected full self-attention within modalities in TCaF, as cross-modal bottleneck attention significantly improves cross-modal generalisation.
Towards Better Video Understanding through Language Guidance · Publikationssystem UB Tuebingen
Considered and rejected
Considered and rejected: Using pre-extracted spatial features to train the temporal attention network was rejected in favor of temporal features (LSTM hidden states), which achieved superior classification performance (AUC 0.77 vs 0.73).
Spatial-temporal attention for video-based assessment of intraoperative surgical skill · JScholarship
Considered and rejected
Considered and rejected: Enforcing self-attention within CRTNet's transformer decoder layers was omitted because it provided no observed performance improvements over cross-attention.
Out-of-Distribution Generalization in Biological and Artificial Intelligence. · Harvard
Considered and rejected
Considered and rejected: Rejected self-attention mechanisms in TAIN's CS module in favor of cross-attention between tentative frame predictions (queries) and input frames (keys/values)
Motion Boundary and Occlusion Reasoning for Video Analysis · DukeSpace
Considered and rejected
Considered and rejected: Rejected causal attention in favor of bidirectional attention because bidirectional attention achieved 3-5% higher PR-AUC while early detection could still be simulated via temporal masking.
TITAN: TRANSFORMER-BASED IDENTIFICATION AND TRACKING OF ANOMALOUS NAVIGATION · Calhoun
Lost to a baseline
Transformer lost to Pointer-Generator on diversity (Div-2) on the Filtered Reddit dataset (0.06% vs 4.00%)
Modeling context and knowledge for dialogue generation · University of Nottingham Repository
Considered and rejected
Considered and rejected: Rejected Transformers for neural deduction and soft unification because their fixed number of encoder layers cannot be iterated for an arbitrary number of time-steps required by multi-step logical inference, and their self-attention benefits were not directly relevant to short sequences (<= 20 symbols).
End-to-end neuro-symbolic learning of logic-based inference · Imperial
Tried and failed
transformer post-training quantization methods applied to state space models. Outcome: worse than baseline. Reason: outlier distributions and architecture differences in state space models cause severe perplexity degradation
Efficient model adaptation and compression for edge intelligence · UT Austin
Tried and failed
replacing dense linear layers with structured sparse approximations applied to transformer blocks. Outcome: worse than baseline. Reason: accumulated approximation error across all layers caused model quality degradation
Structured Sparsity-Aware Hardware-Software Co-Design for Deep Neural Network Acceleration · Georgia Tech
Tried and failed
Transformer architecture for autoencoding applied to moderate-size text autoencoders. Outcome: worse than baseline. Reason: failed to outperform recurrent neural networks on moderate-sized datasets
Lost to a baseline
Transformer without symmetry regularizer lost to Transformer with permutation-invariance regularizer on cyclical target f3 test error (8.14 ± 0.9 vs 0.95 ± 0.13).
Learning to Reason with Neural Networks: Principles and Methods · EPFL
Lost to a baseline
Transformer-RNN Model achieved higher F-beta score (0.967) and lower validation loss (0.04) than H-TST (F-beta 0.952, valid loss 0.05) on CICIoT2023
Optimizing Distributed Denial of Service (DDoS) Detection with Time Series Transformers · Carleton University Institutional Repository
Lost to a baseline
Standard GRU and Transformer slightly outperformed the complex R-Transformer on EVSE charging power misconfiguration detection (F-score 0.47 vs 0.46)
Data driven detection of misconfigurations in power distribution systems · DSpace-CRIS at TU Wien
Lost to a baseline
The dual-Transformer architecture underperformed both single-Transformer and Transformer+LSTM architectures on overall F1 (F1 = 0.71 vs. 0.72 and 0.73) and precision in random hyperparameter searches.
Automatic Analysis of Epistemic Stance-Taking in Academic English Writing: A Systemic Functional Approach · Scholars' Bank
Considered and rejected
Considered and rejected: Rejected pure BERT or more complex NLP transformers in favor of TFIDF logistic regression for interpretability, simplicity, and labeled dictionary generation.
Considered and rejected
Considered and rejected: Rejected Transformer classifier in favor of MLP for practical deployment because Transformer training time was almost 3x longer with identical accuracy (0.977).
A Unified Autotuning Framework for Deep Learning on the Cloud-Edge Continuum · TXST Digital Repository
Considered and rejected
Considered and rejected: Applying PCA dimensionality reduction to transformer embeddings before fitting regressions was tested and rejected because it decreased model performance and reduced interpretability.
Considered and rejected
Considered and rejected: Rejected encoder-decoder Transformer architectures in favor of decoder-only models to avoid computational overhead in autoregressive generation and simplify beam search constraints.
Deep Generative Models for Trajectory Prediction and Mobility Network Forecasting · YorkSpace
Vision transformers struggle with local spatial structures and high sample requirements compared to convolutional networks
Authors found that pure vision transformers lack the spatial inductive biases of convolutional architectures on volumetric, nonstationary, and video data. When trained on small datasets or exposed to large domain gaps, vision transformers suffered from severe overfitting, long training times, and degraded accuracy compared to standard convolutional baselines.
Tried and failed
temporal transformers instead of factorised spatio-temporal convolutions applied to video regression tasks. Outcome: worse than baseline. Reason: None
Deep learning for medical ultrasound video understanding and generation · Imperial
Considered and rejected
Considered and rejected: Rejected Transformers for cardiac ultrasound video LVEF regression in favor of ResNet 2+1D because transformers failed to outperform convolutions along temporal dimensions.
Towards autonomous diagnostic systems with medical imaging · Imperial
Tried and failed
general-purpose self-supervised vision transformers applied to fine-grained visual skill assessment. Outcome: worse than baseline. Reason: recognition-optimized representations fail to capture subtle domain-specific articulatory quality nuances
Advancing Phonology-Based Sign Language Assessment: From Learner to Machine-Generated Videos · EPFL
Tried and failed
pure vision transformer with smaller patch size applied to 3D spatiotemporal microstructural degradation prediction. Outcome: worse than baseline. Reason: Lacks inductive spatial biases of CNNs for complex 3D volumetric sequences
TransVNet: Predicting bone degradation using ViT and virtual dataset of cellular microstructures · Iowa State
Tried and failed
adding intermediate fully-connected layers before classification head applied to vision transformer classification heads. Outcome: did not generalise. Reason: degraded generalization performance compared to direct classification
Transformer Networks for Smart Cities: Framework and Application to Makassar Smart Garden Alleys · Virginia Tech
Tried and failed
increasing input resolution during training applied to vision transformer self-supervised learning. Outcome: worse than baseline. Reason: reduced patch receptive fields degraded global token representation alignment
Exploiting Representation Similarities in Self-Supervised Learning for Vision Tasks · EPFL
Tried and failed
training all hidden layers in vision transformer applied to 3D microstructure sequence prediction. Outcome: worse than baseline. Reason: Full fine-tuning of transformer hidden layers failed to achieve satisfactory segmentation and distance metrics
TransVNet: Predicting bone degradation using ViT and virtual dataset of cellular microstructures · Iowa State
Tried and failed
Vision Transformer for spatial property estimation applied to short-range nonstationary spatial data. Outcome: worse than baseline. Reason: Attention mechanisms struggle to capture very localized spatial correlations compared to standard convolutional architectures under nonstationarity
Deep learning for spatial nonstationarity : evaluation, mitigation, and generation · UT Austin
Tried and failed
dilated sparse attention without dense local attention applied to vision transformers for image classification. Outcome: worse than baseline. Reason: omitting local neighborhood context degrades representations when relying purely on dilated sparse receptive fields
Neighborhood Attention: Fast and Flexible Sparse Attention · Georgia Tech
Tried and failed
self-supervised vision transformer for novel-class verification applied to sub-terahertz imaging object detection. Outcome: did not generalise. Reason: severe domain gap between natural image pre-training and sub-terahertz imagery
Advanced machine learning for object detection in sub-terahertz images · Imperial
Tried and failed
uniform input resolution downsampling applied to vision transformer compute reduction. Outcome: worse than baseline. Reason: reduced spatial detail significantly degraded accuracy compared to dynamic token and patch pruning at matched compute
Scalable transfer learning : enhancing adaptation and efficiency in uni- and multi-modal applications · UT Austin
Lost to a baseline
Swin-T vision transformer had lower brain prediction accuracy across visual cortex than simpler convolutional networks (CNN8, AlexNet, SqueezeNet1.1, ResNet18)
Characterizing human vision through large-scale brain imaging and computational models · MIT
Considered and rejected
Considered and rejected: Vision transformers (ViTs) rejected due to large data requirements and small target dataset size
Deep Learning and Synthetic Imagery for Migratory Bird Species Identification Using Drones · Carleton University Institutional Repository
Considered and rejected
Considered and rejected: Applying test-time OOD detection algorithms (other than simple MSP) to Vision Transformers (ViTs) due to poor performance compared to ConvNets
A Study of Techniques for Robustness to Out-of-Distribution Examples · DalSpace
Considered and rejected
Considered and rejected: Rejected using transformers for medical image analysis within the thesis due to their high data demands and lack of structural inductive bias compared to CNNs.
Considered and rejected
Considered and rejected: Vision Transformer (ViT) model was rejected/not tested on the Project 1640 dataset due to excessive training time required for large 250x250 image sizes.
Using AI to Enable Autonomous Exoplanet Direct Imaging · MIT
Considered and rejected
Considered and rejected: Rejected end-to-end recurrent/deep video architectures (e.g., LSTM, video transformers) for interview classification due to small subject sample sizes and severe overfitting risks
Multimodal assessment of neuropsychiatric disorders using audiovisual recordings · Georgia Tech
Lost to a baseline
Handcrafted/supervised OmniSource Slow-8x8-R101x2 beat thesis's linear transformer-pooling feature representation on HMDB-51 (83.8% vs 81.3%) and UCF-101 (98.6% vs 98.0%)
Learning and interpreting deep representations from multi-modal data · Oxford
Custom attention head steering, gating, and redistribution heuristics degrade representation quality and task performance
Attempts to steer attention heads, invert sink scores, apply multiplicative gates, or selectively redistribute attention weights degraded performance compared to standard softmax attention or uniform baselines. These heuristic interventions frequently removed essential contextual information, caused token distraction, or failed to outperform simple linear or weighted combinations.
Tried and failed
softmax self-attention mechanism applied to linear regression tasks. Outcome: worse than baseline. Reason: nonlinearity introduced representational inefficiencies requiring twice as many attention heads to match linear baseline
Tried and failed
steering all attention heads simultaneously applied to large language model self-attention mechanism. Outcome: worse than baseline. Reason: over-amplified targeted features at the expense of other essential contextual information
On the Efficiency and Steerability of Self-Attention Mechanism of Large Language Models · Georgia Tech
Tried and failed
threshold-clamped inverted normalisation softmax attention applied to vision transformer self-attention. Outcome: worse than baseline. Reason: encourages attention distraction across tokens during self-attention computation
Efficient machine learning software stack from algorithms to compilation · UT Austin
Tried and failed
standard self-attention without smoothness regularization applied to aspect extraction. Outcome: worse than baseline. Reason: attention weights over-concentrated on single words rather than broader context
Novel Algorithms for Understanding Online Reviews · Virginia Tech
Tried and failed
multiplicative gating on self-attention applied to transformer language models. Outcome: worse than baseline. Reason: did not improve downstream GLUE performance over standard self-attention
Designing for Inference in Future Generative Models · Cornell
Tried and failed
learned attention weighting for extractive rationale selection applied to text classification tasks. Outcome: worse than baseline. Reason: learned attention mechanisms did not improve classification performance over uniform attention baselines
Interpreting Neural Networks for and with Natural Language · Georgia Tech
Tried and failed
soft attention in reinforcement learning applied to visual robotic manipulation policy learning. Outcome: worse than baseline. Reason: implicit attention learning without explicit losses failed to focus feature representations effectively
Tightly-coupled manipulation pipelines: Combining traditional pipelines and end-to-end learning · Imperial
Tried and failed
selective attention redistribution to specific prompt components applied to transformer language model inference. Outcome: worse than baseline. Reason: biasing recovered attention exclusively toward questions or choices degraded task performance relative to uniform redistribution
Enhancing Foundation Models with Self-Guided Techniques: From Attention to Adapters to Agents · Georgia Tech
Tried and failed
Allocating asymmetric attention under constant information capacity applied to Sequential binary decision making. Outcome: worse than baseline. Reason: Performance peaked under balanced attention allocation rather than extreme focused attention
Optimal policy for attention-modulated decisions explains human fixation behavior · Harvard
Tried and failed
Attention mechanism over prototype feature representations applied to Metric-based few-shot classification. Outcome: worse than baseline. Reason: Learned attention weights did not outperform simple weighted side information combinations
Few-Shot and Zero-Shot Learning for Information Extraction · Virginia Tech
Tried and failed
increasing attention sink scores via inverse calibration applied to attention redistribution in foundation models. Outcome: worse than baseline. Reason: amplifying attention to non-informative sink tokens exacerbates distraction from semantically relevant tokens
Enhancing Foundation Models with Self-Guided Techniques: From Attention to Adapters to Agents · Georgia Tech
Tried and failed
zeroing scaling coefficient during attention steering applied to attention heads in language models. Outcome: worse than baseline. Reason: completely removed other crucial contexts at the steered attention heads
On the Efficiency and Steerability of Self-Attention Mechanism of Large Language Models · Georgia Tech
Tried and failed
linear attention approximation applied to automatic speech recognition. Outcome: worse than baseline. Reason: linear attention degradation compared to softmax attention
Tried and failed
self-attention mechanism for pairwise interaction modeling applied to permutation-invariant neural network surrogate. Outcome: worse than baseline. Reason: exhibited slow convergence and poorer MSE and gradient accuracy compared to weight-shared MLPs
Considered and rejected
Considered and rejected: Rejected using standard non-masked self-attention exclusively from the start of training in EoMT because it underperformed relative to attention-guided models (53.2 vs 56.0 PQ)
Efficiency Matters: Modern Techniques for Efficient Computer Vision · IRIS - POLITO - prod
Considered and rejected
Considered and rejected: Rejected reformulating the extractive loss over slots in slot attention because attention-weighted slot loss performed significantly better
Interpretable Representation Learning and Evaluation for Abstractive Summarization · EPFL
Considered and rejected
Considered and rejected: Single-head attention for in-context learning of 3-gram statistics because it fails to outperform bigram baselines.
Combinatorial Tasks as Model Systems of Deep Learning · Harvard
Excessive transformer parameter capacity leads to severe overfitting on small or specialized domain datasets
When applied to specialized tasks such as chemical property prediction, mass spectrometry, or small tabular datasets, large transformer models experienced rising validation loss and poor generalization. Simple tree ensembles, shallow neural networks, and smaller transformer variants consistently outperformed heavily parameterized models in these data-constrained settings.
Tried and failed
large transformer architecture for time series forecasting applied to power distribution load forecasting. Outcome: overfit. Reason: excessive parameter count led to severe overfitting on the training data
Advances to forecasting and planning frameworks toward resilient power distribution systems · Iowa State
Tried and failed
increasing hidden dimension in transformer model applied to time series temperature forecasting. Outcome: overfit. Reason: excessive model capacity caused overfitting after prolonged training compared to smaller hidden dimensions
Machine learning on building energy modeling : city and building scale · UT Austin
Lost to a baseline
Transformer language model reaction embeddings failed to outperform Random Forest, Extreme Gradient Boosting, or Feedforward Neural Networks using standard fingerprint or DFT features.
Machine Learning for Chemical Reactivity Prediction: Paradigms, Challenges, and Applications · MIT
Tried and failed
noisy text augmentation for cross-validation applied to transformer language models on small datasets. Outcome: overfit. Reason: augmentations lacked sufficient diversity and heightened model sensitivity to noise, worsening overfitting on small splits
Tried and failed
fine-tuning pretrained transformer for continuous regression applied to short text evaluation scores. Outcome: overfit. Reason: validation loss fluctuated in later epochs, failing to generalize to unseen data
Lost to a baseline
A 45M parameter large Transformer achieved lower mean generalization accuracy (0.20 ± 0.26) than smaller 9.5M (0.35) and 5.3M (0.37) parameter Transformers.
Compositional Linguistic Generalization in Artificial Neural Networks · JScholarship
Considered and rejected
Considered and rejected: Rejected Transformer and hybrid spectrogram Demucs variants due to overfitting and lack of convergence on small sample sizes.
Extraction of Blood Volume Pulse Morphology from Facial Videos Using an LSTM-Based Temporal Encoder-Decoder Model · Virginia Tech
Considered and rejected
Considered and rejected: Rejected MLP classifier with cross-attention for SketchQL's Matcher because it easily overfit training data compared to a cosine similarity metric over learned transformer embeddings
Enabling semantically richer queries over unstructured data · Georgia Tech
Tried and failed
sequence-to-sequence transformer on filtered benchmark dataset applied to chemical reaction product prediction. Outcome: overfit. Reason: uncleaned artifactual training examples producing trivial byproducts caused catastrophic overfitting and unchemical outputs
Tried and failed
overly strict early stopping threshold applied to transformer sequence-to-sequence fine-tuning. Outcome: overfit. Reason: small improvement threshold prevented stopping, leading to excessive training epochs and reduced generalization
Ensemble summarization models to leverage performance on CoronaNet · Texas Tech
Tried and failed
transformer neural network on fingerprint representations applied to mass spectrum intensity prediction. Outcome: overfit. Reason: the model showed rising validation loss during training and failed to generalize to unseen data
Learning to Fragment Molecular Graphs for Mass Spectrometry Prediction · Harvard
Tried and failed
Lowering classification head dropout during fine-tuning applied to audio spectrogram transformer classification. Outcome: overfit. Reason: Insufficient regularization led to higher training accuracy at the expense of test generalization
AI-generated synthetic audio analysis and forensic attribution · Iowa State
Considered and rejected
Considered and rejected: Rejected deep/large transformer neural architectures for tabular similarity feature classification in favor of XGBoost and a shallow 5-layer Sigmoid ANN to prevent overfitting.
Knowledge Graph Extension by Entity Type Recognition · IRIS - UNITN - prod
Considered and rejected
Considered and rejected: Rejected Transformer and GNN context encoders in ModuMorph due to overfitting and training instability, choosing simple MLPs instead.
Left open by the authors
Problems the authors named and did not get to.
Left open
Benchmark and compare the dual temporal attention recurrent model against time-series Transformers for many-to-many sequence prediction. Blocker: None
Recurrence and Temporal Attention Synergy for Optimal Time-Series Modeling and Interpretability · TXST Digital Repository
Left open
Evaluate pretrained transformer models with subword tokenization and transfer learning on scraped job postings to predict bankruptcy and corporate growth. Blocker: Scraped job listings dataset may not be publicly archived with the thesis
Assessing Corporate Growth and Bankruptcy Risk Using Public Data Proxies · Harvard
Left open
Compare Transformer architectures against LSTM networks for financial time-series bubble/jump detection and computer vision document-reading extraction. Blocker: None
Left open
Mitigate error accumulation in multi-step indoor temperature predictions over long context horizons using the Physical Transformer architecture. Blocker: None
Machine learning on building energy modeling : city and building scale · UT Austin
Left open
Format ABI and GMI satellite data into time-series sequences and train transformer models to predict temporally linked outputs. Blocker: None
BAYESIAN PATCH TO PATCH PREDICTION FOR SATELLITE IMAGERY · Calhoun
Left open
Train and evaluate transformer-based architectures instead of LSTMs to improve growing degree unit (GDU) prediction accuracy. Blocker: None
Improving crop productivity through data-driven optimization and hybrid deep learning-based approaches · Iowa State
Left open
Pre-train Transformer-based energy foundation models on web-scale, high-dimensional energy time-series datasets. Blocker: No specific web-scale datasets or architectures are defined beyond a general research direction
Toward Transformer-based Large Energy Models for Smart Energy Management · Virginia Tech
Left open
Implement generative transformer architectures for time-series biosignal classification. Blocker: None
A Skin-Like Sternal Patch to Monitor Autonomic Tone During Cognitive Stress and Sympathetic Arousals in Disordered Sleep · Georgia Tech
Left open
Characterize transformer training dynamics on correlated token sequences for next-token prediction using gradient flow analysis. Blocker: None
Towards Newtonian understanding of deep learning training dynamics · UT Austin
Left open
Establish formal theoretical relationships connecting transformer architectures and diffusion models operating on graphs and manifolds. Blocker: Lacks specific mathematical formulations, conjectures, or defined targets to bridge transformers and diffusion models
MANIFOLD FILTERS AND NEURAL NETWORKS: GEOMETRIC GRAPH SIGNAL PROCESSING IN THE LIMIT · Penn
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.