Chapter Four · failure evidence
What Large Language Models & Prompt Engineering got wrong, from 50 dissertations
Across diverse tasks, large language models and prompt engineering techniques frequently underperformed simpler supervised classifiers, struggled with domain-specific extraction, and failed on complex reasoning or code generation. Iterative prompt modifications often yielded negligible gains over baseline prompts, while fine-tuning and advanced prompting structures encountered high implementation overhead, capacity bottlenecks, and behavioral degradation. These records come from PhD theses at 17 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Zero-shot prompting fails on specialized domain tasks and complex extraction
Zero-shot prompts achieved near-random accuracy or low recall when applied to domain-specific extraction, legal statutory interpretation, and nuanced subjective annotation. Models lacked the necessary domain context, structured reasoning, or fine-tuning to outperform simple baselines like rule-based sentiment analyzers.
Tried and failed
direct zero-shot prompting of large language models applied to clinical note information extraction. Outcome: did not generalise. Reason: failed to adapt and performed poorly on advanced reasoning models without structured reasoning or fine-tuning
Machine Learning Approaches for Drug Combination Discovery · Cornell
Tried and failed
prompt-based zero-shot large language model classification applied to student text behavior classification. Outcome: worse than baseline. Reason: produced substantially lower AUC and more than double the false positives compared to embedding-based classifiers
Tried and failed
zero-shot large language model prompting applied to commonsense logic rule generation. Outcome: did not generalise. Reason: models fail to consistently output valid compositional rules without task-specific fine-tuning
Neurosymbolic Programming in Scallop: Design, Implementation, and Applications · Penn
Tried and failed
zero-shot prompting with basic statutory definitions applied to confidential information redaction in text. Outcome: did not generalise. Reason: Abstract legal definitions lacked sufficient context to identify complex, non-standard sensitive categories accurately.
LLM-Assisted Detecting and Redacting Confidential Information for Government Information Disclosure · Virginia Tech
Tried and failed
large language model zero-shot text classification applied to moral framing text annotation. Outcome: no signal. Reason: LLM predictions failed to align with validated domain-specific dictionary benchmarks across categories
COMPUTATIONAL METHODS ON POLITICAL MORAL LANGUAGE IN ONLINE SPACES · Cornell
Tried and failed
zero-shot large language model generation without exemplars applied to generating software vulnerability exploits. Outcome: no signal. Reason: the model lacked domain-specific context and test patterns needed to synthesize functional exploit payloads
Secure Coding Practice in Java: Automatic Detection, Repair, and Vulnerability Demonstration · Virginia Tech
Tried and failed
zero-shot large language model reasoning applied to temporal knowledge graph link prediction. Outcome: did not generalise. Reason: models fail to understand structured graph entities and temporal relations without specialized fine-tuning
Temporal link prediction in the wild · Imperial
Tried and failed
zero-shot prompting of general-purpose LLMs applied to extracting viral mutations from text. Outcome: no signal. Reason: general-purpose models achieved extremely low recall and F1 scores on domain-specific extraction
Discovering Viral Hosts, Mutations, and Diseases using Machine Learning · Virginia Tech
Tried and failed
zero-shot large language model prompting applied to nuanced text annotation tasks. Outcome: no signal. Reason: models achieved near-random performance on subjective constructs
More Than a Sum of Its Parts: Teamwork Through a Multidimensional Lens · Penn
Tried and failed
direct LLM zero-shot prompting without retrieval applied to biological perturbation effect prediction. Outcome: no signal. Reason: models lacked external experimental context and structured reasoning to predict perturbation outcomes above chance
Practical Algorithms for Modeling Causality to Accelerate Scientific Discovery · MIT
Lost to a baseline
Zero-shot prompting on several LLMs (including Llama3 8B and Mistral 7B) achieved lower accuracy than the rule-based VADER sentiment analysis baseline
Prompted generative models lose to supervised baselines and specialized encoders
Few-shot and in-context prompts underperformed specialized encoder architectures like BERT and simple supervised classifiers on domain text classification and sequence extraction. In several instances, prompted generative models failed to beat naive majority-class baselines or simpler language modeling alternatives.
Tried and failed
few-shot prompting of pretrained language models applied to domain-specific text classification. Outcome: worse than baseline. Reason: general pretraining lacks alignment with specialized, nuanced coding rubrics compared to supervised simple classifiers
Evaluating language models applied to student thinking about experiments · Cornell
Tried and failed
prompting pretrained generative large language models applied to complex domain-specific text classification. Outcome: worse than baseline. Reason: zero-shot and few-shot prompting failed to capture complex domain constructs compared to fine-tuned domain-specific encoders
Modeling Legal Constructs · Cornell
Lost to a baseline
Most LLM prompts for stance extraction across all 1,000 tweets had lower overall accuracy than a naive majority-class baseline
Lost to a baseline
FIST with only prompt input yielded higher perplexity (30.2 word PPL / 25.9 BPE) and lower ROUGE-1 F1 (0.181) than PSA (31.6 word / 21.3 BPE / 0.265 F1) on WritingPrompts.
Towards Effective and Controllable Neural Text Generation · DSpace at SUNY Buffalo
Considered and rejected
Considered and rejected: Rejected relying solely on in-context prompting (zero-shot, few-shot, and Chain-of-Thought) for specialized legal construct classification due to poor performance and severe class imbalance errors.
Modeling Legal Constructs · Cornell
Tried and failed
fine-tuned large language models for sequence extraction applied to named entity recognition. Outcome: worse than baseline. Reason: Generative LLMs underperformed specialized smaller encoder models like BERT on token-level span classification tasks.
Consistency-aware and LLM-assisted methods for named entity recognition · Iowa State
Considered and rejected
Considered and rejected: Rejected Large Language Models (LLMs) in favor of LDA topic models and dictionary-based content analysis.
The Rise of Counterpopulism: a Comparative Analysis of Social Movements Resisting Right-Wing Populism in Power in Italy, the United Kingdom and the United States of America · IRIS - SNS - prod
Prompt engineering and instruction refinement yield negligible performance gains
Manual prompt engineering, perspective augmentations, and formatting prompts with human annotation guidelines failed to provide meaningful improvements over simple baseline prompts. Re-prompting failed to correct persistent model biases, and complex prompt instructions increased overfitting risks without reliably controlling output complexity.
Tried and failed
manual iterative prompt engineering applied to large language model text annotation. Outcome: no signal. Reason: manual refinements yielded only marginal performance gains over baseline prompts
Tried and failed
using human annotation guidelines as LLM prompts applied to text classification and rationale extraction. Outcome: worse than baseline. Reason: instructions optimized for human annotators did not translate effectively to LLM reasoning and yielded no improvement
Towards Bridging and Governing Decentralized Communities · MIT
Tried and failed
First-person perspective text augmentation applied to Clinical text classification. Outcome: no signal. Reason: Perspective shift provided no meaningful performance gain over baseline prompts
Transforming SDOH Screening: Towards a General Framework for Transformer-based Prediction of Social Determinants of Health · Virginia Tech
Tried and failed
prompt engineering on large language models applied to zero-shot tone classification of transcripts. Reason: re-prompting failed to correct the model's conservative classification bias, causing poor agreement with human annotators
Considered and rejected
Considered and rejected: Rejected complex prompt instructions for closed-book QA, sticking to 'Q: <question> A:' because sophisticated prompts did not improve accuracy enough to justify model overfitting risks.
Beyond Scaling: Frontiers of Retrieval-Augmented LMs · ResearchWorks
Considered and rejected
Considered and rejected: Rejected relying solely on grade-level prompt engineering to vary summary complexity because lower grade prompts often generated complex language.
Language as Design: Adapting Language to Different Online Audiences · ResearchWorks
Considered and rejected
Considered and rejected: Rejected using zero-shot LLM prompts with detailed definitions and implicit instructions, choosing concise prompts without added explanations.
Machine Learning Approaches to Understanding Legislative Economic Discourse · Research Repository UCD
Code repair and generation struggle with complex logic and architectural context
Prompting models for code conflict resolution and algorithmic problem solving resulted in meaningless outputs, rule overgeneralization, and failure on multi-step state transitions. Practitioners rejected language models for critical code development and multi-file refactoring because debugging machine-generated logic errors was inefficient and models lacked broader architectural context.
Tried and failed
large language models for code repair applied to software build conflict resolution. Reason: generated meaningless outputs and achieved low resolution accuracy on complex merge conflict tasks
Automating The Detection and Resolution of Build Conflicts in Software Merge for Java Programs · Virginia Tech
Tried and failed
prompting LLMs with explicit domain-specific rules applied to automated code conflict resolution. Outcome: worse than baseline. Reason: models overgeneralized explicit rules, incorrectly misapplying patterns to unrelated code contexts
Automating The Detection and Resolution of Build Conflicts in Software Merge for Java Programs · Virginia Tech
Tried and failed
prompting large language models for algorithmic generation applied to complex dynamic programming problems. Outcome: did not generalise. Reason: models fail on complex algorithmic reasoning requiring multi-step state transitions and optimal substructure
Are Bigger LLMs Always Better? A Study of Open and Closed-Source Models in Code Generation and Translation · Virginia Tech
Considered and rejected
Considered and rejected: Rejected using Large Language Models (ChatGPT) for writing critical code because debugging machine-generated logic errors takes more time than writing manually and models cannot answer contextual 'why' questions.
Code, communication, and culture : an ethnography of open source quantum software communities · UT Austin
Considered and rejected
Considered and rejected: Rejected open-ended optimization prompts (e.g., 'make my code faster') and multi-file refactoring with ChatGPT due to lack of extensive non-code and architectural context
Studying Developers’ Engagement with ChatGPT for Issue Resolution · HARVEST
Supervised fine-tuning causes behavioral degradation and imposes high training overhead
Supervised fine-tuning collapsed citation grounding, degraded output format compliance, and lacked the self-revision mechanisms available in prompting. Researchers also rejected fine-tuning for prototyping and specialized detection tasks due to extensive annotation requirements, behavioral instability, and prohibitive computational costs.
Tried and failed
supervised fine-tuning for structured reasoning applied to large language models for clinical grounding. Outcome: worse than baseline. Reason: optimizing for diagnostic correctness collapsed citation grounding and degraded structured output format compliance
Causal Inference and Evidence-Grounded Language Models for Trustworthy Personalized Clinical Decision Support · Georgia Tech
Tried and failed
uninformed magnitude pruning with weight regeneration applied to large language models. Outcome: worse than baseline. Reason: causes irreparable knowledge damage on complex tasks that weight regeneration fine-tuning cannot recover
Chasing efficiency in the era of large language models · UT Austin
Considered and rejected
Considered and rejected: Fine-tuning large language models rejected for prototyping due to extensive annotation/training requirements and difficulty maintaining consistent behavioral parameters.
Balancing Performance and Social Considerations for Autonomous Agents Interacting with Humans · Cornell
Considered and rejected
Considered and rejected: Rejected fine-tuning large language models (GPT-4/Gemini) for affective state detection due to prohibitive computational costs and limited annotated data, opting for instruction-based zero-shot prompting.
Advancing Ransomware Detection with Machine Learning: Human Factors, Evolution, and Detection · Research Repository UCD
Tried and failed
Supervised fine-tuning on positive reasoning paths only applied to Mathematical question answering. Outcome: worse than baseline. Reason: Lacked error-correction and self-revision mechanisms present in prompting
Knowledge-Centric Multimodal Intelligence: Understanding, Reasoning, and Generation · Virginia Tech
Mathematical and multi-step reasoning degrades under direct or uninformative prompting
Models struggled with explicit symbolic equations and multi-step mathematical problems, leading to degraded accuracy and higher error rates. Direct prompting for intermediate answers reduced answer consistency compared to resampling techniques, while empty noise exclusion hints worsened downstream reasoning performance.
Tried and failed
prompting language models with explicit symbolic equations applied to synthetic tabular data generation. Outcome: worse than baseline. Reason: LLMs struggle with complex mathematical reasoning, degrading generation accuracy and increasing MSE
New Approaches to Synthetic Tabular Data Generation · Virginia Tech
Lost to a baseline
UnifiedQA zero-shot answering on StrategyQA scored 58.95% accuracy, heavily underperforming GPT-3 zero-shot prompting (58.08% / 63.32% few-shot) and other reasoning baselines when not given gold supporting facts.
Incidental Supervision for Natural Language Understanding · Penn
Tried and failed
direct prompting for intermediate answers applied to multi-step mathematical reasoning. Outcome: worse than baseline. Reason: yields lower answer consistency across reasoning benchmarks than forward-reverse resampling
In the Blink of an Eye: A Unified Theory for Feature Emergence in Generative Models · Harvard
Tried and failed
prompting with an empty noise exclusion hint applied to mathematical reasoning under noisy input. Outcome: worse than baseline. Reason: providing an empty exclusion hint degraded downstream problem-solving performance compared to explicit noise identification
Quantitative reasoning ability of large language models under noisy data · Iowa State
Complex prompting structures overwhelm the capacity of smaller language models
Applying complex reasoning prompts such as tree structures or chain-of-thought to lower-capacity models led to poor performance. The representational overhead caused model confusion, high labeling variability, and frequent classification errors.
Tried and failed
tree-structured prompting applied to small language model reasoning. Outcome: worse than baseline. Reason: representational overhead overwhelmed the capacity of smaller models
Backward Reasoning in LLMs: A Strategy for Identifying Irrelevant Context in Mathematical Word Problems · Virginia Tech
Tried and failed
chain-of-thought prompting with lower-capacity LLM applied to complex multi-class text annotation. Outcome: unstable. Reason: Model exhibited confusion, high labeling variability, and frequent classification errors compared to larger models
Value Expressions in Patents: Their Relationship to Patent Valuation and Technological Orientation · Georgia Tech
Automated evaluation prompts generate generic feedback rather than pinpointing errors
When evaluating student solutions or writing, models produced vague, generic, and repetitive critiques rather than identifying specific error locations. These automated assessments failed to provide the precise diagnostic feedback offered by human instructors.
Tried and failed
large language models for automated grading applied to student engineering derivations and solutions. Reason: Outputted full solutions and generic suggestions rather than pinpointing specific error locations in student work
Developing an Automated Practice Environment and Feedback Engine for Guided Instruction in the Deformable Bodies Course · Virginia Tech
Tried and failed
prompting large language models for critique generation applied to evaluating complex student writing. Outcome: worse than baseline. Reason: generated critiques were vague, generic, and repetitive compared to human instructor assessments
Addressing teachers’ needs for integrating generative AI into a pedagogical feedback model · Iowa State
Left open by the authors
Problems the authors named and did not get to.
Left open
Evaluate zero-shot, one-shot, and few-shot prompting on LLM annotation performance for pedagogical discourse moves in classroom transcripts. Blocker: Access to the annotated elementary classroom transcript dataset.
Left open
Apply the PARSE framework to evaluate LLM perturbation bias on prompts in healthcare, legal guidance, and job matching domains. Blocker: None
PARSE: A Framework for Evaluating Linguistic and Typographical Perturbation Bias in Large Language Models · YorkSpace
Left open
Benchmark general-purpose large language models against specialized predictive models on reaction condition tasks such as solvent prediction. Blocker: None
Left open
Compare empirical neural scaling statistics in large language models against predictions from prior analytical scaling models. Blocker: None
Left open
Implement and evaluate early-fusion n-gram LM prediction prompting or conditioning for neural LLM text generation. Blocker: None
Navigating the Ocean of Language Model Training Data · ResearchWorks
Left open
Adapt summarization evaluation prompting methods to handle documents exceeding context limits and structured tabular inputs. Blocker: None
Fine-grained evaluation for text summarization · UT Austin
Left open
Apply advanced large language models and hybrid quantum-classical algorithms to source code representations for software defect prediction. Blocker: The unfinished work lacks concrete specifications, specific model architectures, or target benchmarks
Left open
Apply lifelong learning fine-tuning techniques to large language models for generating automated code review comments across evolving codebases. Blocker: None
Towards Sustainable AI for Continuous Integration Quality Gates · Queens University Institutional Repository
Left open
Apply large language models to resolve grounding and policy errors in compositional decision-making tasks. Blocker: None
Left open
Finetune large language models directly on large annotated hardware design codebases instead of using demonstration libraries. Blocker: None
Harnessing Large Language Models Towards More Accessible Hardware Accelerator Design · Georgia Tech
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.