Chapter Four · failure evidence
What Zero-Shot Learning got wrong, from 93 dissertations
Across these thesis studies, zero-shot learning methods frequently underperform supervised alternatives, fail under distribution shifts, and struggle with specialized reasoning tasks. They also suffer from severe prediction biases and fail to capture subjective human qualities or fine-grained physiological and acoustic patterns. These records come from PhD theses at 21 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Zero-shot methods lag behind supervised models and simple baselines across benchmarks
Zero-shot foundation models and prompting configurations consistently achieve lower accuracy, recall, and precision than fully trained or fine-tuned baselines across vision and language tasks. In several evaluations, zero-shot models even underperformed basic rule-based systems, random baselines, or few-shot prompts.
Lost to a baseline
On Multi30K German-to-image search, thesis English-only zero-shot model (31.4% R@1) lost to supervised MHA-D (40.3% R@1)
Learning and interpreting deep representations from multi-modal data · Oxford
Lost to a baseline
All zero-shot models achieved lower raw Recall@K across all cutoffs compared to fully trained models (IMP, GPSNet, Motifs, VCTree, OpenPSG)
Zero-Shot Scene Graph Relationship Prediction using VLMs · Virginia Tech
Lost to a baseline
UnifiedQA zero-shot answering on StrategyQA scored 58.95% accuracy, heavily underperforming GPT-3 zero-shot prompting (58.08% / 63.32% few-shot) and other reasoning baselines when not given gold supporting facts.
Incidental Supervision for Natural Language Understanding · Penn
Lost to a baseline
Vanilla zero-shot style transfer failed to return a valid response 25.4% of the time, compared to 0.6% for augmented zero-shot learning.
Understanding the Limitations of Using Large Language Models for Text Generation · Penn
Lost to a baseline
Zero-shot foundation model mPLUG-Owl achieved inferior or comparable precision/recall (0.20 precision / 0.20 recall on multiple groundings) compared to naive single/multiple baselines on the Single Answer Grounding Challenge.
Visual question answering : representing authentic use cases and ambiguity · UT Austin
Lost to a baseline
Zero-shot pretrained models (R-AUROC ~0.57-0.60) lost to many-shot ResNet-50 trained directly on downstream target classes (R-AUROC 0.68)
Uncertainties of Latent Representations in Computer Vision · Publikationssystem UB Tuebingen
Lost to a baseline
In-domain zero-shot German question formation for mBART failed to correctly transform the full sequence (0% sequence accuracy), underperforming mT5.
Emergent Syntactic Behaviors and Mechanisms in Neural Language Models · JScholarship
Lost to a baseline
Zero-shot CLIP OOD methods (MCM and Cosine) achieved 67.1% to 79.8% FPR on NINCO, failing to beat standard IN-1K fine-tuned classifiers
Out-of-Distribution Detection, Sharpness, and Unlearning: Advancing Robust and Trustworthy Deep Learning · Publikationssystem UB Tuebingen
Tried and failed
domain fine-tuned zero-shot language model applied to time series forecasting. Outcome: worse than baseline. Reason: exhibited performance degradation across all horizons and metrics compared to standard deep learning models
Deep Learning Methods for Built Environment Operational Management · Virginia Tech
Tried and failed
zero-shot prompting with abstract category definitions applied to program semantic reasoning. Outcome: worse than baseline. Reason: lacking concrete few-shot examples degraded performance relative to direct prediction
CodeSense: A real-world benchmark and dataset for code semantic reasoning · Iowa State
Lost to a baseline
Zero-shot and few-shot Llama 3 in-context learning achieved F1-scores of only 0.29–0.40, falling well behind the non-LLM BiLSTM baseline F1 of 0.67.
Surgical Site Infection (SSI) Identification Across Multiple Facilities and Surgery Types Using Multimodal Data and Deep Learning · ResearchWorks
Lost to a baseline
Zero-shot prompting on several LLMs (including Llama3 8B and Mistral 7B) achieved lower accuracy than the rule-based VADER sentiment analysis baseline
Lost to a baseline
On C-GQA open-world compositional zero-shot learning with zero-shot CLIP, GloVe baseline beat FLM on unseen class accuracy (3.92% vs 2.62%) and AUC (0.20% vs 0.16%).
Learning under Limited Supervision for Better Generalization · Publikationssystem UB Tuebingen
Lost to a baseline
On ActivityNet-QA VQA, supervised finetuned JustAsk (FT) achieved 64.667% vs this method's zero-shot 61.168%.
Learning Generalizable Systems by Learning Composable Energy Landscapes · MIT
Lost to a baseline
DiffuCOMET-Fact (RA-F1 43.51) lost to Beam-COMET baseline (RA-F1 45.88) on out-of-domain MovieSummaries zero-shot testing.
Lost to a baseline
Zero-shot Kaleido 3B achieved 71.5% on ETHICS Commonsense, losing to few-shot ChatGPT (80.3%).
Steps Towards the Pluralistic Alignment of Language Models · ResearchWorks
Tried and failed
pretrained vision-language and self-supervised models for keypoint detection applied to deformable object semantic keypoint localization. Outcome: worse than baseline. Reason: zero-shot VLM and DINOv2 backbones lacked fine spatial localization precision compared to supervised ResNets
Deformable Object Manipulation with a Tactile Reactive Gripper · MIT
Lost to a baseline
On few-shot HICO classification (Few@1, Few@5, Few@10), zero-shot HTS ViT-B/32 (33.6, 36.3, 37.2 mAP) was beaten by supervised HTS* ViT-B/32 (49.6, 52.6, 52.9 mAP).
Crossing the Chasm: Bridging Perception and Generative Models for Enhanced Vision-Language Understanding · ResearchWorks
Lost to a baseline
Zero-shot and decoder-fit DNNs achieved lower accuracy (barely above 8.33% chance at 35% fragments) compared to humans (~50% accuracy).
Differences in Visual Perception in Humans and Deep Neural Networks · EPFL
Zero-shot models struggle with specialized domain knowledge and structured reasoning
Pretrained models fail when deployed on highly technical tasks including medical grounding, chemical reaction prediction, code synthesis, and temporal knowledge reasoning. Without domain-specific supervision or intermediate feedback, they produce ungrounded predictions, generate invalid program code, or trigger safety guardrails.
Tried and failed
zero-shot vision-language contrastive classification applied to fine-grained multi-class pathology subtyping. Outcome: did not generalise. Reason: lack of supervision causes zero-shot models to fail on complex, fine-grained multi-class classification tasks
Data-Driven General Purpose Foundation Models for Computational Pathology · MIT
Tried and failed
zero-shot text-prompted segmentation models applied to medical image finding grounding. Outcome: no signal. Reason: pretrained models failed completely on complex free-text clinical finding descriptions
AI Systems for Understanding and Grounding Radiology Reports · Harvard
Tried and failed
direct zero-shot diagnostic prompting of multimodal models applied to medical image classification. Reason: unengineered direct diagnostic queries triggered commercial safety guardrails, causing high refusal rates
Augmenting medical image classifiers with synthetic data across populations · Harvard
Tried and failed
zero-shot transformer reaction outcome predictor applied to unseen chemical reaction classes. Outcome: did not generalise. Reason: Models fail to infer transformation mechanisms for novel reaction types absent from training data
Tried and failed
direct zero-shot prompting of large language models applied to clinical note information extraction. Outcome: did not generalise. Reason: failed to adapt and performed poorly on advanced reasoning models without structured reasoning or fine-tuning
Machine Learning Approaches for Drug Combination Discovery · Cornell
Tried and failed
zero-shot sequence-to-sequence pretrained model classification applied to claim verification verdict classification. Outcome: worse than baseline. Reason: lack of task-specific fine-tuning severely limited classification capability compared to fine-tuned alternatives
Explainable Neural Claim Verification Using Rationalization · Virginia Tech
Tried and failed
instruction-tuned vision-language model applied to zero-shot visual question answering. Outcome: worse than baseline. Reason: prior instruction tuning biased the model toward coarse-grained answers rather than fine-grained predictions
Extracting Knowledge with Multimodal and Multilingual Intelligent Systems · Georgia Tech
Tried and failed
zero-shot transformer reranker on beam search candidates applied to neural code generation candidates. Outcome: worse than baseline. Reason: compounded overfitting from using an un-finetuned reranker
Neuro-symbolic program generation and execution for hybrid reasoning · Iowa State
Considered and rejected
Considered and rejected: Rejected relying solely on in-context prompting (zero-shot, few-shot, and Chain-of-Thought) for specialized legal construct classification due to poor performance and severe class imbalance errors.
Modeling Legal Constructs · Cornell
Tried and failed
zero-shot vision-language models applied to robot trajectory planning and motion reasoning. Outcome: did not generalise. Reason: models exhibit severe forward-motion and deceleration biases and fail at temporal motion reasoning
Building Intelligence that can Interact with the Physical World · MIT
Tried and failed
zero-shot code generation using pretrained LLM applied to high-level synthesis code generation. Reason: pretrained model lacked specialized domain knowledge and constructs required for valid high-level synthesis code
Harnessing Large Language Models Towards More Accessible Hardware Accelerator Design · Georgia Tech
Tried and failed
zero-shot large language model generation without exemplars applied to generating software vulnerability exploits. Outcome: no signal. Reason: the model lacked domain-specific context and test patterns needed to synthesize functional exploit payloads
Secure Coding Practice in Java: Automatic Detection, Repair, and Vulnerability Demonstration · Virginia Tech
Tried and failed
LLM zero-shot reward function code generation applied to dexterous robotic hand manipulation tasks. Reason: Generated reward code failed to execute consistently across multiple random seeds.
Personalized, Safe, and Interactive Robot Programming via Human Demonstrations · Georgia Tech
Tried and failed
zero-shot direct prompting of large language models applied to resolving semantic ambiguities in text. Reason: Failed to resolve complex relational and temporal dependencies without intermediate feedback or attributes
Transforming Free-Form Sentences into Sequence of Unambiguous Sentences with Large Language Model · Virginia Tech
Lost to a baseline
Zero-shot BLINK achieved 54.4% accuracy on NLM-Chem, missing domain-specific performance
Entity Linking in Low-Annotation Data Settings · JScholarship
Tried and failed
zero-shot in-context learning with large language models applied to entity linking disambiguation. Outcome: worse than baseline. Reason: struggled with fine-grained candidate entity disambiguation without fine-tuning
Information extraction with weak supervision · Iowa State
Tried and failed
zero-shot large language model reasoning applied to temporal knowledge graph link prediction. Outcome: did not generalise. Reason: models fail to understand structured graph entities and temporal relations without specialized fine-tuning
Temporal link prediction in the wild · Imperial
Models fail to capture subjective, affective, and nuanced human evaluations
Zero-shot prompting produces near-random accuracy or negative correlation when evaluating subjective qualities, emotional states, and rhetorical attributes. Across persona modeling, stance detection, and qualitative thematic extraction, unadapted models fail to match human consensus ratings and benchmark standards.
Tried and failed
zero-shot autoregressive language model generation and classification applied to logical fallacy prediction. Outcome: did not generalise. Reason: models fail to identify and distinguish unseen fallacy categories in zero-shot settings
Extending Provenance for Understanding Claims and Data Analyses · Penn
Tried and failed
Zero-shot prompting of large language models applied to thematic analysis across study groups. Outcome: did not generalise. Reason: Prompts alone struggled to synthesize generalized themes and perform cross-group comparisons across qualitative data.
Embodied Virtual Reality: The Impacts of Human-Nature Connection During Engineering Design · Virginia Tech
Tried and failed
zero-shot multimodal foundation model classification applied to image emotion recognition. Outcome: worse than baseline. Reason: general pre-trained embeddings lack domain-specific fine-tuning on nuanced emotion distributions
Three essays on online word-of-mouth and user behavior in online environments · Iowa State
Tried and failed
zero-shot large language model classification applied to persona relation extraction. Outcome: worse than baseline. Reason: Individual model predictions were unreliable compared to aggregated human majority voting.
Tried and failed
zero-shot large language model generation applied to automated customer review responses. Outcome: worse than baseline. Reason: generated text significantly impaired customer perceived helpfulness across varied sampling temperatures
Strategizing in Response to Environmental Uncertainty in the Hospitality Industry: A Data-Analytical Approach · Virginia Tech
Lost to a baseline
In conversational Turing tests, zero-shot ChatGPT performed poorly as an AI judge and was outperformed by one-shot prompted ChatGPT and trained SVM classifiers.
Lost to a baseline
Zero-shot persona gpt-4o-mini performed worse than the theoretical random baseline (Qini -0.167 vs 0.000).
Language Models as Opinion Models: Techniques and Applications · MIT
Lost to a baseline
Zero-shot text2text gpt-4o-mini produced worse-than-random HTE prediction (Qini -0.052 vs baseline 0.183).
Language Models as Opinion Models: Techniques and Applications · MIT
Considered and rejected
Considered and rejected: Direct zero-shot classification with off-the-shelf VLMs without fine-tuning, rejected due to poor performance on affective tasks
Three essays on online word-of-mouth and user behavior in online environments · Iowa State
Considered and rejected
Considered and rejected: Rejected directly using LLM-generated zero-shot stance predictions in favor of extracting linguistic/subsequence rationales
Advancing stance detection and fine-grained content analysis for socially relevant domains · Leibniz Universität Hannover Repository
Tried and failed
zero-shot prompt classification of isolated short phrases applied to narrative frame identification. Outcome: no signal. Reason: isolated phrases lacked sufficient context, making direct frame elicitation too vague and unreliable
Tried and failed
zero-shot LLM numerical attribute scoring applied to rhetorical attribute extraction from text. Outcome: no signal. Reason: LLM scores for nuanced subjective attributes failed to correlate with human benchmark ratings
Preaching to the Choir: An AI-Based Analysis of Religious Demand in U.S. Church Sermons, 2000-2023 · Harvard
Tried and failed
zero-shot vision-language classification with abstract keywords applied to visual art genre classification. Reason: abstract aesthetic keywords caused semantic leakage and false positive over-assignment across unrelated categories
Tried and failed
zero-shot large language model prompting applied to nuanced text annotation tasks. Outcome: no signal. Reason: models achieved near-random performance on subjective constructs
More Than a Sum of Its Parts: Teamwork Through a Multidimensional Lens · Penn
Lost to a baseline
Zero-shot text2text Pythia-410m (untrained) produced an individual-dataset Qini (-0.063) that was worse than the ensemble baseline (0.183).
Language Models as Opinion Models: Techniques and Applications · MIT
Lost to a baseline
Unsupervised SimCSEbase and supervised FlanT5large achieved negative Spearman correlation (-3.0 and -0.43) on zero-shot Conditional Semantic Text Similarity.
Methods and Challenges In Inference Across Textual Sources · Penn
Lost to a baseline
Zero-shot GPT-4 achieved only 44% accuracy matching human ratings across rating constructs.
Procedural Justice for CS Hiring: Towards Fair and Equitable Resume Matching in AI Era · Virginia Tech
Severe domain shifts degrade zero-shot transfer across visual and physical environments
Vision and spatial models trained on standard distributions fail when evaluated on histopathology, satellite, aerial, aquatic, or simulated robotic environments. The lack of task-specific adaptation causes severe false negatives, unstable segmentations, and semantic confusion across disparate domains.
Tried and failed
zero-shot segmentation using pre-trained network applied to out-of-distribution histopathology images. Outcome: did not generalise. Reason: produced unstable nuclear segmentations and unreliable classification without fine-tuning and threshold calibration
Integrated Molecular and Histopathological Studies of Testicular Germ Cell Tumor Biology · Harvard
Tried and failed
zero-shot vision-language model classification applied to satellite imagery object detection. Outcome: no signal. Reason: poor zero-shot alignment on domain-specific satellite imagery led to zero true positives
Adaptive Learning for AI Systems: AI Methods for Scientific Domains under Limited Supervision · Cornell
Tried and failed
deep learning PDE solver direct zero-shot inference applied to out-of-distribution physical scattering fields. Outcome: did not generalise. Reason: test inputs had structures completely uncorrelated with the synthetic training distribution
Exploring Novel Modalities for Optical Diffraction Tomography · EPFL
Tried and failed
zero-shot transfer of task-finetuned conditional adapters applied to unseen visual generation tasks. Outcome: did not generalise. Reason: specialized finetuning destroyed zero-shot generalization capability on out-of-distribution conditioning tasks
Enhancing generative model efficiency and control for data generation, reinforcement learning, and robotics · UT Austin
Tried and failed
unprompted zero-shot foundation model segmentation applied to complex morphology cell microscopy images. Outcome: did not generalise. Reason: zero-shot segmentation failed on complex star-shaped developmental morphologies without task-specific prompting or fine-tuning
IDCC-SAM: Automated cell counting in immunocytochemistry using the segment anything model · Iowa State
Tried and failed
zero-shot pretrained cell segmentation model applied to unseen tissue histopathology classification. Outcome: did not generalise. Reason: the specific target cancer type was underrepresented in the pretraining dataset
Spatially aware deep learning for clear cell renal cell carcinoma characterization and discovery · Harvard
Tried and failed
zero-shot object detection using off-the-shelf models applied to floating aquatic debris identification. Outcome: did not generalise. Reason: Domain shift between standard training imagery and aquatic surface imagery caused near-zero detection accuracy.
Autonomous System for Identifying and Capturing Floating Waste · Georgia Tech
Tried and failed
zero-shot transfer of real-trained segmentation models applied to synthetic simulation environments. Outcome: did not generalise. Reason: Domain shift between real-world training imagery and synthetic rendering distributions.
Perception Enabled Planning for Autonomous Systems · Cornell
Tried and failed
zero-shot pre-trained instance segmentation applied to indoor robotic object detection. Outcome: did not generalise. Reason: distribution shift between pre-training dataset and deployment environment caused significant false negatives on visible objects
Autonomous exploration and object reconstruction with an MAV · Imperial
Tried and failed
zero-shot transfer of self-supervised image clustering applied to satellite imagery across geographic regions. Outcome: did not generalise. Reason: domain shift between geographic locations caused semantic confusion across visually distinct feature clusters
Tried and failed
direct zero-shot cross-dataset transfer of object detectors applied to multi-object tracking benchmark evaluation. Outcome: did not generalise. Reason: domain shift between source and target datasets degraded detection accuracy without target fine-tuning
Improved 2D Camera-Based Multi-Object Tracking for Autonomous Vehicles · Virginia Tech
Tried and failed
zero-shot image-to-image translation GAN applied to cross-sensor aerial imagery. Outcome: did not generalise. Reason: sensor and domain distribution shifts between satellite training data and aerial test imagery
Deep Learning Emulators for Accessible Climate Projections · MIT
Tried and failed
Pre-trained video moment retrieval model applied to cross-domain video moment queries. Outcome: did not generalise. Reason: Severe zero-shot domain mismatch between pre-training distribution and evaluation moment queries.
User-centered Programmatic Data Labeling · Georgia Tech
Tried and failed
zero-shot transfer of pre-trained document layout models applied to long-form academic dissertations. Outcome: did not generalise. Reason: Distribution shift between short research paper training sets and complex long-form dissertation formats.
Analyzing and Navigating Electronic Theses and Dissertations · Virginia Tech
Tried and failed
pretrained cross-modal feature matching networks applied to zero-shot airborne 2D-3D data. Outcome: did not generalise. Reason: extreme domain shift between ground training data and airborne perspectives yielding negligible inlier ratios
Concurrent adjustment of active and passive optical sensors with GNSS and raw inertial data · EPFL
Tried and failed
zero-shot policy transfer across network topologies applied to optimization algorithm parameter tuning. Outcome: did not generalise. Reason: policies trained on one graph structure fail on differently sized or structured graphs without retraining
Designing policy optimization algorithms for multi-agent reinforcement learning · Georgia Tech
Signal and physiological variability cause zero-shot transfer to fail in acoustic and biological domains
Zero-shot models experience drastic performance drops when transferring across different human subjects, atypical vocal patterns, or unseen language phonetics. Generic pretrained representations fail to capture subject-specific functional features or subtle acoustic distinctions without individualized tuning.
Tried and failed
zero-shot cross-subject classification from decomposed signals applied to dexterous movement intent decoding. Outcome: did not generalise. Reason: inter-individual variability in physiological signal representations degraded generalization performance across subjects
Decoding peripheral neural correlates of dexterous movements · Imperial
Tried and failed
zero-shot cross-subject transfer of neural networks applied to fMRI brain response prediction. Outcome: did not generalise. Reason: models trained exclusively on other individuals failed to capture subject-specific functional representations without individualized readouts
Characterizing human vision through large-scale brain imaging and computational models · MIT
Lost to a baseline
Direct zero-shot Seq2Seq transfer yielded near or exceeding 100% WER, severely underperforming modular hybrid DNN-HMM systems.
Automatic Speech Recognition without Transcribed Speech or Pronunciation Lexicons · JScholarship
Tried and failed
zero-shot transfer learning with pretrained audio embeddings applied to idiosyncratic atypical vocalization classification. Outcome: no signal. Reason: generic pretrained representations failed to capture idiosyncratic acoustic patterns of atypical vocalizations
Foundations of Cognitive, Affective, and Communicative Systems for Neurodiverse Individuals · MIT
Tried and failed
zero-shot multimodal large language models applied to audio event extraction. Outcome: no signal. Reason: Models completely failed to extract audio events, resulting in zero F1 performance across all tests
M3EC: A Multimodal Multidocument Benchmark for Event Extraction and Coreference Resolution · Virginia Tech
Tried and failed
zero-shot cross-lingual transfer of sub-phonetic representations applied to unseen target language speech recognition. Outcome: did not generalise. Reason: source language lacked phonetic distinctions (like voicing contrasts) required to resolve target language syllables
Modeling of Language-Universal Speech Attributes for Multilingual Speech Recognition and Processing · Georgia Tech
Lost to a baseline
In zero-shot multilingual SLU intent classification, the proposed zero-shot Speech-LLM (16.24%–47.38%) lost to the text NLU baseline (70%–80%) and cascaded SLU across all evaluated languages.
Multilingual Spoken Language Understanding: Efficient Speech Dataset Collection, Architectural Exploration, and Zero-Shot SLU · IRIS - UNITN - prod
Lost to a baseline
Generic zero-shot AudioSet transfer learning achieved only 51.1% accuracy on self-talk, performing barely above chance
Lost to a baseline
WaveNet slightly outperformed DiT in zero-shot SECS (0.307 vs 0.299) and CER (4.06% vs 4.43%) on GenerSpeech
Considered and rejected
Considered and rejected: Decided against using zero-shot automatic phoneme recognition models (e.g. XLS-R FAIR) because they learn surface-level acoustics without linguistic/phonemic contrasts.
Investigating the Corpus Phonetics Pipeline Applied to Diverse Speech Data · ResearchWorks
Lost to a baseline
In zero-shot gaze prediction on whole signals, high-privacy autoencoder signals (AE-0.1, AE-0.2) achieved worse prediction error (2.27°, 2.52°) than the baseline no-prediction approach.
Toward Privacy-Preserving Eye Tracking with Applications in Cross-Platform and Extended Reality Environments · TXST Digital Repository
Models exhibit severe class prediction bias, label imbalance, and high false alarm rates
Zero-shot classifiers and natural language inference pipelines frequently collapse into majority-class predictions or over-predict neutral categories. Uncalibrated thresholding and prompt sensitivity also lead to elevated false positive rates and poor recognition of rare classes.
Tried and failed
NLI-based zero-shot classification applied to thematic categorization of domain-specific text. Outcome: did not generalise. Reason: the model exhibited severe label bias, classifying the vast majority of samples into a single category
Capturing Elusive Metrics in Economics · unevada
Considered and rejected
Considered and rejected: Using purely zero-shot NLP classification without deterministic adjustments; rejected because models favored brief answers and scored unrelated phrases.
Weaving AI into Society: The Co-Evolution of Artificial Intelligence and the social fabric of organizations · IRIS - LUISS - prod
Tried and failed
prompt-based zero-shot large language model classification applied to student text behavior classification. Outcome: worse than baseline. Reason: produced substantially lower AUC and more than double the false positives compared to embedding-based classifiers
Tried and failed
zero-shot prompt-based LLM classification applied to multi-class text categorization. Outcome: worse than baseline. Reason: None
Evaluating Human-LLM Alignment in ETD Subject Classification New Trends in Theory and Practice of Digital Libraries, TPDL 2025 · Virginia Tech
Tried and failed
zero-shot prompting of large language models applied to time series anomaly detection. Outcome: worse than baseline. Reason: produced high false alarm rates or generic text and code instead of precise anomaly intervals
Machine Learning Systems for Unsupervised Time Series Anomaly Detection · MIT
Tried and failed
Adding predicate confidence score to scoring function applied to parse tree simplification for event extraction. Outcome: worse than baseline. Reason: Model over-relied on learned language priors rather than generalising in the zero-shot setting
Towards Explainable Event Detection and Extraction · Virginia Tech
Tried and failed
zero-shot stance detection using general-purpose embedding models applied to political and congressional speech text. Outcome: did not generalise. Reason: lacked domain-specific adaptation, causing models to heavily over-predict neutral labels
Tried and failed
zero-shot LLM classification on rare classes applied to noisy text transcript classification. Outcome: data insufficient. Reason: low prevalence and high semantic ambiguity resulted in poor F1 scores
Local Television News in a Changing Media Environment · Penn
Tried and failed
zero-shot NLI for text classification applied to domain-specific technical requirements classification. Outcome: did not generalise. Reason: severe class prediction bias caused poor performance on specialized domain text
Standardization of Engineering Requirements using Large Language Models · Georgia Tech
Tried and failed
zero-shot co-training with soft prompting applied to weakly supervised question answering. Outcome: worse than baseline. Reason: extremely high label noise rate prevented pseudo-label refinement from improving over the initial zero-shot model
Learning from Weak Supervision: Theory, Methods, and Applications · MIT
Left open by the authors
Problems the authors named and did not get to.
Left open
Develop zero-shot or transfer learning AI prognostic systems for newly deployed engineering assets lacking operational or failure data. Blocker: The task is an open-ended research direction without specific target metrics, assets, or concrete methodologies proposed.
Left open
Evaluate generalized zero-shot learning for cross-modal IMU activity recognition by classifying both seen and unseen classes at test time. Blocker: None
Cross-modal learning from visual information for activity recognition on inertial sensors · Oxford
Left open
Develop and evaluate methods to resolve zero-shot classification failure of BioTrove-CLIP models on the Confounding-Species benchmark. Blocker: None
Curating the world’s largest biodiversity dataset for AI · Iowa State
Left open
Reduce the zero-shot scene graph relationship prediction performance gap between small and large vision-language models for resource-constrained environments. Blocker: None
Zero-Shot Scene Graph Relationship Prediction using VLMs · Virginia Tech
Left open
Improve zero-shot vision-language models for scene graph generation to match or exceed supervised recall benchmarks while preserving generalization. Blocker: None
Zero-Shot Scene Graph Relationship Prediction using VLMs · Virginia Tech
Left open
Evaluate the effect of representation dimension and dataset coverage on Proto-Successor Measure zero-shot RL performance. Blocker: None
Reinforcement learning beyond rewards : decision-making in the language of visitation distributions · UT Austin
Left open
Evaluate zero-shot, one-shot, and few-shot prompting on LLM annotation performance for pedagogical discourse moves in classroom transcripts. Blocker: Access to the annotated elementary classroom transcript dataset.
Left open
Evaluate the zero-shot stance detection pipeline across additional domains and distinct NLP task settings beyond climate change. Blocker: None
Advancing stance detection and fine-grained content analysis for socially relevant domains · Leibniz Universität Hannover Repository
Left open
Evaluate alternative distance metrics, prototype encodings, and transductive zero-shot learning settings in Logic Tensor Networks with vision backbones. Blocker: None
Robust machine learning models for high dimensional data interpretation · IRIS - POLITO - prod
Left open
Apply few-shot, zero-shot, and one-shot learning methods to improve classification performance on tail classes in NSL-KDD, UNSW-NB15, and CIC-IDS-2018. Blocker: None
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.