Chapter Four · failure evidence
What Sensor Fusion & Multimodal Integration got wrong, from 94 dissertations
Across numerous multimodal integration studies, adding modalities or expanding sensor networks frequently degraded performance compared to standalone unimodal baselines. Failures stemmed from naive feature concatenation, sensory noise, synchronization offsets, and physical deployment constraints that corrupted joint representations and tracking stability. These records come from PhD theses at 26 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Multi-sensor fusion underperforms single-sensor or unimodal baselines due to sensory noise and feature redundancy
Fusing all available sensor channels often increases estimation error and degrades classification compared to relying on a single dominant sensor modality. Uncorrelated signals, redundant inputs, and cross-sensor interference frequently degrade accuracy and harm cross-subject generalization.
Tried and failed
multi-sensor fusion with additional inertial measurement units applied to locomotion speed estimation. Outcome: worse than baseline. Reason: additional sensors did not reduce error and degraded user-independent cross-subject generalization
Sensor Fusion Representation of Locomotion Biomechanics with Applications in the Control of Lower Limb Prostheses · Georgia Tech
Tried and failed
multimodal fusion using all available sensors applied to beam prediction in wireless networks. Outcome: worse than baseline. Reason: excessive sensor integration introduced redundant or noisy features that slightly degraded accuracy
Reimagining wireless networks: Generative AI for integrated sensing and communications · Iowa State
Tried and failed
expanding multimodal sensor fusion for regression applied to cross-subject kinematic speed estimation. Outcome: worse than baseline. Reason: adding sensors beyond minimal IMU setup plateaued performance and increased unfiltered estimation error
Improving Intelligence of Robotic Lower-Limb Prostheses to Enhance Mobility for Individuals with Limb Loss · Georgia Tech
Tried and failed
multimodal sensor fusion using PCA applied to heterogeneous time-series signal classification. Outcome: worse than baseline. Reason: force sensor signals negatively interfered with current sensor features during fusion
Development of deep meta-learning framework for cross-domain multisensory systems · Georgia Tech
Tried and failed
multimodal feature fusion with explicit posteriors applied to video action quality assessment. Outcome: worse than baseline. Reason: multimodal fusion interference degraded performance compared to single-modality baseline
Advancing Phonology-Based Sign Language Assessment: From Learner to Machine-Generated Videos · EPFL
Tried and failed
feature-level multi-band sensor fusion applied to ship wake detection in SAR imagery. Outcome: worse than baseline. Reason: multi-modal feature fusion introduced noise and degraded classification performance compared to single best band
Machine Learning and Data Fusion of Simulated Remote Sensing Data · Virginia Tech
Tried and failed
multi-axis sensor feature concatenation in naive Bayes applied to accelerometer-based state classification. Outcome: worse than baseline. Reason: additional axes introduced noise or violated conditional independence, lowering classification accuracy compared to single axis
Improved Vehicle Dynamics Sensing during Cornering for Trajectory Tracking using Robust Control and Intelligent Tires · Virginia Tech
Tried and failed
multichannel sensor fusion for anomaly detection applied to multichannel physiological time series. Outcome: worse than baseline. Reason: distant or localized anomalous activity across channels increased false positive rate during fusion
Real-time Personalized Monitoring of Neurological Disorders on Wearable Systems · EPFL
Lost to a baseline
Single combined SVM outperformed sensor fusion on PC and camera events where only 1-2 dominant sensors (power meters) existed
Lost to a baseline
Single Low-Speed HS3 sensor slightly outperformed the fused MT1 60 + Low-Speed HS3 BNN model due to MT1 60 inconsistency at low power.
Uncertainty quantification of faults in rotating machines · Texas Tech
Lost to a baseline
In vacuum chicken batch testing, single-sensor MSI (0.7933 RMSEp) outperformed all early fusion configurations (e.g., MSI-FTIR early fusion at 1.2745 RMSEp).
Sensor Approaches for the Non-Destructive Assessment of Food Safety and Authenticity · Cranfield
Lost to a baseline
Single vibration sensor (S3, 96.89%) beat dual-sensor fusion MS5 (VS+TS, 96.44%) under KAT B1 working condition.
Development of deep meta-learning framework for cross-domain multisensory systems · Georgia Tech
Lost to a baseline
Single PDR modality alone (4.14 brpm MAE) beat all multi-modal fusion methods (MA-U-Net 5.14 brpm, SE-U-Net 4.44 brpm) during running
Enabling Accurate Cardiopulmonary Monitoring Using Machine Learning and a Chest-Worn Wearable Patch · Georgia Tech
Lost to a baseline
In combined aerobic+vacuum chicken batch testing, single-sensor MSI (0.8090 RMSEp) beat early-fusion MSI-FTIR (1.2714 RMSEp) and early-fusion MSI-FTIR-MSIF (1.0544 RMSEp).
Sensor Approaches for the Non-Destructive Assessment of Food Safety and Authenticity · Cranfield
Lost to a baseline
Under LOSO, the multimodal fusion model DualLSTM-F (0.857 mean rho) did not outperform the unimodal vision model CNN-V (0.898 mean rho) or CNN-LSTM-V (0.861 mean rho)
Human-robot cooperation for teleoperation in robotic surgery · Imperial
Lost to a baseline
Wrist Accelerometer alone baseline achieved 0.81 F1 score, outperforming Wrist Gyroscope + Hip Gyroscope (0.75 F1) and multi-sensor combinations containing gyroscopes.
Towards Improving the Real Time Performance of Smartfall System · TXST Digital Repository
Lost to a baseline
In chicken thigh batch-on-batch validation, single-sensor MSI (0.6147 RMSEp aerobic, 0.7933 RMSEp vacuum) outperformed all single-sensor FTIR (1.6040 RMSEp aerobic, 1.2924 RMSEp vacuum) and MSIF models (2.2473 RMSEp aerobic, 1.0976 RMSEp vacuum).
Sensor Approaches for the Non-Destructive Assessment of Food Safety and Authenticity · Cranfield
Considered and rejected
Considered and rejected: Rejected early fusion (concatenating DLG directly to input) because it degraded performance compared to range-image-only baselines
Pavement Crack Segmentation with Dense Local Geometry Features and Boundary Enhancement Loss · Georgia Tech
Tried and failed
multimodal RGB-D input fusion applied to 2D geometric thickness estimation. Outcome: worse than baseline. Reason: RGB modality introduced distracting appearance features that degraded depth-based geometric predictions
Multi-object scene completion using view-based network predictions · Imperial
Tried and failed
multimodal feature fusion via partial least squares applied to crop yield prediction. Reason: additional soil and field features provided redundant information already captured by remote sensing and weather
Soybean production management and remote sensing performance for grain yield prediction · Iowa State
Lost to a baseline
Full multimodal feature concatenation without SHAP feature selection performed worse or barely better than unimodal models (e.g., unaware MW RF F1 was 0.057 vs unimodal Video RF F1 of 0.081).
Multimodal Machine Learning for Automated Assessment of Attention-Related Processes during Learning · Publikationssystem UB Tuebingen
Lost to a baseline
BERT+LSTM + ResNet50 multimodal model (70.00% F1-score) performed worse than unimodal VGG19 visual classification alone (75.30% F1-score).
AI for social good: social media mining of migration discourse · Leibniz Universität Hannover Repository
Lost to a baseline
Under LOUO, multimodal DualLSTM-F (0.671 mean rho) was beaten by unimodal vision model CNN-LSTM-V (0.843 mean rho)
Human-robot cooperation for teleoperation in robotic surgery · Imperial
Tried and failed
multimodal feature concatenation for classification applied to infant gait pattern classification. Outcome: worse than baseline. Reason: integrating additional sensor features introduced noise that degraded model performance
Walking and talking: gait and its role in early development in autism · Penn
Tried and failed
retaining uncorrelated sensor features in prognostic models applied to remaining useful life prediction. Outcome: worse than baseline. Reason: uncorrelated and non-informative sensor channels introduced noise and degraded prognostic accuracy
Uncertainty quantification of faults in rotating machines · Texas Tech
Tried and failed
multimodal classification omitting primary kinematic modality applied to repetitive behavior classification. Outcome: worse than baseline. Reason: physiological signals lacked sufficient discriminative power without dominant motion features
Explainable and Robust Data-Driven Machine Learning Methods for Digital Healthcare Monitoring · Virginia Tech
Tried and failed
expanding feature set with noisy sensor modalities applied to interruptibility classification models. Outcome: worse than baseline. Reason: inclusion of noisy body orientation and audio angle features degraded classifier performance
Facilitating Reliable Autonomy with Human-Robot Interaction · Georgia Tech
Lost to a baseline
In normal conditions without sensor failure, standalone INS/GPS achieved slightly lower position RMSE (1.77 m) and velocity RMSE (0.73 m/s) compared to INS/VO/GPS DEKF (2.01 m position RMSE and 0.73 m/s velocity RMSE).
Multi-Sensor Fusion for Navigation of Ground Vehicles · Carleton University Institutional Repository
Lost to a baseline
In Motion 2 for Subject 1, event-based-only tracking was more accurate along x and y axes than the Kalman sensor fusion output (fusion error was ~24% higher for forearm).
Effective and safe framework for human-robot interaction · IRIS - POLITO - prod
Lost to a baseline
Modified stereo vision approach achieved 2.8 cm translation error, beating the proposed camera+IMU fusion method (4.7 cm), though stereo was twice as slow and lacked robust rotation.
Sensor Fusion and Stroke Learning in Robotic Table Tennis · Publikationssystem UB Tuebingen
Naive early feature concatenation and uniform integration cause modality dominance and representation collapse
Concatenating raw multimodal representations without adaptive weighting allows high-dimensional or dominant modalities to overshadow informative inputs. This uniform integration leads to representation space collapse, corrupted cross-modal alignment, and degraded downstream predictive accuracy.
Tried and failed
multimodal feature concatenation for anomaly detection applied to mechanical fault detection. Outcome: worse than baseline. Reason: redundant multimodal features degraded anomaly detection performance compared to single-sensor baselines
Practical and generally applicable condition based maintenance (CBM) system for mud pump · UT Austin
Tried and failed
direct feature addition or concatenation across modalities applied to RGB-D salient object detection. Outcome: worse than baseline. Reason: depth map noise corrupted multi-scale feature representations during naive early/mid fusion
Effective deep leaning methodologies for salient object detection · Imperial
Tried and failed
early fusion via raw feature concatenation applied to multimodal regression with high dimensional imbalance. Outcome: worse than baseline. Reason: large modality dimensionality disparities caused higher-dimensional inputs to dominate predictions
Sensor Approaches for the Non-Destructive Assessment of Food Safety and Authenticity · Cranfield
Tried and failed
concatenating multiple heterogeneous visual feature representations applied to multimodal regression prediction. Outcome: worse than baseline. Reason: feature redundancy and multicollinearity degraded predictive performance across models
Housing Price Prediction with Computer Vision and Image Features · Harvard
Tried and failed
direct multimodal imaging feature fusion applied to multimodal medical image classification. Outcome: worse than baseline. Reason: increased data complexity without clinical context failed to outperform unimodal baselines
Advancing Personalized Medicine Through Generative Artificial Intelligence · Georgia Tech
Tried and failed
concatenation feature fusion for multimodal representations applied to visual question answering. Outcome: worse than baseline. Reason: underperformed multiplicative and bilinear fusion methods across language model architectures
Tried and failed
multimodal network fed duplicate single modality applied to multimodal fusion models. Outcome: overfit. Reason: duplicate input branches increased capacity without added information, leading to memorization rather than ensembling
Multimodal and Context-Aware Computational Pathology · Harvard
Tried and failed
early fusion of visual modalities before language cross-attention applied to vision-and-language robot navigation. Outcome: worse than baseline. Reason: joint multimodal visual representation degraded cross-modal alignment with language compared to separate per-modality cross-attention
Learning 3D Robotics Perception using Inductive Priors · Georgia Tech
Tried and failed
concatenation followed by linear projection applied to multimodal feature map fusion. Reason: increased parameter count without significant performance gains over adaptive fusion
Tried and failed
naive concatenation of multimodal features applied to few-shot relation extraction. Outcome: worse than baseline. Reason: irrelevant visual noise degraded representation quality without selective multimodal fusion
Few-Shot and Zero-Shot Learning for Information Extraction · Virginia Tech
Tried and failed
multimodal fusion with uniform layer learning rates applied to multimodal imitation learning. Reason: Dominant image modalities caused network inattention to fused scalar inputs without differential learning rates.
Towards Improving and Extending Traditional Robot Autonomy with Human Guided Machine Learning · Virginia Tech
Considered and rejected
Considered and rejected: Rejected naive channel concatenation for multi-modal feature fusion between objects and RGB because it weights channels uniformly instead of selectively attending to motion.
Model-driven and Data-driven Methods for Recognizing Compositional Interactions from Videos · JScholarship
Considered and rejected
Considered and rejected: Rejected naive early fusion (feature vector concatenation) and late fusion (decision averaging) due to information loss on heterogeneous symbolic data and inability to capture cross-view correlations.
Novel methods for multi-view learning with applications in cyber security · Imperial
Considered and rejected
Considered and rejected: Rejected completely shared multimodal feature spaces/complete modality fusion because it introduces heavy bias towards concrete words and harms abstract concepts
Language Grounding in Vision · Publikationssystem UB Tuebingen
Considered and rejected
Considered and rejected: Rejected shallow modality separation (modality-specific FFNs only with shared QKV) in LLaMaFusion because shared QKV projections corrupted text representation spaces during image training.
Breaking the language model monolith · ResearchWorks
Considered and rejected
Considered and rejected: Rejected joint multi-modal fusion strategies (Linear, Bilinear, Concat, and Vanilla SA) for audio-visual egocentric gaze anticipation due to camera motion and audio-gaze reaction latency.
Multimodal Human Behavior Modeling: From Understanding to Generation · Georgia Tech
Considered and rejected
Considered and rejected: Rejected Late Fusion (sharing predicted bounding boxes) due to lack of contextual feature representations and heavy dependence on single-agent accuracy.
Enhancing Perception for Autonomous Vehicles · Queens University Institutional Repository
Tried and failed
incorporating distinct non-shared features without adaptive weighting applied to multimodal single-cell data integration. Outcome: worse than baseline. Reason: distinct features were excessively noisy relative to shared features, degrading alignment accuracy
Tried and failed
multimodal joint fusion with single unified loss applied to multimodal clinical decision support. Reason: modality competition caused representation space collapse
Data-driven multimodal learning towards safer clinical decision support · Imperial
Multi-sensor wearable and edge deployments fail due to hardware overhead and inter-device variability
Complex multi-sensor wearable setups encounter significant practical hurdles including packet dropouts, inter-device variability, and high calibration complexity. Deploying multiple physical sensors also incurs prohibitive computational burdens and reduces participant compliance compared to simpler single-sensor configurations.
Considered and rejected
Considered and rejected: Rejected early and intermediate feature-level sensor fusion across vehicles due to high bandwidth constraints and incompatibility across diverse vehicle sensor suites.
Enhancing Perception Systems using V2V Sensor Fusion · Virginia Tech
Considered and rejected
Considered and rejected: Rejected using multi-IMU sensor fusion for model position in favor of a single master IMU to avoid inter-sensor noise and calibration discrepancies
Improving Usability for Novices in the Design of Mechatronic Devices: A Study Using Arduino Modules · Queens University Institutional Repository
Considered and rejected
Considered and rejected: Rejected activating all vehicle sensors simultaneously due to rapid battery depletion without coverage improvement.
Convergence Results for Ergodic Control of Ensembles via Iterated Function Systems · Research Repository UCD
Tried and failed
integrating heterogeneous consumer wearable sensor streams applied to remote physical activity tracking. Reason: inter-device measurement variability undermined data comparability across study participants
Tried and failed
raw wearable sensor telemetry for sequence modeling applied to joint moment estimation. Reason: sensor data suffered from packet dropout and signal saturation during dynamic movement
A Framework for Autonomous Exoskeleton Assistance Independent of Activity · Georgia Tech
Tried and failed
wearable sensor heuristic activity classification applied to travel mode and physical activity detection. Outcome: worse than baseline. Reason: built-in proprietary algorithms showed low accuracy compared to validated self-reported travel diaries
Applications of causal inference in environmental policy and transport studies · Imperial
Lost to a baseline
Wrist-worn sensors underperformed chest-worn sensors across all 5 holding assessment scenarios (0.738 vs 0.870 accuracy in Scenario 1)
Leveraging pervasive data to study and support mother-infant dyads in the wild · UT Austin
Considered and rejected
Considered and rejected: Rejected signal resampling and zero-padding for multiresolution data fusion due to excessive computational and memory overhead in wearable edge devices
Multi-sensor data fusion for ambulatory health monitoring: signal processing and deep learning techniques · Research Repository UCD
Considered and rejected
Considered and rejected: Rejected multi-sensor IMU systems (gyroscopes/magnetometers) in favor of standalone wrist accelerometers because multi-sensor setups introduce sensor drift, synchronization errors, and reduce patient compliance.
Considered and rejected
Considered and rejected: Rejected placing dual accelerometers on both hip and chest/wrist to avoid reducing patient compliance with multi-sensor wear.
Improving the surgical patient care pathway through use of activity monitoring · Oxford
Considered and rejected
Considered and rejected: Rejected using wearable sensor spectral monitoring data and activity diaries for quantitative prior-light-history tracking due to sensor inaccuracies/malfunctioning and inconsistent participant reporting.
Alertness in work environments : on the role of indoor daylight exposure · EPFL
Considered and rejected
Considered and rejected: Rejected cloth-based Bioharness sensor due to poor fit and low data quality.
Considered and rejected
Considered and rejected: Rejected continuous high-density EMG for closed-loop peripheral phase tracking in favor of single-sensor triaxial accelerometry to minimize instrumentation and improve wearable compliance.
Considered and rejected
Considered and rejected: Rejected bi-directional finger with two sensors design due to higher calibration complexity, sensor performance inconsistency, and mechanical interference/strain.
Capacitive Strain Sensor System for Soft-Rigid Hybrid Robotic Grippers · Harvard
Inertial navigation and kinematic state estimation diverge from compounding sensor drift and dynamic motion noise
Combining inertial measurement units and odometry data without robust compensation leads to rapid tracking divergence caused by compounding sensor bias and drift. Aggressive dynamic maneuvers, unmodeled surface friction, and environmental magnetic interference further corrupt state estimates.
Tried and failed
onboard camera and IMU sensor fusion applied to mobile robot state estimation and tracking. Outcome: no signal. Reason: low-cost onboard sensors suffered from excessive noise and drift for accurate positioning
Risk-aware and robust decision making for autonomous vehicles with reinforcement learning · Imperial
Tried and failed
redundant sensor fusion scaling applied to inertial navigation state estimation. Outcome: no signal. Reason: sensor systematic biases dominated over random noise, rendering additional sensors unhelpful
Identification and Integration of Aerodynamics into Fixed-Wing Drone Navigation · EPFL
Tried and failed
extended Kalman filter wheel odometry fusion applied to mobile robot track navigation. Outcome: unstable. Reason: unmodeled variable frictional slip between drive wheels and contact surfaces degraded state estimation
A novel railway maintenance robot for inspection and repair · Cranfield
Tried and failed
open-loop IMU integration without bias estimation applied to visual-inertial motion tracking. Outcome: unstable. Reason: sensor drift and integration errors compound with lower frame rates and lower-quality inertial sensors
Multi-modal 3D Gaussian Splatting for SLAM · UT Austin
Tried and failed
Doppler-inertial sensor fusion applied to satellite constellation radio navigation. Outcome: unstable. Reason: insufficient line-of-sight velocity vector diversity from co-aligned polar orbit trajectories
Navigation using Radio-Frequency Observables from LEO Constellations with Possible Aiding from an Inertial Navigation System · Virginia Tech
Tried and failed
direct LiDAR-inertial odometry fusion without redundancy applied to aerial state estimation. Outcome: unstable. Reason: intermittent sensor packet loss caused catastrophic estimation divergence without secondary vision-inertial fallback
Semantics-Driven Active Perception and Navigation with Aerial Robots · Penn
Tried and failed
standard complementary and Kalman filtering applied to low-cost mobile inertial sensors. Outcome: worse than baseline. Reason: gyroscope drift and high accelerometer noise during dynamic vehicle motions corrupted orientation estimation
Crowd-sourced Road Geometry and Accurate Vehicle State Estimation Using Mobile Devices · DSpace at SUNY Buffalo
Tried and failed
rolling statistics and sensor fusion orientation features applied to inertial sensor activity classification. Outcome: worse than baseline. Reason: engineered features degraded deep temporal model classification performance compared to raw sensor data
Tried and failed
rigidly mounted visual-inertial sensors applied to dynamic robotic trajectory tracking. Outcome: worse than baseline. Reason: dynamic maneuvers induced severe motion blur and higher tracking error compared to active mechanical stabilization
Calibration and estimation for aerial robots with application to additive manufacturing · Imperial
Tried and failed
Track-to-track fusion using covariance intersection applied to multimodal multi-sensor object tracking. Outcome: worse than baseline. Reason: extreme divergence between individual sensor track error and estimated covariance
Radar and LiDAR Fusion for Scaled Vehicle Sensing · Virginia Tech
Lost to a baseline
Global-odometry (EKF fusion) had a higher standard deviation of absolute pose error (0.5667 m) compared to RTAB-Map-odometry (0.5367 m)
Autonomous localization and navigation for a railway inspection and repair system · Cranfield
Considered and rejected
Considered and rejected: Rejected inclusion of magnetometers in IMU sensor fusion due to distortion from ferromagnetic industrial equipment and the robot itself
Human upper body motion tracking for human-machine interaction in industrial applications · IRIS - POLITO - prod
Considered and rejected
Considered and rejected: Rejected IMU yaw orientation angles and thigh acceleration signals from multi-modal GRF estimation pipelines due to sensor drift and packet loss across subjects.
Fusion algorithms and aggregation rules degrade under outliers, process noise, or sensory corruption
Common fusion rules such as sample mean aggregation or modular filtering break down when individual sensor measurements contain biases or non-zero mean noise. Centralized fusion architectures also experience severe detection failures when sensors encounter adversarial blinding attacks or severe domain distribution shifts.
Tried and failed
global sample mean aggregation applied to redundant sensor fusion with bias. Outcome: worse than baseline. Reason: vulnerable to outlier or biased sensor measurements without robust filtering
Statistical methods to obtain accurate estimates of measured process variables for redundant sensor measurements · Iowa State
Tried and failed
averaging closest pair of redundant estimates applied to redundant sensor data fusion. Outcome: worse than baseline. Reason: yielded higher mean squared error than using the simple sample median
Statistical methods to obtain accurate estimates of measured process variables for redundant sensor measurements · Iowa State
Tried and failed
past traversal visibility fusion offline adaptation applied to unsupervised 3D LiDAR object detection. Outcome: worse than baseline. Reason: caused performance drops for visible cars and non-occluded pedestrians relative to the baseline
PERCEPTION FOR AUTONOMOUS VEHICLES IN CHALLENGING WEATHERS AND OCCLUDED ENVIRONMENTS · Cornell
Lost to a baseline
Modular Bayesian fusion achieved lower accuracy than non-modular combined KNN when all sensors were fully available and trained together (accuracy ratio ~0.8-0.9).
Lost to a baseline
FAIEKF lost to FAUKF in multi-sensor fusion innovation sequence drift and accuracy under non-zero mean process noise (mu = 1 m).
Vision based Real-Time Navigation with Unknown and Uncooperative Space Target · Carleton University Institutional Repository
Lost to a baseline
Concatenated early multimodal feature fusion (AUROC 0.53–0.63) performed worse than unimodal HMM heart rate dynamics alone (AUROC 0.68–0.75) and late fusion (AUROC 0.72–0.82)
Multimodal assessment of neuropsychiatric disorders using audiovisual recordings · Georgia Tech
Lost to a baseline
Existing sensor fusion algorithms (F-PointNet, MV3D, AVOD) experienced perception detection failures across 98% to 99% of bundles under camera blinding attacks, and 79% to 84% under LIDAR rotation error attacks.
Secure and reliable deep learning in signal processing · Virginia Tech
Considered and rejected
Considered and rejected: Centralized sensor fusion architectures rejected in favor of federated architectures due to vulnerability/lower resilience to single-sensor failures
Robust autonomous navigation for UAVS in urban environments using machine learning · Cranfield
Considered and rejected
Considered and rejected: Rejected raw data-level 1-D stacking/matrix fusion without dimensionality reduction due to lack of error correction and inability to handle asymmetric/heterogeneous multi-modal sensors.
Development of deep meta-learning framework for cross-domain multisensory systems · Georgia Tech
Tried and failed
multimodal video feature fusion applied to cross-dataset persuasion strategy prediction. Outcome: did not generalise. Reason: large visual domain gaps between datasets impaired out-of-domain performance
Multimodal Human Behavior Modeling: From Understanding to Generation · Georgia Tech
Tried and failed
adding perceptual loss to multimodal generative model applied to sequential view synthesis from incomplete point clouds. Outcome: did not generalise. Reason: perceptual loss caused RGB-depth inconsistency, degrading multi-step rollout quality despite better single-step visual fidelity
Towards multi-modal AI systems with open-world cognition · Georgia Tech
Tried and failed
multimodal training with mutual information loss applied to cross-modality deformable image registration. Outcome: worse than baseline. Reason: unified cross-modality models underperformed compared to training separate modality-specific models
Deep Learning Methods to Process and Analyse MRI Images · Cornell
Tried and failed
multimodal context fusion with noisy visual features applied to human motion prediction. Outcome: did not converge. Reason: noisy background visual signals and inaccurate automated pose annotations caused training divergence
On the motion and action prediction using deep graph models · UT Austin
Multi-sensor integration fails when measurements suffer from timing mismatches, spatial misalignment, and association errors
Fusing observations across asynchronous or spatially separated sensors degrades tracking accuracy when timestamps and arrival rates do not align. Underestimated timing offsets, association errors, and view registration discrepancies result in the integration of out-of-sync or misaligned data.
Tried and failed
late-stage Kalman filter multi-sensor fusion applied to collaborative multi-object tracking. Outcome: worse than baseline. Reason: noisy local sensor measurements corrupted precise communicated baseline data during fusion
Enhancing Perception Systems using V2V Sensor Fusion · Virginia Tech
Tried and failed
synchronous sensor fusion state estimation applied to multi-sensor 3D state tracking. Outcome: worse than baseline. Reason: measurement arrival rate differences and timestamp mismatches across sensors degraded accuracy
3D INDOOR STATE ESTIMATION FOR RFID-BASED MOTION-CAPTURE SYSTEMS · Georgia Tech
Tried and failed
temporal multi-frame fusion with ego-motion odometry applied to monocular 3D object pose estimation. Outcome: worse than baseline. Reason: sparse low-frequency keypoint tracking and ego-motion noise degraded performance compared to single-frame geometric priors
Reshaping Perception for Autonomous Driving with Semantic Keypoints · EPFL
Tried and failed
Multi-sensor fusion across overlapping peripheral fields of view applied to Cross-traffic object state estimation. Reason: Sensor alignment and detection association failures degraded position and velocity estimation in lateral viewing angles
Assessing Effects of Object Detection Performance on Simulated Crash Outcomes for an Automated Driving System · Virginia Tech
Tried and failed
search window expansion for synchronization error applied to distributed sensor fusion. Outcome: worse than baseline. Reason: underestimating timing offsets causes fusion of corrupted, out-of-sync measurements as reliable data
Statistical Multistatic Radar with Imperfect Time Synchronization · Virginia Tech
Considered and rejected
Considered and rejected: Multi-sensor coordination frameworks: rejected due to complex cross-sensor time synchronization, bias estimation latencies, and registration errors.
Hypersonic: Real-Time Software Architecture for EM-based Radar Signal Processing and Tracking · Georgia Tech
Left open by the authors
Problems the authors named and did not get to.
Left open
Benchmark MW-FGO sensor fusion against traditional Kalman filtering across various sliding window sizes using public pedestrian GNSS/IMU datasets and GTSAM. Blocker: None
Advanced Inertial/GNSS Sensor Fusion for Smart Devices Using Factor Graph Optimization · DeustoTeka
Left open
Develop multi-sensor fusion combining IMU and FSR sensor data to reduce false positives in real-time step detection. Blocker: Requires physical wearable hardware equipped with synchronized IMU and FSR sensors or proprietary dual-sensor gait data.
Left open
Develop a multi-sensor fusion system combining radar or optical sensors for coarse target acquisition with 3D LiDAR for precision tracking of small UAS. Blocker: Requires multi-modal physical sensor hardware (LiDAR, radar, optical cameras) and synchronized real-world UAS tracking data.
DETECTION OF SMALL UNMANNED AERIAL SYSTEMS USING A 3D LIDAR SENSOR · Calhoun
Left open
Integrate fault-tolerant mechanisms and sensor fusion into the EMMA autonomous driving framework to handle missing state information and sensor failures. Blocker: None
Collaborative and safe autonomous driving through multi-agent deep reinforcement learning · Imperial
Left open
Develop and evaluate end-to-end multi-modal sensor fusion architectures for vehicle crash prediction using CARLA simulation data. Blocker: None
Vehicle crash prediction using transformer networks · Institutional Repository University of Moratuwa
Left open
Incorporate machine learning algorithms into the multi-robot localization fault detection module to detect sensor faults and failure modes. Blocker: Lack of specific ML architecture, failure mode definitions, or baseline fault detection code/dataset
From Slip Estimation to Collaborative Localization of Multi-Robot Systems Using Multi-Sensor Networks · Carleton University Institutional Repository
Left open
Evaluate whether connected vehicle sensor data alone provides sufficient accuracy for agency pavement management decision-making. Blocker: Requires access to proprietary connected vehicle fleet data and DOT pavement management standards
Pavement Surface Characteristics Evaluation Using Vehicle-Based Data Collection · Virginia Tech
Left open
Benchmark alternative machine learning models against gradient boosting for IMU/GNSS sensor fusion during GNSS outages. Blocker: No specific machine learning architectures, target metrics, or standardized benchmark datasets are defined
Improved IMU/GNSS EKF fusion using Machine Learning · Carleton University Institutional Repository
Left open
Investigate using autonomous vehicles as mobile traffic sensors to estimate traffic state and replace loop detectors or cameras. Blocker: None
Modelling mixed traffic flow of autonomous vehicles and human-driven vehicles · Imperial
Left open
Evaluate vision-based tracking using multi-sensor driving data collected from Cornell's autonomous vehicle platform. Blocker: Requires private sensor data collected from the Cornell autonomous test vehicle
PERCEPTION FOR AUTONOMOUS VEHICLES IN CHALLENGING WEATHERS AND OCCLUDED ENVIRONMENTS · Cornell
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.