e-ISSN: Pending

Browse the failure-mode index

694 real negative results, null findings, and replication failures in Computer Science · Negative / Null Result Report. Search the index →

WASTE indexes published research — it does not host or republish full papers. Each entry is a metadata record compiled from open scholarly databases; the abstract is shown in full only where the paper is openly licensed, otherwise a short excerpt under fair use. Classifications are automated and approximate.

Negative / Null Result ReportOpen accessComputer Science

Tri-Bench: Stress-Testing VLM Reliability on Spatial Reasoning under Camera Tilt and Object Interference

Amit Bendkhale · 2025 · arXiv

Verifiable geometric reasoning is a critical component for trustworthy and controllable agentic AI. Despite impressive capabilities, Vision-Language Models (VLMs) often fail under realistic scene changes. We present Tri-Bench, a compact benchmark of planar triangle problems that isolates relative geometric reasoning while stressing two deployment-critical factors: camera pose (planar vs. tilted) and scene context via object interference (10 everyday objects). To test verifiability and control, we evaluate four recent VLMs using a single, fixed prompt whose guardrail explicitly describes a surr

View details →
Negative / Null Result ReportOpen accessComputer Science

Exploring Gender Bias in Remote Pair Programming among Software Engineering Students: The twincode Original Study and First External Replication

Amador Durán, Pablo Fernández, Beatriz Bernárdez et al. · 2023 · arXiv

Context. Software Engineering (SE) has low female representation due to gender bias that men are better at programming. Pair programming (PP) is common in industry and can increase student interest in SE, especially women; but if gender bias affects PP, it may discourage women from joining the field. Objective. We explore gender bias in PP. In a remote setting where students cannot see their peers' gender, we study how perceived productivity, technical competency and collaboration/interaction behaviors of SE students vary by perceived gender of their remote partner. Method. We developed an onl

View details →
Negative / Null Result ReportOpen accessComputer Science

Trust in Generative AI for Health Information Consumption and the Effect of Learned Dependency: An Experimental Investigation

Arif Ahmed, Gondy Leroy, Agrim Sachdeva et al. · 2026 · arXiv

Background: Generative artificial intelligence (GenAI) is increasingly used for health information, yet its influence on users' trust calibration remains unclear. Objective: This study examines whether learned dependency on GenAI influences trust in AI-generated health information and whether text highlighting reduces overreliance on incorrect outputs. Methods: Two randomized controlled experiments were conducted with 338 college students and 563 Amazon Mechanical Turk participants. Both experiments used a 2 by 2 between-subjects design manipulating information accuracy (correct versus incorre

View details →
Negative / Null Result ReportOpen accessComputer Science

Don't Leave Me Alone: Retrospective Think Aloud supported by Real-time Monitoring of Participant's Physiology

Alexandros Liapis, Christos Katsanos, Michalis Xenos · 2018 · arXiv

Think aloud protocols are widely applied in user experience studies. In this paper, the effect of two different applications of the Retrospective Think Aloud (RTA) protocol on the number of user-reported usability issues is examined. To this end, 30 users were asked to use the National Cadastre and Mapping Agency web application and complete a set of tasks, such as measuring the land area of a square in their hometown. The order of tasks was randomized per participant. Next, participants were involved in RTA sessions. Each participant was involved in two different RTA modes: (a) the strict gui

View details →
Negative / Null Result ReportOpen accessComputer Science

Navigating with Haptic Gloves: Investigating Strategies for Horizontal and Vertical Movement Guidance

Mahdis Tajdari, Jason Forsyth, Sol Lim · 2025 · arXiv

Navigating peripersonal space requires reaching targets in both horizontal (e.g., desks) and vertical (e.g., shelves) layouts with high precision. We developed a haptic glove to aid peri-personal target navigation and investigated the effectiveness of different feedback delivery methods. Twenty-two participants completed target navigation tasks under various conditions, including scene layout (horizontal or vertical), guidance approach (two-tactor or worst-axis first), guidance metaphor (push or pull), and intensity mode (linear or zone) for conveying distance cues. Task completion time, hand

View details →
Negative / Null Result ReportOpen accessComputer Science

BSQ Conserved Charges in Relativistic Viscous Hydrodynamics solved with Smoothed Particle Hydrodynamics

Christopher Plumberg, Dekrayat Almaalol, Travis Dore et al. · 2024 · arXiv

Conservation laws play a crucial role in the modeling of heavy-ion collisions, including the those for charges such as baryon number (B), strangeness (S), and electric charge (Q). In this study, we present a new 2+1 relativistic viscous hydrodynamic code called CCAKE which uses the Smoothed Particle Hydrodynamics (SPH) formalism to locally conserve BSQ charges, together with an extended description of the multi-dimensional equation of state (EoS) obtained from lattice Quantum Chromodynamics. Initial conditions for CCAKE are supplied by the ICCING model, which samples gluon splittings into quar

View details →
Negative / Null Result ReportOpen accessComputer Science

An Empirical Security Evaluation of LLM-Generated Cryptographic Rust Code

Mohamed Elsayed, Kenneth Fulton, Jeong Yang · 2026 · arXiv

Developers and organizations are using Large Language Models (LLMs) to generate security-critical code more frequently than ever, including cryptographic solutions for their products. This study presents an empirical evaluation of cryptographic security in 240 Rust code samples for two crypto algorithms (AES-256-GCM and ChaCha20-Poly1305) generated by three LLMs (Gemini 2.5 Pro, GPT-4o, and DeepSeek Coder) using four different prompt strategies. For each successfully compiled code sample, CodeQL static analysis and our rule-based crypto-specific analyzer were used to detect vulnerabilities, wh

View details →
Negative / Null Result ReportOpen accessComputer Science

ConvMemory: A Lightweight Learned Memory Reranker, a Negative Attribution Result, and a Research-Preview Conflict Editor

Taiheng Pan · 2026 · arXiv

We describe ConvMemory, a small 3.6M-parameter learned reranker for conversational long-term memory retrieval, trained with cross-encoder teacher supervision over fused dense and lexical features. On the LongMemEval memory family, ConvMemory operates above the BGE-large cross-encoder in Recall@10 at 12-47x lower latency, remains within 0.025 Recall@10 of mxbai-rerank-large-v1 on Clean500 while running 28x cheaper; under Stress1000 distractors the Recall@10 gap widens to 0.081 but ConvMemory still operates at 117x lower latency; these LongMemEval numbers are single-run or single-seed and are re

View details →
Negative / Null Result ReportOpen accessComputer Science

Cooperation and Contagion in Web-Based, Networked Public Goods Experiments

Siddharth Suri, Duncan J. Watts · 2010 · arXiv

A longstanding idea in the literature on human cooperation is that cooperation should be reinforced when conditional cooperators are more likely to interact. In the context of social networks, this idea implies that cooperation should fare better in highly clustered networks such as cliques than in networks with low clustering such as random networks. To test this hypothesis, we conducted a series of web-based experiments, in which 24 individuals played a local public goods game arranged on one of five network topologies that varied between disconnected cliques and a random regular graph. In c

View details →
Negative / Null Result ReportOpen accessComputer Science

Towards the Design of Effective Freehand Gestural Interaction for Interactive TV

Gang Ren, Wenbin Li, Eamonn O'Neill · 2016 · arXiv

As interactive devices become pervasive, people are beginning to looking for more advanced interaction with televisions in the living room. Interactive television has the potential to offer a very engaging experience. But most common user tasks are still challenging with such systems, such as menu selection or text input. And little work has been done on understanding and sup-porting the effective design of freehand interaction with an TV in the living room. In this paper, we perform two studies investi-gating freehand gestural interaction with a consumer level sensor, which is suitable for TV

View details →
Negative / Null Result ReportOpen accessComputer Science

A Modular Transradial Bypass Socket for Surface Myoelectric Prosthetic Control in Non-Amputees

Michael D. Paskett, Nathaniel R. Olsen, Jacob A. George et al. · 2019 · arXiv

Bypass sockets allow researchers to perform tests of prosthetic systems from the prosthetic user's perspective. We designed a modular upper-limb bypass socket with 3D-printed components that can be easily modified for use with a variety of terminal devices. Our bypass socket preserves access to forearm musculature and the hand, which are necessary for surface electromyography and to provide substituted sensory feedback. Our bypass socket allows a sufficient range of motion to complete tasks in the frontal working area, as measured on non-amputee participants. We examined the performance of non

View details →
Negative / Null Result ReportOpen accessComputer Science

End-to-End Fidelity Analysis of Quantum Circuit Optimization: From Gate-Level Transformations to Pulse-Level Control

Rylan Malarchick · 2026 · arXiv

We present an analysis of quantum circuit fidelity across the full compilation stack, from high-level gate optimization through pulse-level control. We connect a C++ circuit optimizer to a per-gate Lindblad master-equation fidelity model whose decoherence channels are cross-validated against qiskit-dynamics and whose absolute predictions are benchmarked against execution on real hardware. Across a campaign of 4,452 experiment runs over 371 benchmark circuits, gate cancellation provides the dominant improvement ($d = 1.66$, 72% of circuits improved), while circuit size and pulse duration are th

View details →
Negative / Null Result ReportOpen accessComputer Science

Quark resonances and high E_t jets

Myron Bander · 1996 · arXiv

Possible spin-3/2 quark resonances would have a significant effect on high E$_{\mbox{\rm t}}$ jet production through their contribution to the subprocess $q+{\bar q}\rightarrow g+g$. Such enhancements are compared to a, recently reported, anomaly in inclusive jet production at the CDF detector.

View details →
Negative / Null Result ReportOpen accessComputer Science

Roles of $\bar{D}^{*}K^{*}$ and $D^*\bar{D}$ molecular states in decay $B^+ \to D^{*+} D^- K^+$

Zuo-Ming Ding, Qi Huang, Jun He · 2025 · arXiv

This study investigates the three-body decay process $B^+ \to D^{*+} D^- K^+$, aiming to explore the possible origins of $T^*_{\bar{c}\bar{s}0}(2870)^0$ and $χ_{c1}(3872)$ as intermediate states. Within the molecular state framework, $T^*_{\bar{c}\bar{s}0}(2870)^0$ and $χ_{c1}(3872)$ are considered as possible $\bar{D}^{*}K^{}$ and $D^*\bar{D}$ molecular states, respectively. Using effective Lagrangians, the interaction kernels of the $\bar{D}^{*}K^{*}$ and $D^*\bar{D}$ systems are constructed within the one-boson-exchange model. The corresponding rescattering amplitudes and pole positions are

View details →
Negative / Null Result ReportOpen accessComputer Science

Investigating the Robustness of Sequential Recommender Systems Against Training Data Perturbations

Filippo Betello, Federico Siciliano, Pushkar Mishra et al. · 2023 · arXiv

Sequential Recommender Systems (SRSs) are widely employed to model user behavior over time. However, their robustness in the face of perturbations in training data remains a largely understudied yet critical issue. A fundamental challenge emerges in previous studies aimed at assessing the robustness of SRSs: the Rank-Biased Overlap (RBO) similarity is not particularly suited for this task as it is designed for infinite rankings of items and thus shows limitations in real-world scenarios. For instance, it fails to achieve a perfect score of 1 for two identical finite-length rankings. To address

View details →
Negative / Null Result ReportOpen accessComputer Science

Impact of geolocation data on augmented reality usability: A comparative user test

Julien Mercier, N. Chabloz, G. Dozot et al. · 2023 · arXiv

Abstract. While the use of location-based augmented reality (AR) for education has demonstrated benefits on participants' motivation, engagement, and on their physical activity, geolocation data inaccuracy causes augmented objects to jitter or drift, which is a factor in downgrading user experience. We developed a free and open source web AR application and conducted a comparative user test (n = 54) in order to assess the impact of geolocation data on usability, exploration, and focus. A control group explored biodiversity in nature using the system in combination with embedded GNSS data, and

View details →
Negative / Null Result ReportOpen accessComputer Science

The effect of supersymmetric CP phases on Chargino-Pair Production via Drell-Yan Process at the LHC

Kerem Cankocak, Aytekin Aydemir, Ramazan Sever · 2004 · arXiv

We compute the rates for pp annihilation into chargino-pairs via Drell-Yan process taking into account the effects of supersymmetric soft phases, at proton-proton collider. In particular, the phase of the mu parameter gains direct accessibility via the production of dissimilar charginos. The phases of the trilinear soft masses do not have a significant effect on the cross sections.

View details →
Negative / Null Result ReportOpen accessComputer Science

More Rounds, More Noise: Why Multi-Turn Review Fails to Improve Cross-Context Verification

Song Tae-Eun · 2026 · arXiv

Cross-Context Review (CCR) improves LLM verification by separating production and review into independent sessions. A natural extension is multi-turn review: letting the reviewer ask follow-up questions, receive author responses, and review again. We call this Dynamic Cross-Context Review (D-CCR). In a controlled experiment with 30 artifacts and 150 injected errors, we tested four D-CCR variants against the single-pass CCR baseline. Single-pass CCR (F1 = 0.376) significantly outperformed all multi-turn variants, including D-CCR-2b with question-and-answer exchange (F1 = 0.303, $p < 0.001$, $d

View details →
Negative / Null Result ReportOpen accessComputer Science

An Analysis of BPE Vocabulary Trimming in Neural Machine Translation

Marco Cognetta, Tatsuya Hiraoka, Naoaki Okazaki et al. · 2024 · arXiv

We explore threshold vocabulary trimming in Byte-Pair Encoding subword tokenization, a postprocessing step that replaces rare subwords with their component subwords. The technique is available in popular tokenization libraries but has not been subjected to rigorous scientific scrutiny. While the removal of rare subwords is suggested as best practice in machine translation implementations, both as a means to reduce model size and for improving model performance through robustness, our experiments indicate that, across a large space of hyperparameter settings, vocabulary trimming fails to improv

View details →
Negative / Null Result ReportOpen accessComputer Science

Learning to Learn End-to-End Goal-Oriented Dialog From Related Dialog Tasks

Janarthanan Rajendran, Jonathan K. Kummerfeld, Satinder Singh · 2021 · arXiv

For each goal-oriented dialog task of interest, large amounts of data need to be collected for end-to-end learning of a neural dialog system. Collecting that data is a costly and time-consuming process. Instead, we show that we can use only a small amount of data, supplemented with data from a related dialog task. Naively learning from related data fails to improve performance as the related data can be inconsistent with the target task. We describe a meta-learning based method that selectively learns from the related dialog task data. Our approach leads to significant accuracy improvements in

View details →
Negative / Null Result ReportOpen accessComputer Science

Rethinking Neural Width for Alternating Current Optimal Power Flow Proxies

Dhruvi Khandelwal, Anurag Basistha, Ayushi Jolotia et al. · 2026 · arXiv

Deep learning proxies for Alternating Current Optimal Power Flow (ACOPF) lack systematic methods for determining architectural size. This paper conducts a constructive thought experiment to answer a fundamental inquiry: how wide must a neural network be to almost accurately approximate the ACOPF manifold? We introduce a Loss-Guided Neural Densification (LG-ND) algorithm that incrementally discovers necessary capacity by expanding only when the current deep neural network topology fails to improve further. Empirical results across various IEEE systems show that LG-ND achieves performance parity

View details →
Negative / Null Result ReportOpen accessComputer Science

Meta-Curriculum Learning for Domain Adaptation in Neural Machine Translation

Runzhe Zhan, Xuebo Liu, Derek F. Wong et al. · 2021 · arXiv

Meta-learning has been sufficiently validated to be beneficial for low-resource neural machine translation (NMT). However, we find that meta-trained NMT fails to improve the translation performance of the domain unseen at the meta-training stage. In this paper, we aim to alleviate this issue by proposing a novel meta-curriculum learning for domain adaptation in NMT. During meta-training, the NMT first learns the similar curricula from each domain to avoid falling into a bad local optimum early, and finally learns the curricula of individualities to improve the model robustness for learning dom

View details →
Negative / Null Result ReportOpen accessComputer Science

Not How Many, But Which: Parameter Placement in Low-Rank Adaptation

Arijit Sehanobish, Charles Lovering · 2026 · arXiv

We study the \textit{parameter placement problem}: given a fixed budget of $k$ trainable entries within the B matrix of a LoRA adapter (A frozen), does the choice of which $k$ matter? Under supervised fine-tuning, random and informed subsets achieve comparable performance. Under GRPO on base models, random placement fails to improve over the base model, while gradient-informed placement recovers standard LoRA accuracy. This regime dependence traces to gradient structure: SFT gradients are low-rank and directionally stable, so any subset accumulates coherent updates; GRPO gradients are high-ran

View details →
Negative / Null Result ReportOpen accessComputer Science

Joint Beamforming and Speaker-Attributed ASR for Real Distant-Microphone Meeting Transcription

Can Cui, Imran Ahamad Sheikh, Mostafa Sadeghi et al. · 2024 · arXiv

Distant-microphone meeting transcription is a challenging task. State-of-the-art end-to-end speaker-attributed automatic speech recognition (SA-ASR) architectures lack a multichannel noise and reverberation reduction front-end, which limits their performance. In this paper, we introduce a joint beamforming and SA-ASR approach for real meeting transcription. We first describe a data alignment and augmentation method to pretrain a neural beamformer on real meeting data. We then compare fixed, hybrid, and fully neural beamformers as front-ends to the SA-ASR model. Finally, we jointly optimize the

View details →
Negative / Null Result ReportOpen accessComputer Science

InsCLR: Improving Instance Retrieval with Self-Supervision

Zelu Deng, Yujie Zhong, Sheng Guo et al. · 2021 · arXiv

This work aims at improving instance retrieval with self-supervision. We find that fine-tuning using the recently developed self-supervised (SSL) learning methods, such as SimCLR and MoCo, fails to improve the performance of instance retrieval. In this work, we identify that the learnt representations for instance retrieval should be invariant to large variations in viewpoint and background etc., whereas self-augmented positives applied by the current SSL methods can not provide strong enough signals for learning robust instance-level representations. To overcome this problem, we propose InsCL

View details →
Negative / Null Result ReportOpen accessComputer Science

On the Transferability of Minimal Prediction Preserving Inputs in Question Answering

Shayne Longpre, Yi Lu, Christopher DuBois · 2020 · arXiv

Recent work (Feng et al., 2018) establishes the presence of short, uninterpretable input fragments that yield high confidence and accuracy in neural models. We refer to these as Minimal Prediction Preserving Inputs (MPPIs). In the context of question answering, we investigate competing hypotheses for the existence of MPPIs, including poor posterior calibration of neural models, lack of pretraining, and "dataset bias" (where a model learns to attend to spurious, non-generalizable cues in the training data). We discover a perplexing invariance of MPPIs to random training seed, model architecture

View details →
Negative / Null Result ReportOpen accessComputer Science

Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQL

Hanbing Liu, Haoyang Li, Xiaokang Zhang et al. · 2025 · arXiv

Direct Preference Optimization (DPO) has proven effective in complex reasoning tasks like math word problems and code generation. However, when applied to Text-to-SQL datasets, it often fails to improve performance and can even degrade it. Our investigation reveals the root cause: unlike math and code tasks, which naturally integrate Chain-of-Thought (CoT) reasoning with DPO, Text-to-SQL datasets typically include only final answers (gold SQL queries) without detailed CoT solutions. By augmenting Text-to-SQL datasets with synthetic CoT solutions, we achieve, for the first time, consistent and

View details →
Negative / Null Result ReportOpen accessComputer Science

The Adverse Effects of Omitting Records in Differential Privacy: How Sampling and Suppression Degrade the Privacy--Utility Tradeoff (Long Version)

Àlex Miranda-Pascual, Javier Parra-Arnau, Thorsten Strufe · 2026 · arXiv

Sampling is renowned for its privacy amplification in differential privacy (DP), and is often assumed to improve the utility of a DP mechanism by allowing a noise reduction. In this paper, we further show that this last assumption is flawed: When measuring utility at equal privacy levels, sampling as preprocessing consistently yields penalties due to utility loss from omitting records over all canonical DP mechanisms -- Laplace, Gaussian, exponential, and report noisy max -- , as well as recent applications of sampling, such as clustering. Extending this analysis, we investigate suppression as

View details →
Negative / Null Result ReportOpen accessComputer Science

Towards Geo-Culturally Grounded LLM Generations

Piyawat Lertvittayakumjorn, David Kinney, Vinodkumar Prabhakaran et al. · 2025 · arXiv

Generative large language models (LLMs) have demonstrated gaps in diverse cultural awareness across the globe. We investigate the effect of retrieval augmented generation and search-grounding techniques on LLMs' ability to display familiarity with various national cultures. Specifically, we compare the performance of standard LLMs, LLMs augmented with retrievals from a bespoke knowledge base (i.e., KB grounding), and LLMs augmented with retrievals from a web search (i.e., search grounding) on multiple cultural awareness benchmarks. We find that search grounding significantly improves the LLM p

View details →
Negative / Null Result ReportOpen accessComputer Science

Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?

Fırat Öncel, Matthias Bethge, Beyza Ermis et al. · 2024 · arXiv

In the last decade, the generalization and adaptation abilities of deep learning models were typically evaluated on fixed training and test distributions. Contrary to traditional deep learning, large language models (LLMs) are (i) even more overparameterized, (ii) trained on unlabeled text corpora curated from the Internet with minimal human intervention, and (iii) trained in an online fashion. These stark contrasts prevent researchers from transferring lessons learned on model generalization and adaptation in deep learning contexts to LLMs. To this end, our short paper introduces empirical ob

View details →