e-ISSN: Pending

Browse the failure-mode index

703 real negative results, null findings, and replication failures in Computer Science · Negative / Null Result Report. Search the index →

WASTE indexes published research — it does not host or republish full papers. Each entry is a metadata record compiled from open scholarly databases; the abstract is shown in full only where the paper is openly licensed, otherwise a short excerpt under fair use. Classifications are automated and approximate.

Negative / Null Result ReportOpen accessComputer Science

The Devil is in Fine-tuning and Long-tailed Problems:A New Benchmark for Scene Text Detection

Tianjiao Cao, Jiahao Lyu, Weichao Zeng et al. · 2025 · arXiv

Scene text detection has seen the emergence of high-performing methods that excel on academic benchmarks. However, these detectors often fail to replicate such success in real-world scenarios. We uncover two key factors contributing to this discrepancy through extensive experiments. First, a \textit{Fine-tuning Gap}, where models leverage \textit{Dataset-Specific Optimization} (DSO) paradigm for one domain at the cost of reduced effectiveness in others, leads to inflated performances on academic benchmarks. Second, the suboptimal performance in practical settings is primarily attributed to the

View details →
Negative / Null Result ReportOpen accessComputer Science

Penalizing Confident Predictions on Largely Perturbed Inputs Does Not Improve Out-of-Distribution Generalization in Question Answering

Kazutoshi Shinoda, Saku Sugawara, Akiko Aizawa · 2022 · arXiv

Question answering (QA) models are shown to be insensitive to large perturbations to inputs; that is, they make correct and confident predictions even when given largely perturbed inputs from which humans can not correctly derive answers. In addition, QA models fail to generalize to other domains and adversarial test sets, while humans maintain high accuracy. Based on these observations, we assume that QA models do not use intended features necessary for human reading but rely on spurious features, causing the lack of generalization ability. Therefore, we attempt to answer the question: If the

View details →
Negative / Null Result ReportOpen accessComputer Science

Discrimination of two channels by adaptive methods and its application to quantum system

Masahito Hayashi · 2008 · arXiv

The optimal exponential error rate for adaptive discrimination of two channels is discussed. In this problem, adaptive choice of input signal is allowed. This problem is discussed in various settings. It is proved that adaptive choice does not improve the exponential error rate in these settings. These results are applied to quantum state discrimination.

View details →
Negative / Null Result ReportOpen accessComputer Science

AFP Algorithm and a Canonical Normal Form for Horn Formulas

Ruhollah Majdoddin · 2014 · arXiv

AFP Algorithm is a learning algorithm for Horn formulas. We show that it does not improve the complexity of AFP Algorithm, if after each negative counterexample more that just one refinements are performed. Moreover, a canonical normal form for Horn formulas is presented, and it is proved that the output formula of AFP Algorithm is in this normal form.

View details →
Negative / Null Result ReportOpen accessComputer Science

Nuclear Data Adjustment for Nonlinear Applications in the OECD/NEA WPNCS SG14 Benchmark -- A Bayesian Inverse UQ-based Approach for Data Assimilation

Christopher Brady, Xu Wu · 2025 · arXiv

The Organization for Economic Cooperation and Development (OECD) Working Party on Nuclear Criticality Safety (WPNCS) proposed a benchmark exercise to assess the performance of current nuclear data adjustment techniques applied to nonlinear applications and experiments with low correlation to applications. This work introduces Bayesian Inverse Uncertainty Quantification (IUQ) as a method for nuclear data adjustments in this benchmark, and compares IUQ to the more traditional methods of Generalized Linear Least Squares (GLLS) and Monte Carlo Bayes (MOCABA). Posterior predictions from IUQ showed

View details →
Negative / Null Result ReportOpen accessComputer Science

Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion

Meimingwei Li, Yuanhao Ding, Esteban Garces Arias et al. · 2026 · arXiv

Recent work has identified a counterintuitive phenomenon termed "Hyperfitting", where fine-tuning Large Language Models (LLMs) to near-zero training loss on small datasets surprisingly enhances open-ended generation quality and mitigates repetition in greedy decoding. While effective, the underlying mechanism remains poorly understood, with the extremely low-entropy output distributions suggesting a potential equivalence to simple temperature scaling. In this work, we demonstrate that this phenomenon is fundamentally distinct from distribution sharpening; entropy-matched control experiments re

View details →
Negative / Null Result ReportOpen accessComputer Science

Beyond Geometry: Artistic Disparity Synthesis for Immersive 2D-to-3D

Ping Chen, Zezhou Chen, Xingpeng Zhang et al. · 2026 · arXiv

Current 2D-to-3D conversion methods achieve geometric accuracy but are artistically deficient, failing to replicate the immersive and emotionally resonant experience of professional 3D cinema. This is because geometric reconstruction paradigms mistake deliberate artistic intent, such as strategic zero-plane shifts for pop-out effects and local depth sculpting, for data noise or ambiguity. This paper argues for a new paradigm: Artistic Disparity Synthesis, shifting the goal from physically accurate disparity estimation to artistically coherent disparity synthesis. We propose Art3D, a preliminar

View details →
Negative / Null Result ReportOpen accessComputer Science

People readily follow personal advice from AI but it does not improve their well-being

Lennart Luettgau, Vanessa Cheung, Magda Dubois et al. · 2025 · arXiv

People increasingly seek personal advice from large language models (LLMs), yet whether humans follow their advice, and its consequences for their well-being, remains unknown. In a longitudinal randomised controlled trial with a representative UK sample (N = 6,474), we found that up to 79% of participants who had a 20-minute discussion with one of three AI chatbots (GPT-4o, LLama-3.3-70B, Gemini 3 Pro) about health, careers or relationships subsequently reported following its advice. Advice-following remained above 60% even for high-stakes recommendations, suggesting that users only weakly cal

View details →
Negative / Null Result ReportOpen accessComputer Science

A New Method for Employing Feedback to Improve Coding Performance

Aaron B. Wagner, Nirmal V. Shende, Yücel Altuğ · 2019 · arXiv

We introduce a novel mechanism, called timid/bold coding, by which feedback can be used to improve coding performance. For a certain class of DMCs, called compound-dispersion channels, we show that timid/bold coding allows for an improved second-order coding rate compared with coding without feedback. For DMCs that are not compound dispersion, we show that feedback does not improve the second-order coding rate. Thus we completely determine the class of DMCs for which feedback improves the second-order coding rate. An upper bound on the second-order coding rate is provided for compound-dispersi

View details →
Negative / Null Result ReportOpen accessComputer Science

PUB: An LLM-Enhanced Personality-Driven User Behaviour Simulator for Recommender System Evaluation

Chenglong Ma, Ziqi Xu, Yongli Ren et al. · 2025 · arXiv

Traditional offline evaluation methods for recommender systems struggle to capture the complexity of modern platforms due to sparse behavioural signals, noisy data, and limited modelling of user personality traits. While simulation frameworks can generate synthetic data to address these gaps, existing methods fail to replicate behavioural diversity, limiting their effectiveness. To overcome these challenges, we propose the Personality-driven User Behaviour Simulator (PUB), an LLM-based simulation framework that integrates the Big Five personality traits to model personalised user behaviour. PU

View details →
Negative / Null Result ReportOpen accessComputer Science

Does Editing Improve Answer Quality on Stack Overflow? A Data-Driven Investigation

Saikat Mondal, Chanchal K. Roy · 2025 · arXiv

High-quality answers in technical Q&A platforms like Stack Overflow (SO) are crucial as they directly influence software development practices. Poor-quality answers can introduce inefficiencies, bugs, and security vulnerabilities, and thus increase maintenance costs and technical debt in production software. To improve content quality, SO allows collaborative editing, where users revise answers to enhance clarity, correctness, and formatting. Several studies have examined rejected edits and identified the causes of rejection. However, prior research has not systematically assessed whether acce

View details →
Negative / Null Result ReportOpen accessComputer Science

When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning

Vasilis Niarchos, Constantinos Papageorgakis, Alexander G. Stapleton et al. · 2026 · arXiv

As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical question emerges: How does the interaction between researchers and agents affect the results? We study this using SCALAR (Structured Critic--Actor Loop for AI Reasoning), an Actor--Critic--Judge pipeline applied to quantum field theory and string theory problems. The Actor proposes solutions, the Critic provides iterative feedback, and an independent Judge evaluates the transcript against reference solutions. We vary the Actor persona, the Critic fee

View details →
Negative / Null Result ReportOpen accessComputer Science

Does Multimodality Improve Recommender Systems as Expected? A Critical Analysis and Future Directions

Hongyu Zhou, Yinan Zhang, Aixin Sun et al. · 2025 · arXiv

Multimodal recommendation systems are increasingly popular for their potential to improve performance by integrating diverse data types. However, the actual benefits of this integration remain unclear, raising questions about when and how it truly enhances recommendations. In this paper, we propose a structured evaluation framework to systematically assess multimodal recommendations across four dimensions: Comparative Efficiency, Recommendation Tasks, Recommendation Stages, and Multimodal Data Integration. We benchmark a set of reproducible multimodal models against strong traditional baseline

View details →
Negative / Null Result ReportOpen accessComputer Science

Quantifying task-relevant representational similarity using decision variable correlation

Yu Eric Qian, Wilson S. Geisler, Xue-Xin Wei · 2025 · arXiv

Previous studies have compared neural activities in the visual cortex to representations in deep neural networks trained on image classification. Interestingly, while some suggest that their representations are highly similar, others argued the opposite. Here, we propose a new approach to characterize the similarity of the decision strategies of two observers (models or brains) using decision variable correlation (DVC). DVC quantifies the image-by-image correlation between the decoded decisions based on the internal neural representations in a classification task. Thus, it can capture task-rel

View details →
Negative / Null Result ReportOpen accessComputer Science

Disturbance-Injected Robust Imitation Learning with Task Achievement

Hirotaka Tahara, Hikaru Sasaki, Hanbit Oh et al. · 2022 · arXiv

Robust imitation learning using disturbance injections overcomes issues of limited variation in demonstrations. However, these methods assume demonstrations are optimal, and that policy stabilization can be learned via simple augmentations. In real-world scenarios, demonstrations are often of diverse-quality, and disturbance injection instead learns sub-optimal policies that fail to replicate desired behavior. To address this issue, this paper proposes a novel imitation learning framework that combines both policy robustification and optimal demonstration learning. Specifically, this combinato

View details →
Negative / Null Result ReportOpen accessComputer Science

A wearable anti-gravity supplement to therapy does not improve arm function in chronic stroke: a randomized trial

Courtney Celian, Partha Ryali, Valentino Wilson et al. · 2024 · arXiv

Background: Gravity confounds arm movement ability in post-stroke hemiparesis. Reducing its influence allows effective practice leading to recovery. Yet, there is a scarcity of wearable devices suitable for personalized use across diverse therapeutic activities in the clinic. Objective: In this study, we investigated the safety, feasibility, and efficacy of anti-gravity therapy using the ExoNET device in post-stroke participants. Methods: Twenty chronic stroke survivors underwent six, 45-minute occupational therapy sessions while wearing the ExoNET, randomized into either the treatment (ExoNET

View details →
Negative / Null Result ReportOpen accessComputer Science

Dispersion of Klauder's temporally stable coherent states for the hydrogen atom

Paolo Bellomo, C. R. Stroud, · 1998 · arXiv

We study the dispersion of the "temporally stable" coherent states for the hydrogen atom introduced by Klauder. These are states which under temporal evolution by the hydrogen atom Hamiltonian retain their coherence properties. We show that in the hydrogen atom such wave packets do not move quasi-classically; i.e., they do not follow with no or little dispersion the Keplerian orbits of the classical electron. The poor quantum-classical correspondence does not improve in the semiclassical limit.

View details →
Negative / Null Result ReportOpen accessComputer Science

The Wisdom of Deliberating AI Crowds: Does Deliberation Improve LLM-Based Forecasting?

Paul Schneider, Amalie Schramm · 2025 · arXiv

Structured deliberation has been found to improve the performance of human forecasters. This study investigates whether a similar intervention, i.e. allowing LLMs to review each other's forecasts before updating, can improve accuracy in large language models (GPT-5, Claude Sonnet 4.5, Gemini Pro 2.5). Using 202 resolved binary questions from the Metaculus Q2 2025 AI Forecasting Tournament, accuracy was assessed across four scenarios: (1) diverse models with distributed information, (2) diverse models with shared information, (3) homogeneous models with distributed information, and (4) homogene

View details →
Negative / Null Result ReportOpen accessComputer Science

R^2 Dark Matter

Jose A. R. Cembranos · 2010 · arXiv

There is a non-trivial four-derivative extension of the gravitational spectrum that is free of ghosts and phenomenologically viable. It is the so called $R^2$-gravity since it is defined by the only addition of a term proportional to the square of the scalar curvature. Just the presence of this term does not improve the ultraviolet behaviour of Einstein gravity but introduces one additional scalar degree of freedom that can account for the dark matter of our Universe.

View details →
Negative / Null Result ReportOpen accessComputer Science

Indefinite causal order strategy does not improve the estimation of group action

Masahito Hayashi · 2025 · arXiv

We consider estimation of unknown unitary operation when the set of possible unitary operations is given by a projective unitary representation of a compact group. We show that neither indefinite causal order strategy nor adaptive strategy improves the performance of this estimation when error function satisfies group covariance. That is, the optimal parallel strategy gives the optimal performance even under indefinite causal order strategy and adaptive strategy. To study this problem, we newly introduce the concept of generalized positive operator valued measure (GPOVM), and its convariance c

View details →
Negative / Null Result ReportOpen accessComputer Science

Magnetic Moments of the Octet Baryons in the Colour-Dielectric Model

Moo-Sung Bae, Judith A. McGovern · 1995 · arXiv

Baryon magnetic moments are calculated in the colour-dielectric model with pion and kaon loops. The only free parameter of the model is determined from the nucleon isoscalar radius, and all SU(3) symmetry breaking, including that in the quark sector, is determined by mesonic masses and decay constants. Good agreement with experiment is obtained for the ratios of the magnetic moments, but the inclusion of kaons does not improve the results. The results obtained in this approach are significantly better than any that have been obtained in hedgehog-based models.

View details →
Negative / Null Result ReportOpen accessComputer Science

Fusion of Graph Neural Networks via Optimal Transport

Weronika Ormaniec, Michael Vollenweider, Elisa Hoskovec · 2025 · arXiv

In this paper, we explore the idea of combining GCNs into one model. To that end, we align the weights of different models layer-wise using optimal transport (OT). We present and evaluate three types of transportation costs and show that the studied fusion method consistently outperforms the performance of vanilla averaging. Finally, we present results suggesting that model fusion using OT is harder in the case of GCNs than MLPs and that incorporating the graph structure into the process does not improve the performance of the method.

View details →
Negative / Null Result ReportOpen accessComputer Science

Can long-range interactions stabilize quantum memory at nonzero temperature?

Olivier Landon-Cardinal, Beni Yoshida, David Poulin et al. · 2015 · arXiv

A two-dimensional topologically ordered quantum memory is well protected against error if the energy gap is large compared to the temperature, but this protection does not improve as the system size increases. We review and critique some recent proposals for improving the memory time by introducing long-range interactions among anyons, noting that instability with respect to small local perturbations of the Hamiltonian is a generic problem for such proposals. We also discuss some broader issues regarding the prospects for scalable quantum memory in two-dimensional systems.

View details →
Negative / Null Result ReportOpen accessComputer Science

What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis

Xirui Li, Ming Li, Tianyi Zhou · 2026 · arXiv

Reinforcement learning (RL) with verifiable rewards has become a standard post-training stage for boosting visual reasoning in vision-language models, yet it remains unclear what capabilities RL actually improves compared with supervised fine-tuning as cold-start initialization (IN). End-to-end benchmark gains conflate multiple factors, making it difficult to attribute improvements to specific skills. To bridge the gap, we propose a Frankenstein-style analysis framework including: (i) functional localization via causal probing; (ii) update characterization via parameter comparison; and (iii) t

View details →
Negative / Null Result ReportOpen accessComputer Science

Neutral evolution and turnover over centuries of English word popularity

Damian Ruck, R. Alexander Bentley, Alberto Acerbi et al. · 2017 · arXiv

Here we test Neutral models against the evolution of English word frequency and vocabulary at the population scale, as recorded in annual word frequencies from three centuries of English language books. Against these data, we test both static and dynamic predictions of two neutral models, including the relation between corpus size and vocabulary size, frequency distributions, and turnover within those frequency distributions. Although a commonly used Neutral model fails to replicate all these emergent properties at once, we find that modified two-stage Neutral model does replicate the static a

View details →
Negative / Null Result ReportOpen accessComputer Science

Interview-Informed Generative Agents for Product Discovery: A Validation Study

Zichao Wang, Alexa Siu · 2026 · arXiv

Large language models (LLMs) have shown strong performance on standardized social science instruments, but their value for product discovery remains unclear. We investigate whether interview-informed generative agents can simulate user responses in concept testing scenarios. Using in-depth workflow interviews with knowledge workers, we created personalized agents and compared their evaluations of novel AI concepts against the same participants' responses. Our results show that agents are distribution-calibrated but identity-imprecise: they fail to replicate the specific individual they are gro

View details →
Negative / Null Result ReportOpen accessComputer Science

Generalization of short coherent control pulses: extension to arbitrary rotations

S. Pasini, G. S. Uhrig · 2008 · arXiv

We generalize the problem of the coherent control of small quantum systems to the case where the quantum bit (qubit) is subject to a fully general rotation. Following the ideas developed in Pasini et al (2008 Phys. Rev. A 77, 032315), the systematic expansion in the shortness of the pulse is extended to the case where the pulse acts on the qubit as a general rotation around an axis of rotation varying in time. The leading and the next-leading corrections are computed. For certain pulses we prove that the general rotation does not improve on the simpler rotation with fixed axis.

View details →
Negative / Null Result ReportOpen accessComputer Science

Scale Alone Does not Improve Mechanistic Interpretability in Vision Models

Roland S. Zimmermann, Thomas Klein, Wieland Brendel · 2023 · arXiv

In light of the recent widespread adoption of AI systems, understanding the internal information processing of neural networks has become increasingly critical. Most recently, machine vision has seen remarkable progress by scaling neural networks to unprecedented levels in dataset and model size. We here ask whether this extraordinary increase in scale also positively impacts the field of mechanistic interpretability. In other words, has our understanding of the inner workings of scaled neural networks improved as well? We use a psychophysical paradigm to quantify one form of mechanistic inter

View details →
Negative / Null Result ReportOpen accessComputer Science

Does Diversity Improve the Test Suite Generation for Mobile Applications?

Thomas Vogel, Chinh Tran, Lars Grunske · 2019 · arXiv

In search-based software engineering we often use popular heuristics with default configurations, which typically lead to suboptimal results, or we perform experiments to identify configurations on a trial-and-error basis, which may lead to better results for a specific problem. To obtain better results while avoiding trial-and-error experiments, a fitness landscape analysis is helpful in understanding the search problem, and making an informed decision about the heuristics. In this paper, we investigate the search problem of test suite generation for mobile applications (apps) using SAPIENZ w

View details →