e-ISSN: Pending
Failure-mode index

Search what already failed

A searchable index of real negative results, null findings, and replication failures from the published literature — so you can learn what didn't work before repeating it.

WASTE indexes published research — it does not host or republish full papers. Each entry is a metadata record (title, authors, DOI) compiled from open scholarly databases, with the abstract shown in full only where the paper is openly licensed (e.g. Creative Commons); otherwise a short excerpt is shown for reference under fair use. WASTE classifies each work by failure type; classifications are automated and approximate.

1757 results in Negative / Null Result Report for "Mpro" · page 21 of 59

Negative / Null Result ReportOpen accessComputer Science

RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution

Kaiyuan Li, Jing-Cheng Pang, Yang Yu · 2026 · arXiv

Reinforcement learning from verifiable rewards (RLVR) stimulates the thinking processes of large language models (LLMs), substantially enhancing their reasoning abilities on verifiable tasks. It is often assumed that similar gains should transfer to general question answering (GQA), but this assumption has not been thoroughly validated. To assess whether RLVR automatically improves LLM performance on GQA, we propose a Cross-Generation evaluation framework that measures the quality of intermediate reasoning by feeding the generated thinking context into LLMs of varying capabilities. Our evaluat

Negative / Null Result ReportOpen accessComputer Science

I see artifacts: ICA-based EEG artifact removal does not improve deep network decoding across three BCI tasks

Taeho Kang, Yiyu Chen, Christian Wallraven · 2026 · arXiv

In this paper, we conduct a detailed investigation on the effect of independent component (IC)-based noise rejection methods in neural network classifier-based decoding of electroencephalography (EEG) data in different task datasets. We apply a pipeline matrix of two popular different independent component (IC) decomposition methods (Infomax and Adaptive Mixture Independent Component Analysis (AMICA)) with three different component rejection strategies (none, ICLabel, and multiple artifact rejection algorithm [MARA]) on three different EEG datasets (motor imagery, long-term memory formation, a

Negative / Null Result ReportOpen accessComputer Science

When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning

Vasilis Niarchos, Constantinos Papageorgakis, Alexander G. Stapleton et al. · 2026 · arXiv

As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical question emerges: How does the interaction between researchers and agents affect the results? We study this using SCALAR (Structured Critic--Actor Loop for AI Reasoning), an Actor--Critic--Judge pipeline applied to quantum field theory and string theory problems. The Actor proposes solutions, the Critic provides iterative feedback, and an independent Judge evaluates the transcript against reference solutions. We vary the Actor persona, the Critic fee

Negative / Null Result ReportOpen accessComputer Science

Penalizing Confident Predictions on Largely Perturbed Inputs Does Not Improve Out-of-Distribution Generalization in Question Answering

Kazutoshi Shinoda, Saku Sugawara, Akiko Aizawa · 2022 · arXiv

Question answering (QA) models are shown to be insensitive to large perturbations to inputs; that is, they make correct and confident predictions even when given largely perturbed inputs from which humans can not correctly derive answers. In addition, QA models fail to generalize to other domains and adversarial test sets, while humans maintain high accuracy. Based on these observations, we assume that QA models do not use intended features necessary for human reading but rely on spurious features, causing the lack of generalization ability. Therefore, we attempt to answer the question: If the

Negative / Null Result ReportOpen accessComputer Science

Discrimination of two channels by adaptive methods and its application to quantum system

Masahito Hayashi · 2008 · arXiv

The optimal exponential error rate for adaptive discrimination of two channels is discussed. In this problem, adaptive choice of input signal is allowed. This problem is discussed in various settings. It is proved that adaptive choice does not improve the exponential error rate in these settings. These results are applied to quantum state discrimination.

Negative / Null Result ReportOpen accessComputer Science

AFP Algorithm and a Canonical Normal Form for Horn Formulas

Ruhollah Majdoddin · 2014 · arXiv

AFP Algorithm is a learning algorithm for Horn formulas. We show that it does not improve the complexity of AFP Algorithm, if after each negative counterexample more that just one refinements are performed. Moreover, a canonical normal form for Horn formulas is presented, and it is proved that the output formula of AFP Algorithm is in this normal form.

Negative / Null Result ReportOpen accessComputer Science

People readily follow personal advice from AI but it does not improve their well-being

Lennart Luettgau, Vanessa Cheung, Magda Dubois et al. · 2025 · arXiv

People increasingly seek personal advice from large language models (LLMs), yet whether humans follow their advice, and its consequences for their well-being, remains unknown. In a longitudinal randomised controlled trial with a representative UK sample (N = 6,474), we found that up to 79% of participants who had a 20-minute discussion with one of three AI chatbots (GPT-4o, LLama-3.3-70B, Gemini 3 Pro) about health, careers or relationships subsequently reported following its advice. Advice-following remained above 60% even for high-stakes recommendations, suggesting that users only weakly cal

Negative / Null Result ReportOpen accessComputer Science

A New Method for Employing Feedback to Improve Coding Performance

Aaron B. Wagner, Nirmal V. Shende, Yücel Altuğ · 2019 · arXiv

We introduce a novel mechanism, called timid/bold coding, by which feedback can be used to improve coding performance. For a certain class of DMCs, called compound-dispersion channels, we show that timid/bold coding allows for an improved second-order coding rate compared with coding without feedback. For DMCs that are not compound dispersion, we show that feedback does not improve the second-order coding rate. Thus we completely determine the class of DMCs for which feedback improves the second-order coding rate. An upper bound on the second-order coding rate is provided for compound-dispersi

Negative / Null Result ReportOpen accessAgricultural and Biological Sciences

Modeling the mobility of living organisms in heterogeneous landscapes: Does memory improve foraging success?

Denis Boyer, Peter D. Walsh · 2010 · arXiv

Thanks to recent technological advances, it is now possible to track with an unprecedented precision and for long periods of time the movement patterns of many living organisms in their habitat. The increasing amount of data available on single trajectories offers the possibility of understanding how animals move and of testing basic movement models. Random walks have long represented the main description for micro-organisms and have also been useful to understand the foraging behaviour of large animals. Nevertheless, most vertebrates, in particular humans and other primates, rely on sophistic

Negative / Null Result ReportOpen accessComputer Science

Does Editing Improve Answer Quality on Stack Overflow? A Data-Driven Investigation

Saikat Mondal, Chanchal K. Roy · 2025 · arXiv

High-quality answers in technical Q&A platforms like Stack Overflow (SO) are crucial as they directly influence software development practices. Poor-quality answers can introduce inefficiencies, bugs, and security vulnerabilities, and thus increase maintenance costs and technical debt in production software. To improve content quality, SO allows collaborative editing, where users revise answers to enhance clarity, correctness, and formatting. Several studies have examined rejected edits and identified the causes of rejection. However, prior research has not systematically assessed whether acce

Negative / Null Result ReportOpen accessComputer Science

Does Multimodality Improve Recommender Systems as Expected? A Critical Analysis and Future Directions

Hongyu Zhou, Yinan Zhang, Aixin Sun et al. · 2025 · arXiv

Multimodal recommendation systems are increasingly popular for their potential to improve performance by integrating diverse data types. However, the actual benefits of this integration remain unclear, raising questions about when and how it truly enhances recommendations. In this paper, we propose a structured evaluation framework to systematically assess multimodal recommendations across four dimensions: Comparative Efficiency, Recommendation Tasks, Recommendation Stages, and Multimodal Data Integration. We benchmark a set of reproducible multimodal models against strong traditional baseline

Negative / Null Result ReportOpen accessPhysics

Does H2O improve the catalytic activity of Au1-4/MgO towards CO oxidation?

Martin Amft, Natalia V. Skorodumova · 2011 · arXiv

The present density functional theory study addresses the question whether the presence of H2O influences the catalytic activity of small gold clusters, Au1-4/MgO(100), towards the oxidation of carbon monoxide. To this end, we studied the (co-)adsorption of H2O and CO/O2 on these gold clusters. The ground state structures in the presence of all three molecular species, that we found, are Au1O2/MgO and Au2-4CO/MgO with H2O adsorbed on the surface in the proximity of the clusters-molecule complex. In this configuration the catalytic activity of Au1-4/MgO is indifferent to the presence of H2O. We

Negative / Null Result ReportOpen accessComputer Science

The Wisdom of Deliberating AI Crowds: Does Deliberation Improve LLM-Based Forecasting?

Paul Schneider, Amalie Schramm · 2025 · arXiv

Structured deliberation has been found to improve the performance of human forecasters. This study investigates whether a similar intervention, i.e. allowing LLMs to review each other's forecasts before updating, can improve accuracy in large language models (GPT-5, Claude Sonnet 4.5, Gemini Pro 2.5). Using 202 resolved binary questions from the Metaculus Q2 2025 AI Forecasting Tournament, accuracy was assessed across four scenarios: (1) diverse models with distributed information, (2) diverse models with shared information, (3) homogeneous models with distributed information, and (4) homogene

Negative / Null Result ReportOpen accessPhysics

Forget metamaterial: It does not improve sound absorption performance as it claims

Chao Shen, Yu Liu, Tianquan Tang et al. · 2023 · arXiv

The term `sub-wavelength' is commonly used to describe innovative sound-absorbing structures usually labeled as `metamaterials'. Such structures, however, inherently do not bring groundbreaking advancements. This study addresses the limitations imposed by the thickness criterion of Yang et al. by introducing the concept of equivalent mass-spring-damping parameters within the resonator framework. This innovative approach introduces an index of `half-absorption bandwidth' to effectively overcome the thickness restriction. Four practical cases are then presented to correct prevalent misleading co

Negative / Null Result ReportOpen accessComputer Science

Indefinite causal order strategy does not improve the estimation of group action

Masahito Hayashi · 2025 · arXiv

We consider estimation of unknown unitary operation when the set of possible unitary operations is given by a projective unitary representation of a compact group. We show that neither indefinite causal order strategy nor adaptive strategy improves the performance of this estimation when error function satisfies group covariance. That is, the optimal parallel strategy gives the optimal performance even under indefinite causal order strategy and adaptive strategy. To study this problem, we newly introduce the concept of generalized positive operator valued measure (GPOVM), and its convariance c

Negative / Null Result ReportOpen accessComputer Science

What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis

Xirui Li, Ming Li, Tianyi Zhou · 2026 · arXiv

Reinforcement learning (RL) with verifiable rewards has become a standard post-training stage for boosting visual reasoning in vision-language models, yet it remains unclear what capabilities RL actually improves compared with supervised fine-tuning as cold-start initialization (IN). End-to-end benchmark gains conflate multiple factors, making it difficult to attribute improvements to specific skills. To bridge the gap, we propose a Frankenstein-style analysis framework including: (i) functional localization via causal probing; (ii) update characterization via parameter comparison; and (iii) t

Negative / Null Result ReportOpen accessComputer Science

Scale Alone Does not Improve Mechanistic Interpretability in Vision Models

Roland S. Zimmermann, Thomas Klein, Wieland Brendel · 2023 · arXiv

In light of the recent widespread adoption of AI systems, understanding the internal information processing of neural networks has become increasingly critical. Most recently, machine vision has seen remarkable progress by scaling neural networks to unprecedented levels in dataset and model size. We here ask whether this extraordinary increase in scale also positively impacts the field of mechanistic interpretability. In other words, has our understanding of the inner workings of scaled neural networks improved as well? We use a psychophysical paradigm to quantify one form of mechanistic inter

Negative / Null Result ReportOpen accessPhysics

Why material slow light does not improve cavity-enhanced atom detection

B. Megyeri, A. Lampis, G. Harvie et al. · 2017 · arXiv

We discuss the prospects for enhancing absorption and scattering of light from a weakly coupled atom in a high-finesse optical cavity by adding a medium with large, positive group index of refraction. The slow-light effect is known to narrow the cavity transmission spectrum and increase the photon lifetime, but the quality factor of the cavity may not be increased in a metrologically useful sense. Specifically, detection of the weakly coupled atom through either cavity ringdown measurements or the Purcell effect fails to improve with the addition of material slow light. A single-atom model of

Negative / Null Result Report

Modified string test to improve and confirm by molecular characterization for bacterial identification

Muhammad Dawood Mian, Saadullah Jan Khan, Rehana Rani et al. · 2026 · Access Microbiology

Rapid and reliable identification of bacteria is essential in clinical and environmental microbiology. Gram staining remains a widely used method for preliminary classification; however, it may require additional steps and can be difficult…

View details →DOI: 10.1099/acmi.0.000965.v3
Negative / Null Result Report

Abstract 17917: Is Small Change Significant? Association of Small Differences in Life's Simple 7 and Mortality: the Reasons for Geographic and Racial Differences in Stroke (REGARDS) Cohort

Mary Cushman, Suzanne E Judd, Virginia J Howard et al. · 2011 · Circulation

Background. The AHA 2020 Goal includes improving cardiovascular health, defined using a metric consisting of 7 health factors, Life's Simple 7. A central concept of the goal is that small improvements in behavior / lifestyle factors at the…

View details →DOI: 10.1161/circ.124.suppl_21.a17917
Negative / Null Result ReportOpen accessComputer Science

Does Diversity Improve the Test Suite Generation for Mobile Applications?

Thomas Vogel, Chinh Tran, Lars Grunske · 2019 · arXiv

In search-based software engineering we often use popular heuristics with default configurations, which typically lead to suboptimal results, or we perform experiments to identify configurations on a trial-and-error basis, which may lead to better results for a specific problem. To obtain better results while avoiding trial-and-error experiments, a fitness landscape analysis is helpful in understanding the search problem, and making an informed decision about the heuristics. In this paper, we investigate the search problem of test suite generation for mobile applications (apps) using SAPIENZ w

Negative / Null Result ReportOpen accessComputer Science

KNN-LM Does Not Improve Open-ended Text Generation

Shufan Wang, Yixiao Song, Andrew Drozdov et al. · 2023 · arXiv

In this paper, we study the generation quality of interpolation-based retrieval-augmented language models (LMs). These methods, best exemplified by the KNN-LM, interpolate the LM's predicted distribution of the next word with a distribution formed from the most relevant retrievals for a given prefix. While the KNN-LM and related methods yield impressive decreases in perplexity, we discover that they do not exhibit corresponding improvements in open-ended generation quality, as measured by both automatic evaluation metrics (e.g., MAUVE) and human evaluations. Digging deeper, we find that interp

Negative / Null Result Report

Dampak Sosial Ekonomi pada Keluaga Penerima Manfaat (KPM) Program Keluarga Harapan (PKH) Exit Mandiri di Kecamatan Pagelaran Kabuoaten Pringsewu dalam Perspektif The Most Significant Change Technique (MSCt)

Ainun Oktavia Sari, Rahayu Sulistyowati, Ita Prihantika · 2020 · Administrativa: Jurnal Birokrasi, Kebijakan dan Pelayanan Publik

The Conditional Cash Transfer (CCT) is a conditional social cash transfer program that provides assistance to Very Poor Households (RTSM) appointed as participants in the Conditional Cash Transfer program which is related to improving the…

View details →DOI: 10.23960/administrativa.v2i3.51
Negative / Null Result ReportOpen accessComputer Science

Haptic human-human interaction does not improve individual visuomotor adaptation

Niek Beckers, Edwin van Asseldonk, Herman van der Kooij · 2020 · arXiv

Haptic interaction between two humans, for example, a physiotherapist assisting a patient regaining the ability to grasp a cup, likely facilitates motor skill acquisition. Haptic human-human interaction has been shown to enhance individual performance improvement in a tracking task with a visuomotor rotation perturbation. These results are remarkable given that haptically assisting or guiding an individual rarely benefits their individual improvement when the assistance is removed. We, therefore, replicated a study that reported that haptic interaction between humans was beneficial for individ

Negative / Null Result ReportOpen accessComputer Science

Understanding Why Generalized Reweighting Does Not Improve Over ERM

Runtian Zhai, Chen Dan, Zico Kolter et al. · 2022 · arXiv

Empirical risk minimization (ERM) is known in practice to be non-robust to distributional shift where the training and the test distributions are different. A suite of approaches, such as importance weighting, and variants of distributionally robust optimization (DRO), have been proposed to solve this problem. But a line of recent work has empirically shown that these approaches do not significantly improve over ERM in real applications with distribution shift. The goal of this work is to obtain a comprehensive theoretical understanding of this intriguing phenomenon. We first posit the class o

Negative / Null Result ReportOpen accessMathematics

Does preregistration improve the credibility of research findings?

Mark Rubin · 2020 · arXiv

Preregistration entails researchers registering their planned research hypotheses, methods, and analyses in a time-stamped document before they undertake their data collection and analyses. This document is then made available with the published research report to allow readers to identify discrepancies between what the researchers originally planned to do and what they actually ended up doing. This historical transparency is supposed to facilitate judgments about the credibility of the research findings. The present article provides a critical review of 17 of the reasons behind this argument.

Negative / Null Result Report

Improvement in Long‐term Household Food Security among Indiana Households with Children did not Differ between Rural and Urban Counties after a Supplemental Nutrition Assistance Program‐Education Intervention

Rebecca L Rivera, Melissa K Maulding, Angela R Abbott et al. · 2016 · The FASEB Journal

Objective To determine the relationship of rural and urban county household status to long‐term food security among households with children in Indiana after a Supplemental Nutrition Assistance Program‐Education (SNAP‐Ed) intervention.…

View details →DOI: 10.1096/fasebj.30.1_supplement.674.26
Negative / Null Result ReportOpen accessComputer Science

Does Weighting Improve Matrix Factorization for Recommender Systems?

Alex Ayoub, Samuel Robertson, Dawen Liang et al. · 2025 · arXiv

Matrix factorization is a widely used approach for top-N recommendation and collaborative filtering. When implemented on implicit feedback data (such as clicks), a common heuristic is to upweight the observed interactions. This strategy has been shown to improve performance for certain algorithms. In this paper, we conduct a systematic study of various weighting schemes and matrix factorization algorithms. Somewhat surprisingly, we find that training with unweighted data can perform comparably to, and sometimes outperform, training with weighted data, especially for large models. This observat