e-ISSN: Pending
Failure-mode index

Search what already failed

A searchable index of real negative results, null findings, and replication failures from the published literature — so you can learn what didn't work before repeating it.

WASTE indexes published research — it does not host or republish full papers. Each entry is a metadata record (title, authors, DOI) compiled from open scholarly databases, with the abstract shown in full only where the paper is openly licensed (e.g. Creative Commons); otherwise a short excerpt is shown for reference under fair use. WASTE classifies each work by failure type; classifications are automated and approximate.

1896 results for "Mpro" · page 23 of 64

Negative / Null Result ReportOpen accessComputer Science

KNN-LM Does Not Improve Open-ended Text Generation

Shufan Wang, Yixiao Song, Andrew Drozdov et al. · 2023 · arXiv

In this paper, we study the generation quality of interpolation-based retrieval-augmented language models (LMs). These methods, best exemplified by the KNN-LM, interpolate the LM's predicted distribution of the next word with a distribution formed from the most relevant retrievals for a given prefix. While the KNN-LM and related methods yield impressive decreases in perplexity, we discover that they do not exhibit corresponding improvements in open-ended generation quality, as measured by both automatic evaluation metrics (e.g., MAUVE) and human evaluations. Digging deeper, we find that interp

Negative / Null Result Report

Modified string test to improve and confirm by molecular characterization for bacterial identification

Muhammad Dawood Mian, Saadullah Jan Khan, Rehana Rani et al. · 2026 · Access Microbiology

Rapid and reliable identification of bacteria is essential in clinical and environmental microbiology. Gram staining remains a widely used method for preliminary classification; however, it may require additional steps and can be difficult…

View details →DOI: 10.1099/acmi.0.000965.v3
Negative / Null Result Report

Abstract 17917: Is Small Change Significant? Association of Small Differences in Life's Simple 7 and Mortality: the Reasons for Geographic and Racial Differences in Stroke (REGARDS) Cohort

Mary Cushman, Suzanne E Judd, Virginia J Howard et al. · 2011 · Circulation

Background. The AHA 2020 Goal includes improving cardiovascular health, defined using a metric consisting of 7 health factors, Life's Simple 7. A central concept of the goal is that small improvements in behavior / lifestyle factors at the…

View details →DOI: 10.1161/circ.124.suppl_21.a17917
Negative / Null Result ReportOpen accessComputer Science

More Rounds, More Noise: Why Multi-Turn Review Fails to Improve Cross-Context Verification

Song Tae-Eun · 2026 · arXiv

Cross-Context Review (CCR) improves LLM verification by separating production and review into independent sessions. A natural extension is multi-turn review: letting the reviewer ask follow-up questions, receive author responses, and review again. We call this Dynamic Cross-Context Review (D-CCR). In a controlled experiment with 30 artifacts and 150 injected errors, we tested four D-CCR variants against the single-pass CCR baseline. Single-pass CCR (F1 = 0.376) significantly outperformed all multi-turn variants, including D-CCR-2b with question-and-answer exchange (F1 = 0.303, $p < 0.001$, $d

Negative / Null Result ReportOpen accessComputer Science

Haptic human-human interaction does not improve individual visuomotor adaptation

Niek Beckers, Edwin van Asseldonk, Herman van der Kooij · 2020 · arXiv

Haptic interaction between two humans, for example, a physiotherapist assisting a patient regaining the ability to grasp a cup, likely facilitates motor skill acquisition. Haptic human-human interaction has been shown to enhance individual performance improvement in a tracking task with a visuomotor rotation perturbation. These results are remarkable given that haptically assisting or guiding an individual rarely benefits their individual improvement when the assistance is removed. We, therefore, replicated a study that reported that haptic interaction between humans was beneficial for individ

Failed Experiment ReportOpen accessComputer Science

Evaluation and Improvement of Laruelle-Widgrén Inverse Banzhaf Approximation

Frits de Nijs, Daan Wilmer · 2012 · arXiv

The goal of this paper is to critically evaluate a heuristic algorithm for the Inverse Banzhaf Index problem by Laruelle and Widgrén. Few qualitative results are known about the approximation quality of the heuristics for this problem. The intuition behind the operation of this approximation algorithm is analysed and evaluated. We found that the algorithm can not handle general inputs well, and often fails to improve inputs. It is also shown to diverge after only tens of iterations. We present three alternative extensions of the algorithm that do not alter the complexity but can result in up t

Failed Experiment ReportOpen accessPhysics

Self-interaction correction schemes for non-collinear spin-density-functional theory

Nicolas Tancogne-Dejean, Martin Lüders, Carsten A. Ullrich · 2023 · arXiv

We extend some of the well established self-interaction correction (SIC) schemes of density-functional theory to the case of systems with noncollinear magnetism. Our proposed SIC schemes are tested on a set of molecules and metallic clusters in combination with the widely used local spin-density approximation. As expected from the collinear SIC, we show that the averaged-density SIC works well for improving ionization energies but fails to improve more subtle quantities like the dipole moments of polar molecules. We investigate the exchange-correlation magnetic field produced by our extension

Negative / Null Result ReportOpen accessComputer Science

Learning to Learn End-to-End Goal-Oriented Dialog From Related Dialog Tasks

Janarthanan Rajendran, Jonathan K. Kummerfeld, Satinder Singh · 2021 · arXiv

For each goal-oriented dialog task of interest, large amounts of data need to be collected for end-to-end learning of a neural dialog system. Collecting that data is a costly and time-consuming process. Instead, we show that we can use only a small amount of data, supplemented with data from a related dialog task. Naively learning from related data fails to improve performance as the related data can be inconsistent with the target task. We describe a meta-learning based method that selectively learns from the related dialog task data. Our approach leads to significant accuracy improvements in

Negative / Null Result ReportOpen accessComputer Science

Understanding Why Generalized Reweighting Does Not Improve Over ERM

Runtian Zhai, Chen Dan, Zico Kolter et al. · 2022 · arXiv

Empirical risk minimization (ERM) is known in practice to be non-robust to distributional shift where the training and the test distributions are different. A suite of approaches, such as importance weighting, and variants of distributionally robust optimization (DRO), have been proposed to solve this problem. But a line of recent work has empirically shown that these approaches do not significantly improve over ERM in real applications with distribution shift. The goal of this work is to obtain a comprehensive theoretical understanding of this intriguing phenomenon. We first posit the class o

Negative / Null Result Report

Dampak Sosial Ekonomi pada Keluaga Penerima Manfaat (KPM) Program Keluarga Harapan (PKH) Exit Mandiri di Kecamatan Pagelaran Kabuoaten Pringsewu dalam Perspektif The Most Significant Change Technique (MSCt)

Ainun Oktavia Sari, Rahayu Sulistyowati, Ita Prihantika · 2020 · Administrativa: Jurnal Birokrasi, Kebijakan dan Pelayanan Publik

The Conditional Cash Transfer (CCT) is a conditional social cash transfer program that provides assistance to Very Poor Households (RTSM) appointed as participants in the Conditional Cash Transfer program which is related to improving the…

View details →DOI: 10.23960/administrativa.v2i3.51
Negative / Null Result ReportOpen accessMathematics

Does preregistration improve the credibility of research findings?

Mark Rubin · 2020 · arXiv

Preregistration entails researchers registering their planned research hypotheses, methods, and analyses in a time-stamped document before they undertake their data collection and analyses. This document is then made available with the published research report to allow readers to identify discrepancies between what the researchers originally planned to do and what they actually ended up doing. This historical transparency is supposed to facilitate judgments about the credibility of the research findings. The present article provides a critical review of 17 of the reasons behind this argument.

Negative / Null Result ReportOpen accessComputer Science

InsCLR: Improving Instance Retrieval with Self-Supervision

Zelu Deng, Yujie Zhong, Sheng Guo et al. · 2021 · arXiv

This work aims at improving instance retrieval with self-supervision. We find that fine-tuning using the recently developed self-supervised (SSL) learning methods, such as SimCLR and MoCo, fails to improve the performance of instance retrieval. In this work, we identify that the learnt representations for instance retrieval should be invariant to large variations in viewpoint and background etc., whereas self-augmented positives applied by the current SSL methods can not provide strong enough signals for learning robust instance-level representations. To overcome this problem, we propose InsCL

Negative / Null Result ReportOpen accessComputer Science

Does Weighting Improve Matrix Factorization for Recommender Systems?

Alex Ayoub, Samuel Robertson, Dawen Liang et al. · 2025 · arXiv

Matrix factorization is a widely used approach for top-N recommendation and collaborative filtering. When implemented on implicit feedback data (such as clicks), a common heuristic is to upweight the observed interactions. This strategy has been shown to improve performance for certain algorithms. In this paper, we conduct a systematic study of various weighting schemes and matrix factorization algorithms. Somewhat surprisingly, we find that training with unweighted data can perform comparably to, and sometimes outperform, training with weighted data, especially for large models. This observat

Negative / Null Result ReportOpen accessComputer Science

Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQL

Hanbing Liu, Haoyang Li, Xiaokang Zhang et al. · 2025 · arXiv

Direct Preference Optimization (DPO) has proven effective in complex reasoning tasks like math word problems and code generation. However, when applied to Text-to-SQL datasets, it often fails to improve performance and can even degrade it. Our investigation reveals the root cause: unlike math and code tasks, which naturally integrate Chain-of-Thought (CoT) reasoning with DPO, Text-to-SQL datasets typically include only final answers (gold SQL queries) without detailed CoT solutions. By augmenting Text-to-SQL datasets with synthetic CoT solutions, we achieve, for the first time, consistent and

Negative / Null Result ReportOpen accessMathematics

The Poincaré Inequality does not improve with blow-up

Andrea Schioppa · 2015 · arXiv

For each $β>1$ we construct a family $F_β$ of metric measure spaces which is closed under the operation of taking weak-tangents (i.e.~blow-ups), and such that each element of $F_β$ admits a $(1,P)$-Poincaré inequality if and only if $P>β$.

Negative / Null Result ReportOpen accessEngineering

From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology

Haoyang Li, Yuchen Hu, Chen Chen et al. · 2024 · arXiv

Deep neural network (DNN)-based speech enhancement (SE) usually uses conventional activation functions, which lack the expressiveness to capture complex multiscale structures needed for high-fidelity SE. Group-Rational KAN (GR-KAN), a variant of Kolmogorov-Arnold Networks (KAN), retains KAN's expressiveness while improving scalability on complex tasks. We adapt GR-KAN to existing DNN-based SE by replacing dense layers with GR-KAN layers in the time-frequency (T-F) domain MP-SENet and adapting GR-KAN's activations into the 1D CNN layers in the time-domain Demucs. Results on Voicebank-DEMAND sho

Negative / Null Result ReportOpen accessEconomics, Econometrics and Finance

Does Anxiety Improve Economic Decision-Making?

Ian Crawford, Carl-Emil Pless · 2026 · arXiv

We study the associations between everyday economic decision-making quality and people's emotional states. Using high-frequency, highly disaggregated consumer "scanner" data, we show that the cost of poor decision-making is substantial, on average equal to around half of day-to-day consumption budgets. While material circumstances help explain decision-making quality, how people feel about those circumstances is equally important. Contrary to evidence that stress and worry impair performance in settings where distraction is costly, we find these same feelings are associated with improved decis

Negative / Null Result ReportOpen accessComputer Science

Towards Geo-Culturally Grounded LLM Generations

Piyawat Lertvittayakumjorn, David Kinney, Vinodkumar Prabhakaran et al. · 2025 · arXiv

Generative large language models (LLMs) have demonstrated gaps in diverse cultural awareness across the globe. We investigate the effect of retrieval augmented generation and search-grounding techniques on LLMs' ability to display familiarity with various national cultures. Specifically, we compare the performance of standard LLMs, LLMs augmented with retrievals from a bespoke knowledge base (i.e., KB grounding), and LLMs augmented with retrievals from a web search (i.e., search grounding) on multiple cultural awareness benchmarks. We find that search grounding significantly improves the LLM p

Abandoned Hypothesis

How can I find my own voice through my instrument

Julia Casañas Cast · 2020 · Royal Conservatoire Research Portal

Many classical musicians can suffer from tension and nervousness during solo performance. This research looks at how practicing improvisation and creative body movement, as well as creating one’s own performance together with a dancer, can…

View details →DOI: 10.22501/koncon.792184
Negative / Null Result ReportOpen accessComputer Science

Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems

Adam Byerly, Daniel Khashabi · 2024 · arXiv

Self-consistency (SC) improves the performance of large language models (LLMs) across various tasks and domains that involve short content. However, does this support its effectiveness for long-context problems? We challenge the assumption that SC's benefits generalize to long-context settings, where LLMs often struggle with position bias, the systematic over-reliance on specific context regions-which hinders their ability to utilize information effectively from all parts of their context. Through comprehensive experimentation with varying state-of-the-art models, tasks, and SC formulations, w

Negative / Null Result ReportOpen accessComputer Science

SciRepEval: A Multi-Format Benchmark for Scientific Document Representations

Amanpreet Singh, Mike D'Arcy, Arman Cohan et al. · 2022 · arXiv

Learned representations of scientific documents can serve as valuable input features for downstream tasks without further fine-tuning. However, existing benchmarks for evaluating these representations fail to capture the diversity of relevant tasks. In response, we introduce SciRepEval, the first comprehensive benchmark for training and evaluating scientific document representations. It includes 24 challenging and realistic tasks, 8 of which are new, across four formats: classification, regression, ranking and search. We then use this benchmark to study and improve the generalization ability o

Negative / Null Result ReportOpen accessComputer Science

The Persuasion Paradox: When LLM Explanations Fail to Improve Human-AI Team Performance

Ruth Cohen, Lu Feng, Ayala Bloch et al. · 2026 · arXiv

While natural-language explanations from large language models (LLMs) are widely adopted to improve transparency and trust, their impact on objective human-AI team performance remains poorly understood. We identify a Persuasion Paradox: fluent explanations systematically increase user confidence and reliance on AI without reliably improving, and in some cases undermining, task accuracy. Across three controlled human-subject studies spanning abstract visual reasoning (RAVEN matrices) and deductive logical reasoning (LSAT problems), we disentangle the effects of AI predictions and explanations u

Negative / Null Result ReportOpen accessComputer Science

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity

Andrei Liviu Nicolicioiu, Mohammad Pezeshki, Aaron Courville · 2026 · arXiv

On-policy self-distillation achieves strong pass@1 accuracy by using a single model as both teacher and student, with the teacher conditioned on a correct demonstration to provide dense token-level feedback. We show that this could come at a hidden cost: rollout diversity decreases and pass@k curves flatten (i.e., generating more rollouts fails to improve accuracy). We trace this to compounding biases in the design of self-distillation with sampled demonstrations. The teacher scores each student rollout while conditioned on a sampled correct rollout, channeling its feedback through the model's

Negative / Null Result ReportOpen accessComputer Science

When is dataset cartography ineffective? Using training dynamics does not improve robustness against Adversarial SQuAD

Paul K. Mandal · 2025 · arXiv

In this paper, I investigate the effectiveness of dataset cartography for extractive question answering on the SQuAD dataset. I begin by analyzing annotation artifacts in SQuAD and evaluate the impact of two adversarial datasets, AddSent and AddOneSent, on an ELECTRA-small model. Using training dynamics, I partition SQuAD into easy-to-learn, ambiguous, and hard-to-learn subsets. I then compare the performance of models trained on these subsets to those trained on randomly selected samples of equal size. Results show that training on cartography-based subsets does not improve generalization to

Negative / Null Result ReportOpen accessMathematics

Preregistration does not improve the transparent evaluation of severity in Popper's philosophy of science or when deviations are allowed

Mark Rubin · 2024 · arXiv

One justification for preregistering research hypotheses, methods, and analyses is that it improves the transparent evaluation of the severity of hypothesis tests. In this article, I consider two cases in which preregistration does not improve this evaluation. First, I argue that, although preregistration may facilitate the transparent evaluation of severity in Mayo's error statistical philosophy of science, it does not facilitate this evaluation in Popper's theory-centric approach. To illustrate, I show that associated concerns about Type I error rate inflation are only relevant in the error

Negative / Null Result ReportOpen accessComputer Science

Does Interaction Improve Bayesian Reasoning with Visualization?

Ab Mosca, Alvitta Ottley, Remco Chang · 2021 · arXiv

Interaction enables users to navigate large amounts of data effectively, supports cognitive processing, and increases data representation methods. However, there have been few attempts to empirically demonstrate whether adding interaction to a static visualization improves its function beyond popular beliefs. In this paper, we address this gap. We use a classic Bayesian reasoning task as a testbed for evaluating whether allowing users to interact with a static visualization can improve their reasoning. Through two crowdsourced studies, we show that adding interaction to a static Bayesian reaso

Negative / Null Result ReportOpen accessComputer Science

Batch normalization does not improve initialization

Joris Dannemann, Gero Junike · 2025 · arXiv

Batch normalization is one of the most important regularization techniques for neural networks, significantly improving training by centering the layers of the neural network. There have been several attempts to provide a theoretical justification for batch ormalization. Santurkar and Tsipras (2018) [How does batch normalization help optimization? Advances in neural information rocessing systems, 31] claim that batch normalization improves initialization. We provide a counterexample showing that this claim s not true, i.e., batch normalization does not improve initialization.