e-ISSN: Pending
Failure-mode index

Search what already failed

A searchable index of real negative results, null findings, and replication failures from the published literature — so you can learn what didn't work before repeating it.

WASTE indexes published research — it does not host or republish full papers. Each entry is a metadata record (title, authors, DOI) compiled from open scholarly databases, with the abstract shown in full only where the paper is openly licensed (e.g. Creative Commons); otherwise a short excerpt is shown for reference under fair use. WASTE classifies each work by failure type; classifications are automated and approximate.

1759 results in Negative / Null Result Report for "Mpro" · page 22 of 59

Negative / Null Result ReportOpen accessComputer Science

An Analysis of BPE Vocabulary Trimming in Neural Machine Translation

Marco Cognetta, Tatsuya Hiraoka, Naoaki Okazaki et al. · 2024 · arXiv

We explore threshold vocabulary trimming in Byte-Pair Encoding subword tokenization, a postprocessing step that replaces rare subwords with their component subwords. The technique is available in popular tokenization libraries but has not been subjected to rigorous scientific scrutiny. While the removal of rare subwords is suggested as best practice in machine translation implementations, both as a means to reduce model size and for improving model performance through robustness, our experiments indicate that, across a large space of hyperparameter settings, vocabulary trimming fails to improv

Negative / Null Result ReportOpen accessPhysics

Forget metamaterial: It does not improve sound absorption performance as it claims

Chao Shen, Yu Liu, Tianquan Tang et al. · 2023 · arXiv

The term `sub-wavelength' is commonly used to describe innovative sound-absorbing structures usually labeled as `metamaterials'. Such structures, however, inherently do not bring groundbreaking advancements. This study addresses the limitations imposed by the thickness criterion of Yang et al. by introducing the concept of equivalent mass-spring-damping parameters within the resonator framework. This innovative approach introduces an index of `half-absorption bandwidth' to effectively overcome the thickness restriction. Four practical cases are then presented to correct prevalent misleading co

Negative / Null Result ReportOpen accessComputer Science

More Rounds, More Noise: Why Multi-Turn Review Fails to Improve Cross-Context Verification

Song Tae-Eun · 2026 · arXiv

Cross-Context Review (CCR) improves LLM verification by separating production and review into independent sessions. A natural extension is multi-turn review: letting the reviewer ask follow-up questions, receive author responses, and review again. We call this Dynamic Cross-Context Review (D-CCR). In a controlled experiment with 30 artifacts and 150 injected errors, we tested four D-CCR variants against the single-pass CCR baseline. Single-pass CCR (F1 = 0.376) significantly outperformed all multi-turn variants, including D-CCR-2b with question-and-answer exchange (F1 = 0.303, $p < 0.001$, $d

Negative / Null Result ReportOpen accessComputer Science

End-to-End Fidelity Analysis of Quantum Circuit Optimization: From Gate-Level Transformations to Pulse-Level Control

Rylan Malarchick · 2026 · arXiv

We present an analysis of quantum circuit fidelity across the full compilation stack, from high-level gate optimization through pulse-level control. We connect a C++ circuit optimizer to a per-gate Lindblad master-equation fidelity model whose decoherence channels are cross-validated against qiskit-dynamics and whose absolute predictions are benchmarked against execution on real hardware. Across a campaign of 4,452 experiment runs over 371 benchmark circuits, gate cancellation provides the dominant improvement ($d = 1.66$, 72% of circuits improved), while circuit size and pulse duration are th

Negative / Null Result ReportOpen accessComputer Science

Prompting Science Report 3: I'll pay you or I'll kill you -- but will you care?

Lennart Meincke, Ethan Mollick, Lilach Mollick et al. · 2025 · arXiv

This is the third in a series of short reports that seek to help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. In this report, we investigate two commonly held prompting beliefs: a) offering to tip the AI model and b) threatening the AI model. Tipping was a commonly shared tactic for improving AI performance and threats have been endorsed by Google Founder Sergey Brin (All-In, May 2025, 8:20) who observed that 'models tend to do better if you threaten them,' a claim we subject to empirical testing here. We evaluate model p

Negative / Null Result ReportOpen accessComputer Science

Rethinking Neural Width for Alternating Current Optimal Power Flow Proxies

Dhruvi Khandelwal, Anurag Basistha, Ayushi Jolotia et al. · 2026 · arXiv

Deep learning proxies for Alternating Current Optimal Power Flow (ACOPF) lack systematic methods for determining architectural size. This paper conducts a constructive thought experiment to answer a fundamental inquiry: how wide must a neural network be to almost accurately approximate the ACOPF manifold? We introduce a Loss-Guided Neural Densification (LG-ND) algorithm that incrementally discovers necessary capacity by expanding only when the current deep neural network topology fails to improve further. Empirical results across various IEEE systems show that LG-ND achieves performance parity

Negative / Null Result ReportOpen accessComputer Science

Meta-Curriculum Learning for Domain Adaptation in Neural Machine Translation

Runzhe Zhan, Xuebo Liu, Derek F. Wong et al. · 2021 · arXiv

Meta-learning has been sufficiently validated to be beneficial for low-resource neural machine translation (NMT). However, we find that meta-trained NMT fails to improve the translation performance of the domain unseen at the meta-training stage. In this paper, we aim to alleviate this issue by proposing a novel meta-curriculum learning for domain adaptation in NMT. During meta-training, the NMT first learns the similar curricula from each domain to avoid falling into a bad local optimum early, and finally learns the curricula of individualities to improve the model robustness for learning dom

Negative / Null Result ReportOpen accessComputer Science

Indefinite causal order strategy does not improve the estimation of group action

Masahito Hayashi · 2025 · arXiv

We consider estimation of unknown unitary operation when the set of possible unitary operations is given by a projective unitary representation of a compact group. We show that neither indefinite causal order strategy nor adaptive strategy improves the performance of this estimation when error function satisfies group covariance. That is, the optimal parallel strategy gives the optimal performance even under indefinite causal order strategy and adaptive strategy. To study this problem, we newly introduce the concept of generalized positive operator valued measure (GPOVM), and its convariance c

Negative / Null Result ReportOpen accessComputer Science

What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis

Xirui Li, Ming Li, Tianyi Zhou · 2026 · arXiv

Reinforcement learning (RL) with verifiable rewards has become a standard post-training stage for boosting visual reasoning in vision-language models, yet it remains unclear what capabilities RL actually improves compared with supervised fine-tuning as cold-start initialization (IN). End-to-end benchmark gains conflate multiple factors, making it difficult to attribute improvements to specific skills. To bridge the gap, we propose a Frankenstein-style analysis framework including: (i) functional localization via causal probing; (ii) update characterization via parameter comparison; and (iii) t

Negative / Null Result ReportOpen accessEconomics, Econometrics and Finance

Board gender diversity and emissions performance: Insights from panel regressions, machine learning, and explainable AI

Mohammad Hassan Shakil, Arne Johan Pollestad, Khine Kyaw et al. · 2025 · arXiv

With European Union initiatives mandating gender quotas on corporate boards, a key question arises: Is greater board gender diversity (BGD) associated with better emissions performance (EP)? To answer this question, we examine the influence of BGD on EP across a sample of European firms from 2016 to 2022. Using panel regressions, advanced machine learning algorithms, and explainable AI, we reveal a non-linear relationship. Specifically, EP improves with BGD up to an optimal level of approximately 35 %, beyond which further increases in BGD yield no additional improvement in EP. A minimum BGD t

Negative / Null Result ReportOpen accessComputer Science

Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQL

Hanbing Liu, Haoyang Li, Xiaokang Zhang et al. · 2025 · arXiv

Direct Preference Optimization (DPO) has proven effective in complex reasoning tasks like math word problems and code generation. However, when applied to Text-to-SQL datasets, it often fails to improve performance and can even degrade it. Our investigation reveals the root cause: unlike math and code tasks, which naturally integrate Chain-of-Thought (CoT) reasoning with DPO, Text-to-SQL datasets typically include only final answers (gold SQL queries) without detailed CoT solutions. By augmenting Text-to-SQL datasets with synthetic CoT solutions, we achieve, for the first time, consistent and

Negative / Null Result Report

Improvement in Long‐term Household Food Security among Indiana Households with Children did not Differ between Rural and Urban Counties after a Supplemental Nutrition Assistance Program‐Education Intervention

Rebecca L Rivera, Melissa K Maulding, Angela R Abbott et al. · 2016 · The FASEB Journal

Objective To determine the relationship of rural and urban county household status to long‐term food security among households with children in Indiana after a Supplemental Nutrition Assistance Program‐Education (SNAP‐Ed) intervention.…

View details →DOI: 10.1096/fasebj.30.1_supplement.674.26
Negative / Null Result ReportOpen accessComputer Science

Scale Alone Does not Improve Mechanistic Interpretability in Vision Models

Roland S. Zimmermann, Thomas Klein, Wieland Brendel · 2023 · arXiv

In light of the recent widespread adoption of AI systems, understanding the internal information processing of neural networks has become increasingly critical. Most recently, machine vision has seen remarkable progress by scaling neural networks to unprecedented levels in dataset and model size. We here ask whether this extraordinary increase in scale also positively impacts the field of mechanistic interpretability. In other words, has our understanding of the inner workings of scaled neural networks improved as well? We use a psychophysical paradigm to quantify one form of mechanistic inter

Negative / Null Result ReportOpen accessComputer Science

Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?

Fırat Öncel, Matthias Bethge, Beyza Ermis et al. · 2024 · arXiv

In the last decade, the generalization and adaptation abilities of deep learning models were typically evaluated on fixed training and test distributions. Contrary to traditional deep learning, large language models (LLMs) are (i) even more overparameterized, (ii) trained on unlabeled text corpora curated from the Internet with minimal human intervention, and (iii) trained in an online fashion. These stark contrasts prevent researchers from transferring lessons learned on model generalization and adaptation in deep learning contexts to LLMs. To this end, our short paper introduces empirical ob

Negative / Null Result ReportOpen accessPhysics

Why material slow light does not improve cavity-enhanced atom detection

B. Megyeri, A. Lampis, G. Harvie et al. · 2017 · arXiv

We discuss the prospects for enhancing absorption and scattering of light from a weakly coupled atom in a high-finesse optical cavity by adding a medium with large, positive group index of refraction. The slow-light effect is known to narrow the cavity transmission spectrum and increase the photon lifetime, but the quality factor of the cavity may not be increased in a metrologically useful sense. Specifically, detection of the weakly coupled atom through either cavity ringdown measurements or the Purcell effect fails to improve with the addition of material slow light. A single-atom model of

Negative / Null Result ReportOpen accessComputer Science

Depth Adaptive Efficient Visual Autoregressive Modeling

Chunliang Li, Tianze Cao, Sanyuan Zhao · 2026 · arXiv

Visual Autoregressive (VAR) modeling inefficiently applies a fixed computational depth to each position when generating high-resolution images. While existing methods accelerate inference by pruning tokens using frequency maps, their binary hard-pruning approach is fundamentally limited and fails to improve quality even with better frequency estimation. Observing that VAR models possess significant depth redundancy, we propose a paradigm shift from pruning entire tokens to adaptively allocating per-token computational depth. To this end, we introduce DepthVAR, a training-free framework that dy

Negative / Null Result Report

Modified string test to improve and confirm by molecular characterization for bacterial identification

Muhammad Dawood Mian, Saadullah Jan Khan, Rehana Rani et al. · 2026 · Access Microbiology

Rapid and reliable identification of bacteria is essential in clinical and environmental microbiology. Gram staining remains a widely used method for preliminary classification; however, it may require additional steps and can be difficult…

View details →DOI: 10.1099/acmi.0.000965.v3
Negative / Null Result ReportOpen accessComputer Science

The Persuasion Paradox: When LLM Explanations Fail to Improve Human-AI Team Performance

Ruth Cohen, Lu Feng, Ayala Bloch et al. · 2026 · arXiv

While natural-language explanations from large language models (LLMs) are widely adopted to improve transparency and trust, their impact on objective human-AI team performance remains poorly understood. We identify a Persuasion Paradox: fluent explanations systematically increase user confidence and reliance on AI without reliably improving, and in some cases undermining, task accuracy. Across three controlled human-subject studies spanning abstract visual reasoning (RAVEN matrices) and deductive logical reasoning (LSAT problems), we disentangle the effects of AI predictions and explanations u

Negative / Null Result Report

Abstract 17917: Is Small Change Significant? Association of Small Differences in Life's Simple 7 and Mortality: the Reasons for Geographic and Racial Differences in Stroke (REGARDS) Cohort

Mary Cushman, Suzanne E Judd, Virginia J Howard et al. · 2011 · Circulation

Background. The AHA 2020 Goal includes improving cardiovascular health, defined using a metric consisting of 7 health factors, Life's Simple 7. A central concept of the goal is that small improvements in behavior / lifestyle factors at the…

View details →DOI: 10.1161/circ.124.suppl_21.a17917
Negative / Null Result ReportOpen accessComputer Science

SRA-Seg: Synthetic to Real Alignment for Semi-Supervised Medical Image Segmentation

OFM Riaz Rahman Aranya, Kevin Desai · 2026 · arXiv

Synthetic data, an appealing alternative to extensive expert-annotated data for medical image segmentation, consistently fails to improve segmentation performance despite its visual realism. The reason being that synthetic and real medical images exist in different semantic feature spaces, creating a domain gap that current semi-supervised learning methods cannot bridge. We propose SRA-Seg, a framework explicitly designed to align synthetic and real feature distributions for medical image segmentation. SRA-Seg introduces a similarity-alignment (SA) loss using frozen DINOv2 embeddings to pull s

Negative / Null Result ReportOpen accessEngineering

Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation

Sreyan Ghosh, Mohammad Sadegh Rasooli, Michael Levit et al. · 2024 · arXiv

Generative Error Correction (GEC) has emerged as a powerful post-processing method to enhance the performance of Automatic Speech Recognition (ASR) systems. However, we show that GEC models struggle to generalize beyond the specific types of errors encountered during training, limiting their ability to correct new, unseen errors at test time, particularly in out-of-domain (OOD) scenarios. This phenomenon amplifies with named entities (NEs), where, in addition to insufficient contextual information or knowledge about the NEs, novel NEs keep emerging. To address these issues, we propose DARAG (D

Negative / Null Result ReportMedicine

ClassPass Memberships to Improve Well-Being Among Psychiatry Residents: A Pilot Randomized Controlled Trial.

Yang, Mueller, D'Andrea et al. · 2026 · Academic psychiatry : the journal of the American Association of Directors of Psychiatric Residency Training and the Association for Academic Psychiatry

Resident physicians experience high rates of depression, anxiety, burnout,&#xa0;and loneliness, yet few evidence-based interventions have been evaluated in this population. This randomized controlled pilot trial examined the feasibility,…

View details →DOI: 10.1007/s40596-026-02384-y
Negative / Null Result ReportOpen accessComputer Science

Does Diversity Improve the Test Suite Generation for Mobile Applications?

Thomas Vogel, Chinh Tran, Lars Grunske · 2019 · arXiv

In search-based software engineering we often use popular heuristics with default configurations, which typically lead to suboptimal results, or we perform experiments to identify configurations on a trial-and-error basis, which may lead to better results for a specific problem. To obtain better results while avoiding trial-and-error experiments, a fitness landscape analysis is helpful in understanding the search problem, and making an informed decision about the heuristics. In this paper, we investigate the search problem of test suite generation for mobile applications (apps) using SAPIENZ w

Negative / Null Result ReportOpen accessComputer Science

KNN-LM Does Not Improve Open-ended Text Generation

Shufan Wang, Yixiao Song, Andrew Drozdov et al. · 2023 · arXiv

In this paper, we study the generation quality of interpolation-based retrieval-augmented language models (LMs). These methods, best exemplified by the KNN-LM, interpolate the LM's predicted distribution of the next word with a distribution formed from the most relevant retrievals for a given prefix. While the KNN-LM and related methods yield impressive decreases in perplexity, we discover that they do not exhibit corresponding improvements in open-ended generation quality, as measured by both automatic evaluation metrics (e.g., MAUVE) and human evaluations. Digging deeper, we find that interp

Negative / Null Result Report

Dampak Sosial Ekonomi pada Keluaga Penerima Manfaat (KPM) Program Keluarga Harapan (PKH) Exit Mandiri di Kecamatan Pagelaran Kabuoaten Pringsewu dalam Perspektif The Most Significant Change Technique (MSCt)

Ainun Oktavia Sari, Rahayu Sulistyowati, Ita Prihantika · 2020 · Administrativa: Jurnal Birokrasi, Kebijakan dan Pelayanan Publik

The Conditional Cash Transfer (CCT) is a conditional social cash transfer program that provides assistance to Very Poor Households (RTSM) appointed as participants in the Conditional Cash Transfer program which is related to improving the…

View details →DOI: 10.23960/administrativa.v2i3.51
Negative / Null Result ReportOpen accessComputer Science

Improving Policy Optimization with Generalist-Specialist Learning

Zhiwei Jia, Xuanlin Li, Zhan Ling et al. · 2022 · arXiv

Generalization in deep reinforcement learning over unseen environment variations usually requires policy learning over a large set of diverse training variations. We empirically observe that an agent trained on many variations (a generalist) tends to learn faster at the beginning, yet its performance plateaus at a less optimal level for a long time. In contrast, an agent trained only on a few variations (a specialist) can often achieve high returns under a limited computational budget. To have the best of both worlds, we propose a novel generalist-specialist training framework. Specifically, w

Negative / Null Result ReportOpen accessComputer Science

Haptic human-human interaction does not improve individual visuomotor adaptation

Niek Beckers, Edwin van Asseldonk, Herman van der Kooij · 2020 · arXiv

Haptic interaction between two humans, for example, a physiotherapist assisting a patient regaining the ability to grasp a cup, likely facilitates motor skill acquisition. Haptic human-human interaction has been shown to enhance individual performance improvement in a tracking task with a visuomotor rotation perturbation. These results are remarkable given that haptically assisting or guiding an individual rarely benefits their individual improvement when the assistance is removed. We, therefore, replicated a study that reported that haptic interaction between humans was beneficial for individ

Negative / Null Result ReportOpen accessComputer Science

Understanding Why Generalized Reweighting Does Not Improve Over ERM

Runtian Zhai, Chen Dan, Zico Kolter et al. · 2022 · arXiv

Empirical risk minimization (ERM) is known in practice to be non-robust to distributional shift where the training and the test distributions are different. A suite of approaches, such as importance weighting, and variants of distributionally robust optimization (DRO), have been proposed to solve this problem. But a line of recent work has empirically shown that these approaches do not significantly improve over ERM in real applications with distribution shift. The goal of this work is to obtain a comprehensive theoretical understanding of this intriguing phenomenon. We first posit the class o