e-ISSN: Pending
Failure-mode index

Search what already failed

A searchable index of real negative results, null findings, and replication failures from the published literature — so you can learn what didn't work before repeating it.

WASTE indexes published research — it does not host or republish full papers. Each entry is a metadata record (title, authors, DOI) compiled from open scholarly databases, with the abstract shown in full only where the paper is openly licensed (e.g. Creative Commons); otherwise a short excerpt is shown for reference under fair use. WASTE classifies each work by failure type; classifications are automated and approximate.

19820 results in Negative / Null Result Report · page 172 of 661

Negative / Null Result ReportMedicine

Benchmarking and fine-tuning vision-language models on a visual question answering dataset for myopic maculopathy.

Yip, Xu, Liu et al. · 2026 · Asia-Pacific journal of ophthalmology (Philadelphia, Pa.)

To build a visual question answering (VQA) dataset for fine-tuning and evaluating vision-language models (VLMs) in myopic maculopathy (MM). Cross-sectional study. Colour fundus photographs (CFPs) from two publicly available datasets were…

View details →DOI: 10.1016/j.apjo.2026.100345
Negative / Null Result ReportOpen accessComputer Science

APM: Evaluating Style Personalization in LLMs with Arbitrary Preference Mappings

Philipp Spohn, Leander Girrbach, Zeynep Akata · 2026 · arXiv

Typical LLM responses tend to follow a default style, even though users often have distinct preferences regarding tone, verbosity, and formality that they do not explicitly state in their prompts. Evaluating whether personalization methods can adapt to these implicit preferences is challenging, since users typically provide prompts rather than reference responses, style preferences are not factually verifiable, and reference-free LLM judges may conflate personalization with general response quality. To address these challenges, we introduce the Arbitrary Preference Mapping (APM) benchmark, whi

Negative / Null Result ReportOpen accessEngineering

Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing

Mengqi Wang, Zhan Liu, Zengrui Jin et al. · 2025 · arXiv

Diffusion-based large language models (DLLMs) have recently attracted growing interest as an alternative to autoregressive decoders. In this work, we present an empirical study on using the diffusion-based large language model LLaDA for automatic speech recognition (ASR). We first investigate its use as an external deliberation-based processing module for Whisper-LLaMA transcripts. By leveraging the bidirectional attention and denoising capabilities of LLaDA, we explore random masking, low-confidence masking, and semi-autoregressive strategies, showing that Whisper-LLaDA substantially reduces

Negative / Null Result ReportOpen accessComputer Science

Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning

Minwu Kim, Anubhav Shrestha, Safal Shrestha et al. · 2025 · arXiv

Recent studies have shown that reinforcement learning with verifiable rewards (RLVR) enhances overall accuracy (pass@1) but often fails to improve capability (pass@k) of LLMs in reasoning tasks, while distillation can improve both. In this paper, we investigate the mechanisms behind these phenomena. First, we demonstrate that RLVR struggles to improve capability as it focuses on improving the accuracy of the easier questions to the detriment of the accuracy of the most difficult questions. Second, we show that RLVR does not merely increase the success probability for the easier questions, but

Negative / Null Result ReportOpen accessComputer Science

Toward Understanding Adversarial Distillation: Why Robust Teachers Fail

Hongsin Lee, Hye Won Chung · 2026 · arXiv

Adversarial Distillation aims to enhance student robustness by guiding the student with a robust teacher's soft labels within the min-max adversarial training framework, yet its success is notoriously inconsistent: a more robust teacher often fails to improve, or even harms, the student's robust generalization. In this paper, we identify a key mechanism of this teacher dependency: the misalignment between the teacher's supervisory confidence and the student's representational limitations on a consistent subset of training data -- the Robustly Unlearnable Set. We present a theoretical framework

Negative / Null Result ReportOpen accessComputer Science

Why Limit the Residual Stream to Layers and Not Tokens? Persistent Memory for Continuous Latent Reasoning

Mujtaba Farhan, Maheep Chaudhary · 2026 · arXiv

Large language models (LLMs) have demonstrated remarkable reasoning abilities on mathematical and multi-hop planning tasks. The CoCoNuT (Chain of Continuous Thought) paradigm~\cite{hao2024coconut} extends this by enabling models to reason in latent space, exploring multiple reasoning paths simultaneously rather than committing to a single chain early on. However, we identify a limitation we term the \textbf{concept bottleneck}. At each reasoning pass, intermediate hidden states are overwritten, causing the model to lose critical facts computed in earlier steps as reasoning depth increases. We

Negative / Null Result ReportOpen accessComputer Science

CMAX++ : Leveraging Experience in Planning and Execution using Inaccurate Models

Anirudh Vemula, J. Andrew Bagnell, Maxim Likhachev · 2020 · arXiv

Given access to accurate dynamical models, modern planning approaches are effective in computing feasible and optimal plans for repetitive robotic tasks. However, it is difficult to model the true dynamics of the real world before execution, especially for tasks requiring interactions with objects whose parameters are unknown. A recent planning approach, CMAX, tackles this problem by adapting the planner online during execution to bias the resulting plans away from inaccurately modeled regions. CMAX, while being provably guaranteed to reach the goal, requires strong assumptions on the accuracy

Negative / Null Result ReportOpen accessComputer Science

Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision

Yaowen Ye, Cassidy Laidlaw, Jacob Steinhardt · 2025 · arXiv

Language model (LM) post-training relies on two stages of human supervision: task demonstrations for supervised finetuning (SFT), followed by preference comparisons for reinforcement learning from human feedback (RLHF). As LMs become more capable, the tasks they are given become harder to supervise. Will post-training remain effective under unreliable supervision? To test this, we simulate unreliable demonstrations and comparison feedback using small LMs and time-constrained humans. We find that in the presence of unreliable supervision, SFT still retains some effectiveness, but DPO (a common

Negative / Null Result ReportOpen accessComputer Science

HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application

Yiqian Yang, Tian Lan, Qianghuai Jia et al. · 2025 · arXiv

Effective deep search agents must not only access open-domain and domain-specific knowledge but also apply complex rules-such as legal clauses, medical manuals and tariff rules. These rules often feature vague boundaries and implicit logic relationships, making precise application challenging for agents. However, this critical capability is largely overlooked by current agent benchmarks. To fill this gap, we introduce HSCodeComp, the first realistic, expert-level e-commerce benchmark designed to evaluate deep search agents in hierarchical rule application. In this task, the deep reasoning proc

Negative / Null Result ReportOpen accessEngineering

A Unified Approach to Enforce Non-Negativity Constraint in Neural Network Approximation for Optimal Voltage Regulation

Jiaqi Wu, Jingyi Yuan, Yang Weng et al. · 2025 · arXiv

Power system voltage regulation is crucial to maintain power quality while integrating intermittent renewable resources in distribution grids. However, the system model on the grid edge is often unknown, making it difficult to model physical equations for optimal control. Therefore, previous work proposes structured data-driven methods like input convex neural networks (ICNN) for "optimal" control without relying on a physical model. While ICNNs offer theoretical guarantees based on restrictive assumptions of non-negative neural network parameters, can one improve the approximation power with

Negative / Null Result ReportOpen accessComputer Science

Fully-Dynamic All-Pairs Shortest Paths: Likely Optimal Worst-Case Update Time

Xiao Mao · 2023 · arXiv

The All-Pairs Shortest Paths (APSP) problem is one of the fundamental problems in theoretical computer science. It asks to compute the distance matrix of a given $n$-vertex graph. We revisit the classical problem of maintaining the distance matrix under a fully dynamic setting undergoing vertex insertions and deletions with a fast worst-case running time and efficient space usage. Although an algorithm with amortized update-time $\tilde O(n ^ 2)$ has been known for nearly two decades [Demetrescu and Italiano, STOC 2003], the current best algorithm for worst-case running time with efficient spa

Negative / Null Result ReportOpen accessComputer Science

Understanding and Analyzing Model Robustness and Knowledge-Transfer in Multilingual Neural Machine Translation using TX-Ray

Vageesh Saxena, Sharid Loáiciga, Nils Rethmeier · 2024 · arXiv

Neural networks have demonstrated significant advancements in Neural Machine Translation (NMT) compared to conventional phrase-based approaches. However, Multilingual Neural Machine Translation (MNMT) in extremely low-resource settings remains underexplored. This research investigates how knowledge transfer across languages can enhance MNMT in such scenarios. Using the Tatoeba translation challenge dataset from Helsinki NLP, we perform English-German, English-French, and English-Spanish translations, leveraging minimal parallel data to establish cross-lingual mappings. Unlike conventional meth

Negative / Null Result ReportOpen accessEngineering

Learning-based Axial Video Motion Magnification

Kwon Byung-Ki, Oh Hyun-Bin, Kim Jun-Seong et al. · 2023 · arXiv

Video motion magnification amplifies invisible small motions to be perceptible, which provides humans with a spatially dense and holistic understanding of small motions in the scene of interest. This is based on the premise that magnifying small motions enhances the legibility of motions. In the real world, however, vibrating objects often possess convoluted systems that have complex natural frequencies, modes, and directions. Existing motion magnification often fails to improve legibility since the intricate motions still retain complex characteristics even after being magnified, which may di

Negative / Null Result Report

The P-element Has Not Significant Effect on the Drosophila simulans Viability

L. P. Zakharenko, D. V. Petrovskii, R. A. Bykov · 2023 · Молекулярная биология

Cases of horizontal transfer of transposable elements (TEs) between species are known for the Drosophilidae family. In the middle of the last century, the case of horizontal transfer of the P-element from the Drosophila willistoni to the…

View details →DOI: 10.31857/s0026898423020258
Negative / Null Result Report

Effect of the Extracts from Lichens and Lichenophilic Fungi on in vitro Growth of Clinically Significant Microorganisms

T. A. Pankratov, R. E. Shcherbatov, A. A. Del’tsov · 2023 · Микробиология

Abstract—Activity of the ethanol extracts from lichens (LE), of the cultures of lichenophilic (endobiotic) fungi (LFE), and of ethanol extracts from these cultures was tested using the following test organisms: Escherichia coli, Salmonella…

View details →DOI: 10.31857/s002636562360027x
Negative / Null Result ReportMedicine

vNOTES Pelvic Reconstruction and Presacral-Uterosacral Ligament Compound Suspension for Treatment of Multicompartment Pelvic Organ Prolapse: A Three-Year, Three-Arm, Open-Label, Randomized Controlled Trial.

Peng, Wang, Xiao et al. · 2026 · Journal of minimally invasive gynecology

To compare the effectiveness and safety of transvaginal natural orifice transluminal endoscopic surgery (vNOTES) pelvic reconstruction (PRC), vNOTES presacral-uterosacral ligament suspension (PULS), and sacrospinous ligament fixation…

View details →DOI: 10.1016/j.jmig.2026.06.024
Negative / Null Result ReportOpen accessComputer Science

The added value for MRI radiomics and deep-learning for glioblastoma prognostication compared to clinical and molecular information

D. Abler, O. Pusterla, A. Joye-Kühnis et al. · 2025 · arXiv

Background: Radiomics shows promise in characterizing glioblastoma, but its added value over clinical and molecular predictors has yet to be proven. This study assessed the added value of conventional radiomics (CR) and deep learning (DL) MRI radiomics for glioblastoma prognosis ( 6 months survival) on a large multi-center dataset. Methods: After patient selection, our curated dataset gathers 1152 glioblastoma (WHO 2016) patients from five Swiss centers and one public source. It included clinical (age, gender), molecular (MGMT, IDH), and baseline MRI data (T1, T1 contrast, FLAIR, T2)

Negative / Null Result Report

A Case Report of Korean Medicine Treatment, including Jakyakgamcho-tang, in a Patient with Unilateral Vocal Cord Paralysis That Did Not Improve After Injection Laryngoplasty

Ye-seul Park, Ja-eun Kwak, Ho-ryong Yoo et al. · 2024 · The Journal of Internal Korean Medicine

Background: Unilateral vocal cord paralysis can occur due to various causes. Among them, unilateral vocal cord paralysis that takes place after endotracheal intubation is often caused by damage to the left recurrent nerve branch during…

View details →DOI: 10.22246/jikm.2024.45.6.1309
Negative / Null Result Report

Do Polluters Outperform Non-Polluters?

Yigit Atilgan, K. Ozgur Demirtas, A. Doruk Gunaydin · 2024 · SSRN Electronic Journal

This paper investigates whether industrial companies with higher toxic emissions outperform those with lower toxic emissions. First, the highest polluters do not consistently outperform the lowest polluters in an almost stochastic…

View details →DOI: 10.2139/ssrn.4828452
Negative / Null Result ReportOpen accessPhysics

How to generate a significant effective temperature for cold dark matter, from first principles

Patrick McDonald · 2009 · arXiv

I show how to reintroduce velocity dispersion into perturbation theory (PT) calculations of structure in the Universe, i.e., how to go beyond the pressureless fluid approximation, starting from first principles. This addresses a possible deficiency in uses of PT to compute clustering on the weakly non-linear scales that will be critical for probing dark energy. Specifically, I show how to derive a non-negligible value for the (initially tiny) velocity dispersion of dark matter particles, , where δv is the deviation of particle velocities from the local bulk flow. The calculation is essen

Negative / Null Result ReportMedicine

Material fracture/retention rates of occluso-proximal composite resin restorations with polyethylene fibers in primary molars: an interim randomized controlled trial.

Rocha, Dos Anjos, Pinho et al. · 2026 · Journal of dentistry

This 12-month, double-blind, parallel, randomized controlled trial compared the material fracture/retention rates of occluso-proximal direct composite resin restorations associated with polyethylene fiber (CR+PF; Ribbond®) with…

View details →DOI: 10.1016/j.jdent.2026.106863
Negative / Null Result ReportMedicine

Generative Artificial Intelligence for Orthognathic Planning: Patient-Specific 3-Dimensional Jaw Reference Forms Conditioned on the Cranial Base.

Kim, Yang, Kuang et al. · 2026 · Journal of oral and maxillofacial surgery : official journal of the American Association of Oral and Maxillofacial Surgeons

Conventional cephalometry relies on measurements derived from a limited set of landmarks and population-based norms, which may not fully capture individual skeletal form. Although evaluated independently, many cephalometric variables…

View details →DOI: 10.1016/j.joms.2026.06.005
Negative / Null Result ReportMedicine

Using a hand-held pupillometer to quantify pupil constriction as part of the accommodation reflex.

Olson, Perera, Kamal et al. · 2026 · Journal of clinical neuroscience : official journal of the Neurosurgical Society of Australasia

Quantitative pupillometry (QP) provides a reliable measure of the pupillary light reflex. The pupil accommodation reflex (P-AR) occurs when there is a change in visual focus from near to far or from far to near. The ability to measure the…

View details →DOI: 10.1016/j.jocn.2026.112161
Negative / Null Result ReportMedicine

Clinical outcomes of endovascular treatment in patients with large ischemic strokes selected by non-contrast computed tomography within the extended time window.

Jiang, Zhu, Gan et al. · 2026 · European journal of radiology

This study aimed to investigate the effectiveness and safety of endovascular treatment (EVT) in patients with large ischemic strokes assessed by non-contrast CT (NCCT) within the extended time window. This study was a sub-analysis of a…

View details →DOI: 10.1016/j.ejrad.2026.113034
Negative / Null Result ReportMedicine

The effect of post-success handwashing on insight problem solving.

Zhao, Wang · 2026 · Acta psychologica

Handwashing, as a common physical cleansing behavior, not only serves a physiological cleaning function but is also believed to symbolically "wash away" emotional experiences or mental residues. Previous research has primarily focused on…

View details →DOI: 10.1016/j.actpsy.2026.107323
Negative / Null Result ReportMedicine

Does postoperative gabapentinoid prescription reduce chronic opioid use following short-segment lumbar instrumentation?

Karnati, Wu, Kaghazchi et al. · 2026 · Journal of neurosurgery. Spine

The aim of this study was to evaluate the impact of initial postoperative gabapentinoid prescription on chronic postoperative opioid use in patients who underwent short-segment lumbar instrumentation by using a large multicenter electronic…

View details →DOI: 10.3171/2026.2.spine25802
Negative / Null Result ReportMedicine

Evaluation of Large Language Models for Structured Data Extraction From Interstitial Lung Disease Clinical Notes: Comparative Study.

Chen, Maddali, Langlotz et al. · 2026 · Journal of medical Internet research

Most clinically relevant data are in unstructured clinical notes, which are verbose and imprecise, making structured data extraction a costly bottleneck for screening patients for studies or maintaining health care registries. This…

View details →DOI: 10.2196/90547