e-ISSN: Pending

Browse the failure-mode index

9 real negative results, null findings, and replication failures in Computer Science · Replication Failure. Search the index →

WASTE indexes published research — it does not host or republish full papers. Each entry is a metadata record compiled from open scholarly databases; the abstract is shown in full only where the paper is openly licensed, otherwise a short excerpt under fair use. Classifications are automated and approximate.

Replication FailureOpen accessComputer Science

Leakage and the reproducibility crisis in machine-learning-based science

Sayash Kapoor, Arvind Narayanan · 2023 · Patterns

Machine-learning (ML) methods have gained prominence in the quantitative sciences. However, there are many known methodological pitfalls, including data leakage, in ML-based science. We systematically investigate reproducibility issues in ML-based science. Through a survey of literature in fields that have adopted ML methods, we find 17 fields where leakage has been found, collectively affecting 294 papers and, in some cases, leading to wildly overoptimistic conclusions. Based on our survey, we introduce a detailed taxonomy of eight types of leakage, ranging from textbook errors to open resear

View details →
Replication FailureOpen accessComputer Science

Assessing the Effect of Visualizations on Bayesian Reasoning through Crowdsourcing

Luana Micallef, Pierre Dragicevic, Jean‐Daniel Fekete · 2012 · IEEE Transactions on Visualization and Computer Graphics

People have difficulty understanding statistical information and are unaware of their wrong judgments, particularly in Bayesian reasoning. Psychology studies suggest that the way Bayesian problems are represented can impact comprehension,…

View details →
Replication FailureOpen accessComputer Science

The reliability of acceptability judgments across languages

Tal Linzen, Yohei Oseki · 2018 · Glossa a journal of general linguistics

The reliability of acceptability judgments made by individual linguists has often been called into question. Recent large-scale replication studies conducted in response to this criticism have shown that the majority of published English acceptability judgments are robust. We make two observations about these replication studies. First, we raise the concern that English acceptability judgments may be more reliable than judgments in other languages. Second, we argue that it is unnecessary to replicate judgments that illustrate uncontroversial descriptive facts; rather, candidates for replicatio

View details →
Replication FailureOpen accessComputer Science

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate

Xiqi Hao, Zengqing Wu, Yu-Xuan Qiu et al. · 2026 · arXiv

Multi-agent debate (MAD) is a promising strategy for improving LLM reasoning, but when agents converge on a shared answer, it is unclear whether that convergence reflects genuine deliberation or social compliance. We show that the conventional answer flip rate conflates three distinct mechanisms: spontaneous instability, stance-induced conformity, and reasoning-induced persuasion. Our three-source decomposition framework isolates each through controlled counterfactual conditions. In the primary MMLU-Pro setting, 37% of agent-question observations change under self-reflection alone, while robus

View details →
Replication FailureOpen accessComputer Science

Contrastive Predictive Coding Done Right for Mutual Information Estimation

J. Jon Ryu, Pavan Yeddanapudi, Xiangxiang Xu et al. · 2025 · arXiv

The InfoNCE objective, originally introduced for contrastive representation learning, has become a popular choice for mutual information (MI) estimation, despite its indirect connection to MI. In this paper, we demonstrate why InfoNCE should not be regarded as a valid MI estimator, and we introduce a simple modification, which we refer to as InfoNCE-anchor, for accurate MI estimation. Our modification introduces an auxiliary anchor class, enabling consistent density ratio estimation and yielding a plug-in MI estimator with significantly reduced bias. Beyond this, we generalize our framework us

View details →
Replication FailureOpen accessComputer Science

Quantum Corrections to Baryon Properties in Chiral Soliton Models

Frank Meier, Hans Walliser · 1996 · arXiv

We present a procedure to calculate 1-loop graphs in the soliton sector of chiral Lagrangians and use it to calculate quantum corrections to certain baryon observables in Skyrme-type models. Results generally show an improvement over the values obtained in tree approximation except for the case of the axial coupling g_A.

View details →
Replication FailureOpen accessComputer Science

Spin Transfer Torque Driven Coupled Oscillators for Self-Oscillating RF Mixers

Supriyo Maji · 2017 · arXiv

Spin transfer torque oscillators (STOs) based on magnetic tunnel junction (MTJ) devices are emerging as a possible replacement for complementary metal-oxide semiconductors for radio-frequency (RF) signal generation. Advantages include low power consumption, small device area, and large frequency tunability. But such a single device cannot achieve the necessary noise performance for RF applications. It has been reported lately that a network of globally coupled STOs achieves significant improvement in phase noise. The study here is to propose use of such coupled STOs as self-oscillating RF mixe

View details →
Replication FailureOpen accessComputer Science

Chains That See, Answers That Don't: A Multi-Aspect Evaluation Recipe for Forced Chain-of-Thought on Video-MME

Zhichao Fan, Yanhang Li, Zexin Zhuang · 2026 · arXiv

Forced chain-of-thought (CoT) is widely assumed to make vision-language models more reliable on video question answering. We propose a small three-probe evaluation recipe to test that assumption: paired accuracy across direct, CoT, answer-first, and no-video conditions; a counterfactual video-swap diagnostic over the CoT chains; and a four-rung visual-degradation ladder. Each probe is reported under both a strict and a permissive regex scorer, with multiplicity correction over a manuscript-declared primary family. Applied to Qwen2.5-VL on Video-MME subsets, the recipe returns a two-part findin

View details →