Replication FailureOpen accessComputer Science
Sayash Kapoor, Arvind Narayanan · 2023 · Patterns
Machine-learning (ML) methods have gained prominence in the quantitative sciences. However, there are many known methodological pitfalls, including data leakage, in ML-based science. We systematically investigate reproducibility issues in ML-based science. Through a survey of literature in fields that have adopted ML methods, we find 17 fields where leakage has been found, collectively affecting 294 papers and, in some cases, leading to wildly overoptimistic conclusions. Based on our survey, we introduce a detailed taxonomy of eight types of leakage, ranging from textbook errors to open resear
View details →Replication FailureOpen accessComputer Science
Luana Micallef, Pierre Dragicevic, Jean‐Daniel Fekete · 2012 · IEEE Transactions on Visualization and Computer Graphics
People have difficulty understanding statistical information and are unaware of their wrong judgments, particularly in Bayesian reasoning. Psychology studies suggest that the way Bayesian problems are represented can impact comprehension,…
View details →Replication FailureOpen accessComputer Science
Yang Yang, Wu Youyou, Brian Uzzi · 2020 · Proceedings of the National Academy of Sciences
Replicability tests of scientific papers show that the majority of papers fail replication. Moreover, failed papers circulate through the literature as quickly as replicating papers. This dynamic weakens the literature, raises research…
View details →Replication FailureOpen accessComputer Science
Tal Linzen, Yohei Oseki · 2018 · Glossa a journal of general linguistics
The reliability of acceptability judgments made by individual linguists has often been called into question. Recent large-scale replication studies conducted in response to this criticism have shown that the majority of published English acceptability judgments are robust. We make two observations about these replication studies. First, we raise the concern that English acceptability judgments may be more reliable than judgments in other languages. Second, we argue that it is unnecessary to replicate judgments that illustrate uncontroversial descriptive facts; rather, candidates for replicatio
View details →Replication FailureOpen accessComputer Science
Xiqi Hao, Zengqing Wu, Yu-Xuan Qiu et al. · 2026 · arXiv
Multi-agent debate (MAD) is a promising strategy for improving LLM reasoning, but when agents converge on a shared answer, it is unclear whether that convergence reflects genuine deliberation or social compliance. We show that the conventional answer flip rate conflates three distinct mechanisms: spontaneous instability, stance-induced conformity, and reasoning-induced persuasion. Our three-source decomposition framework isolates each through controlled counterfactual conditions. In the primary MMLU-Pro setting, 37% of agent-question observations change under self-reflection alone, while robus
View details →Replication FailureOpen accessComputer Science
J. Jon Ryu, Pavan Yeddanapudi, Xiangxiang Xu et al. · 2025 · arXiv
The InfoNCE objective, originally introduced for contrastive representation learning, has become a popular choice for mutual information (MI) estimation, despite its indirect connection to MI. In this paper, we demonstrate why InfoNCE should not be regarded as a valid MI estimator, and we introduce a simple modification, which we refer to as InfoNCE-anchor, for accurate MI estimation. Our modification introduces an auxiliary anchor class, enabling consistent density ratio estimation and yielding a plug-in MI estimator with significantly reduced bias. Beyond this, we generalize our framework us
View details →Replication FailureOpen accessComputer Science
Frank Meier, Hans Walliser · 1996 · arXiv
We present a procedure to calculate 1-loop graphs in the soliton sector of chiral Lagrangians and use it to calculate quantum corrections to certain baryon observables in Skyrme-type models. Results generally show an improvement over the values obtained in tree approximation except for the case of the axial coupling g_A.
View details →Replication FailureOpen accessComputer Science
Supriyo Maji · 2017 · arXiv
Spin transfer torque oscillators (STOs) based on magnetic tunnel junction (MTJ) devices are emerging as a possible replacement for complementary metal-oxide semiconductors for radio-frequency (RF) signal generation. Advantages include low power consumption, small device area, and large frequency tunability. But such a single device cannot achieve the necessary noise performance for RF applications. It has been reported lately that a network of globally coupled STOs achieves significant improvement in phase noise. The study here is to propose use of such coupled STOs as self-oscillating RF mixe
View details →Replication FailureOpen accessComputer Science
Zhichao Fan, Yanhang Li, Zexin Zhuang · 2026 · arXiv
Forced chain-of-thought (CoT) is widely assumed to make vision-language models more reliable on video question answering. We propose a small three-probe evaluation recipe to test that assumption: paired accuracy across direct, CoT, answer-first, and no-video conditions; a counterfactual video-swap diagnostic over the CoT chains; and a four-rung visual-degradation ladder. Each probe is reported under both a strict and a permissive regex scorer, with multiplicity correction over a manuscript-declared primary family. Applied to Qwen2.5-VL on Video-MME subsets, the recipe returns a two-part findin
View details →