e-ISSN: Pending

Browse the failure-mode index

19,915 real negative results, null findings, and replication failures · Negative / Null Result Report. Search the index →

WASTE indexes published research — it does not host or republish full papers. Each entry is a metadata record compiled from open scholarly databases; the abstract is shown in full only where the paper is openly licensed, otherwise a short excerpt under fair use. Classifications are automated and approximate.

Negative / Null Result ReportOpen accessComputer Science

The Devil is in Fine-tuning and Long-tailed Problems:A New Benchmark for Scene Text Detection

Tianjiao Cao, Jiahao Lyu, Weichao Zeng et al. · 2025 · arXiv

Scene text detection has seen the emergence of high-performing methods that excel on academic benchmarks. However, these detectors often fail to replicate such success in real-world scenarios. We uncover two key factors contributing to this discrepancy through extensive experiments. First, a \textit{Fine-tuning Gap}, where models leverage \textit{Dataset-Specific Optimization} (DSO) paradigm for one domain at the cost of reduced effectiveness in others, leads to inflated performances on academic benchmarks. Second, the suboptimal performance in practical settings is primarily attributed to the

View details →
Negative / Null Result ReportOpen accessComputer Science

Nuclear Data Adjustment for Nonlinear Applications in the OECD/NEA WPNCS SG14 Benchmark -- A Bayesian Inverse UQ-based Approach for Data Assimilation

Christopher Brady, Xu Wu · 2025 · arXiv

The Organization for Economic Cooperation and Development (OECD) Working Party on Nuclear Criticality Safety (WPNCS) proposed a benchmark exercise to assess the performance of current nuclear data adjustment techniques applied to nonlinear applications and experiments with low correlation to applications. This work introduces Bayesian Inverse Uncertainty Quantification (IUQ) as a method for nuclear data adjustments in this benchmark, and compares IUQ to the more traditional methods of Generalized Linear Least Squares (GLLS) and Monte Carlo Bayes (MOCABA). Posterior predictions from IUQ showed

View details →
Negative / Null Result Report

Between Concrete and Earth

Yu-Han Huang · 2024 · Roadsides

This essay scrutinizes the soil-cement brick (SCB), a half-earthen, half-concrete building material, and its use in U.S.-aided housing projects in Cold War-era Taiwan. Made of cement and natural earth with manually operated ‘brickmaker’…

View details →
Negative / Null Result ReportOpen accessComputer Science

Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion

Meimingwei Li, Yuanhao Ding, Esteban Garces Arias et al. · 2026 · arXiv

Recent work has identified a counterintuitive phenomenon termed "Hyperfitting", where fine-tuning Large Language Models (LLMs) to near-zero training loss on small datasets surprisingly enhances open-ended generation quality and mitigates repetition in greedy decoding. While effective, the underlying mechanism remains poorly understood, with the extremely low-entropy output distributions suggesting a potential equivalence to simple temperature scaling. In this work, we demonstrate that this phenomenon is fundamentally distinct from distribution sharpening; entropy-matched control experiments re

View details →
Negative / Null Result ReportMedicine

Effect of build orientation, sintering schedule, and resin cement on the bond strength of additively manufactured zirconia to dentin.

Sabatini, Molinero-Mourelle, Limones et al. · 2026 · The Journal of prosthetic dentistry

Whether the manufacturing protocol (build orientation and sintering schedule) of additively manufactured (AM) zirconia impacts the bond strength compared with subtractively manufactured (SM) zirconia remains unclear. The purpose of this in…

View details →
Negative / Null Result ReportOpen accessComputer Science

Beyond Geometry: Artistic Disparity Synthesis for Immersive 2D-to-3D

Ping Chen, Zezhou Chen, Xingpeng Zhang et al. · 2026 · arXiv

Current 2D-to-3D conversion methods achieve geometric accuracy but are artistically deficient, failing to replicate the immersive and emotionally resonant experience of professional 3D cinema. This is because geometric reconstruction paradigms mistake deliberate artistic intent, such as strategic zero-plane shifts for pop-out effects and local depth sculpting, for data noise or ambiguity. This paper argues for a new paradigm: Artistic Disparity Synthesis, shifting the goal from physically accurate disparity estimation to artistically coherent disparity synthesis. We propose Art3D, a preliminar

View details →
Negative / Null Result ReportOpen accessAgricultural and Biological Sciences

Emergence of Human-Like Attention in Self-Supervised Vision Transformers: an eye-tracking study

Takuto Yamamoto, Hirosato Akahoshi, Shigeru Kitazawa · 2024 · arXiv

Many models of visual attention have been proposed so far. Traditional bottom-up models, like saliency models, fail to replicate human gaze patterns, and deep gaze prediction models lack biological plausibility due to their reliance on supervised learning. Vision Transformers (ViTs), with their self-attention mechanisms, offer a new approach but often produce dispersed attention patterns if trained with supervised learning. This study explores whether self-supervised DINO (self-DIstillation with NO labels) training enables ViTs to develop attention mechanisms resembling human visual attention.

View details →
Negative / Null Result Report

Evaluating LSP-TEOC.Pro: What we did and what we found out

Joseph Cullen, Gregory Holloway · 2025 · Glottodidactica

This article presents an evaluation of the LSP-TEOC.Pro project. It sets out the evaluation methodology applied, how it was implemented and the key evaluation findings. Given the exploratory nature of the project, the range and complexity…

View details →
Negative / Null Result ReportOpen accessComputer Science

PUB: An LLM-Enhanced Personality-Driven User Behaviour Simulator for Recommender System Evaluation

Chenglong Ma, Ziqi Xu, Yongli Ren et al. · 2025 · arXiv

Traditional offline evaluation methods for recommender systems struggle to capture the complexity of modern platforms due to sparse behavioural signals, noisy data, and limited modelling of user personality traits. While simulation frameworks can generate synthetic data to address these gaps, existing methods fail to replicate behavioural diversity, limiting their effectiveness. To overcome these challenges, we propose the Personality-driven User Behaviour Simulator (PUB), an LLM-based simulation framework that integrates the Big Five personality traits to model personalised user behaviour. PU

View details →
Negative / Null Result ReportOpen accessComputer Science

Disturbance-Injected Robust Imitation Learning with Task Achievement

Hirotaka Tahara, Hikaru Sasaki, Hanbit Oh et al. · 2022 · arXiv

Robust imitation learning using disturbance injections overcomes issues of limited variation in demonstrations. However, these methods assume demonstrations are optimal, and that policy stabilization can be learned via simple augmentations. In real-world scenarios, demonstrations are often of diverse-quality, and disturbance injection instead learns sub-optimal policies that fail to replicate desired behavior. To address this issue, this paper proposes a novel imitation learning framework that combines both policy robustification and optimal demonstration learning. Specifically, this combinato

View details →
Negative / Null Result Report

The ability of VBNC F. tularensis to replicate within THP-1 cells.

Carly Cunningham, Stuart Cantlay, Joseph Horzempa · 2019 · Proceedings of the West Virginia Academy of Science

CARLY CUNNINGHAM, STUART CANTLAY, AND JOSEPH HORZEMPA, Department of Natural Sciences and Mathematics, West Liberty University, West Liberty, WV USA. The ability of VBNC F. tularensis to replicate within THP-1 cells. Francisella…

View details →
Negative / Null Result ReportOpen accessPhysics

Herd Behaviour in Public Goods Games

María Pereda · 2024 · arXiv

The problem of free-riding arises when individuals benefit from a shared resource, service, or public good without contributing proportionately to its provision. This conduct often leads to a collective action problem, as individuals pursue personal gains while relying on the contributions of others. In this study, we present a Bayesian inference model to elucidate the behaviour of participants in a Public Goods Game, a conceptual framework that captures the essence of the free-riding problem. Here, individuals possess information on the distribution of group donations to the public good. Our

View details →
Negative / Null Result ReportOpen accessComputer Science

Neutral evolution and turnover over centuries of English word popularity

Damian Ruck, R. Alexander Bentley, Alberto Acerbi et al. · 2017 · arXiv

Here we test Neutral models against the evolution of English word frequency and vocabulary at the population scale, as recorded in annual word frequencies from three centuries of English language books. Against these data, we test both static and dynamic predictions of two neutral models, including the relation between corpus size and vocabulary size, frequency distributions, and turnover within those frequency distributions. Although a commonly used Neutral model fails to replicate all these emergent properties at once, we find that modified two-stage Neutral model does replicate the static a

View details →
Negative / Null Result ReportOpen accessMathematics

Causal Stability Selection

Falco J. Bargagli-Stoffi, Omar Melikechi · 2026 · arXiv

Identifying covariates that modify treatment effects is a central problem in causal inference. Yet existing data-adaptive procedures do not provide finite-sample control over the expected number of false discoveries, risking spurious findings that fail to replicate. We introduce causal stability selection, an algorithm that combines cross-fitted estimation of conditional average treatment effects with integrated path stability selection. The method accommodates arbitrary treatment effect estimators and arbitrary base selectors, and produces a selection set with an explicit, non-asymptotic boun

View details →
Negative / Null Result Report

Failed execution in inferential pragmatics

Francesco Mercadante · 2026 · Pragmatics and Society

Abstract This study proposes a redefinition of the pragmatic foundations of linguistic communication, based on the notion of failed execution as an autonomous analytical category. Drawing on qualitative analysis of a curated illustrative…

View details →
Negative / Null Result ReportOpen accessComputer Science

Interview-Informed Generative Agents for Product Discovery: A Validation Study

Zichao Wang, Alexa Siu · 2026 · arXiv

Large language models (LLMs) have shown strong performance on standardized social science instruments, but their value for product discovery remains unclear. We investigate whether interview-informed generative agents can simulate user responses in concept testing scenarios. Using in-depth workflow interviews with knowledge workers, we created personalized agents and compared their evaluations of novel AI concepts against the same participants' responses. Our results show that agents are distribution-calibrated but identity-imprecise: they fail to replicate the specific individual they are gro

View details →
Negative / Null Result ReportOpen accessComputer Science

Somatic in the East, Psychological in the West?: Investigating Clinically-Grounded Cross-Cultural Depression Symptom Expression in LLMs

Shintaro Sakai, Jisun An, Migyeong Kang et al. · 2025 · arXiv

Prior clinical psychology research shows that Western individuals with depression tend to report psychological symptoms, while Eastern individuals report somatic ones. We test whether Large Language Models (LLMs), which are increasingly used in mental health, reproduce these cultural patterns by prompting them with Western or Eastern personas. Results show that LLMs largely fail to replicate the patterns when prompted in English, though prompting in major Eastern languages (i.e., Chinese, Japanese, and Hindi) improves alignment in several configurations. Our analysis pinpoints two key reasons

View details →
Negative / Null Result ReportOpen accessComputer Science

Framing Effects on Privacy Concerns about a Home Telepresence Robot

Matthew Rueben, Frank J. Bernieri, Cindy M. Grimm et al. · 2019 · arXiv

Privacy-sensitive robotics is an emerging area of HRI research. Judgments about privacy would seem to be context-dependent, but none of the promising work on contextual "frames" has focused on privacy concerns. This work studies the impact of contextual "frames" on local users' privacy judgments in a home telepresence setting. Our methodology consists of using an online questionnaire to collect responses to animated videos of a telepresence robot after framing people with an introductory paragraph. The results of four studies indicate a large effect of manipulating the robot operator's identit

View details →
Negative / Null Result ReportOpen accessEconomics, Econometrics and Finance

Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina

Yuan Gao, Dokyun Lee, Gordon Burtch et al. · 2024 · arXiv

Recent studies suggest large language models (LLMs) can exhibit human-like reasoning, aligning with human behavior in economic experiments, surveys, and political discourse. This has led many to propose that LLMs can be used as surrogates or simulations for humans in social science research. However, LLMs differ fundamentally from humans, relying on probabilistic patterns, absent the embodied experiences or survival objectives that shape human cognition. We assess the reasoning depth of LLMs using the 11-20 money request game. Nearly all advanced approaches fail to replicate human behavior dis

View details →
Negative / Null Result Report

Patients’ daily reporting of symptoms via mobile application reveals a significant difference between patients’ perceptions and doctors’ interpretations

Cvetka Grašič Kuhar, Nina Privšek, Marjetka Sraka et al. · 2025 · Frontiers in Oncology

Purpose Electronic patient-reported outcomes (ePROs) are gaining importance. The aim of this study was to investigate the difference in the reporting of symptoms between patients via mobile application (m-app) and doctor assessments.…

View details →