e-ISSN: Pending

Browse the failure-mode index

744 real negative results, null findings, and replication failures in Computer Science. Search the index →

WASTE indexes published research — it does not host or republish full papers. Each entry is a metadata record compiled from open scholarly databases; the abstract is shown in full only where the paper is openly licensed, otherwise a short excerpt under fair use. Classifications are automated and approximate.

Negative / Null Result ReportOpen accessComputer Science

CogniPlay: a work-in-progress Human-like model for General Game Playing

Aloïs Rautureau, Éric Piette · 2025 · arXiv

While AI systems have equaled or surpassed human performance in a wide variety of games such as Chess, Go, or Dota 2, describing these systems as truly "human-like" remains far-fetched. Despite their success, they fail to replicate the pattern-based, intuitive decision-making processes observed in human cognition. This paper presents an overview of findings from cognitive psychology and previous efforts to model human-like behavior in artificial agents, discusses their applicability to General Game Playing (GGP) and introduces our work-in-progress model based on these observations: CogniPlay.

View details →
Negative / Null Result ReportOpen accessComputer Science

Towards Optimal Use of Exception Handling Information for Function Detection

Chengbin Pang, Ruotong Yu, Dongpeng Xu et al. · 2021 · arXiv

Function entry detection is critical for security of binary code. Conventional methods heavily rely on patterns, inevitably missing true functions and introducing errors. Recently, call frames have been used in exception-handling for function start detection. However, existing methods have two problems. First, they combine call frames with heuristic-based approaches, which often brings error and uncertain benefits. Second, they trust the fidelity of call frames, without handling the errors that are introduced by call frames. In this paper, we first study the coverage and accuracy of existing a

View details →
Negative / Null Result ReportOpen accessComputer Science

Time-aware Self-Attention Meets Logic Reasoning in Recommender Systems

Zhijian Luo, Zihan Huang, Jiahui Tang et al. · 2022 · arXiv

At the age of big data, recommender systems have shown remarkable success as a key means of information filtering in our daily life. Recent years have witnessed the technical development of recommender systems, from perception learning to cognition reasoning which intuitively build the task of recommendation as the procedure of logical reasoning and have achieve significant improvement. However, the logical statement in reasoning implicitly admits irrelevance of ordering, even does not consider time information which plays an important role in many recommendation tasks. Furthermore, recommenda

View details →
Negative / Null Result ReportOpen accessComputer Science

Does Character-level Information Always Improve DRS-based Semantic Parsing?

Tomoya Kurosawa, Hitomi Yanaka · 2023 · arXiv

Even in the era of massive language models, it has been suggested that character-level representations improve the performance of neural models. The state-of-the-art neural semantic parser for Discourse Representation Structures uses character-level representations, improving performance in the four languages (i.e., English, German, Dutch, and Italian) in the Parallel Meaning Bank dataset. However, how and why character-level information improves the parser's performance remains unclear. This study provides an in-depth analysis of performance changes by order of character sequences. In the exp

View details →
Failed Experiment ReportOpen accessComputer Science

Generalized positivity bounds on chiral perturbation theory

Yu-Jia Wang, Feng-Kun Guo, Cen Zhang et al. · 2020 · arXiv

Recently, a new set of positivity bounds with $t$ derivatives have been discovered. We explore the generic features of these generalized positivity bounds with loop amplitudes and apply these bounds to constrain the parameters in chiral perturbation theory up to the next-to-next-to-leading order. We show that the generalized positivity bounds give rise to stronger constraints on the $\bar l_i$ constants, compared to the existing axiomatic bounds. The parameter space of the $b_i$ constants is constrained by the generalized positivity bounds to be a convex region that is enclosed for many sectio

View details →
Negative / Null Result ReportOpen accessComputer Science

Asymptotic Expansions for Gaussian Channels with Feedback under a Peak Power Constraint

Silas L. Fong, Vincent Y. F. Tan · 2014 · arXiv

This paper investigates the asymptotic expansion for the size of block codes defined for the additive white Gaussian noise (AWGN) channel with feedback under the following setting: A peak power constraint is imposed on every transmitted codeword, and the average error probability of decoding the transmitted message is non-vanishing as the blocklength increases. It is well-known that the presence of feedback does not increase the first-order asymptotics (i.e., capacity) in the asymptotic expansion for the AWGN channel. The main contribution of this paper is a self-contained proof of an upper bo

View details →
Negative / Null Result ReportOpen accessComputer Science

Renyi Differential Privacy of the Subsampled Shuffle Model in Distributed Learning

Antonious M. Girgis, Deepesh Data, Suhas Diggavi · 2021 · arXiv

We study privacy in a distributed learning framework, where clients collaboratively build a learning model iteratively through interactions with a server from whom we need privacy. Motivated by stochastic optimization and the federated learning (FL) paradigm, we focus on the case where a small fraction of data samples are randomly sub-sampled in each round to participate in the learning process, which also enables privacy amplification. To obtain even stronger local privacy guarantees, we study this in the shuffle privacy model, where each client randomizes its response using a local different

View details →
Failed Experiment ReportOpen accessComputer Science

A Coding-Theoretic Application of Baranyai's Theorem

Liang Feng Zhang · 2013 · arXiv

Baranyai's theorem is a well-known theorem in the theory of hypergraphs. A corollary of this theorem says that one can partition the family of all $u$-subsets of an $n$-element set into ${n-1\choose u-1}$ sub-families such that each sub-family form a partition of the $n$-element set, where $n$ is divisible by $u$. In this paper, we present a coding-theoretic application of Baranyai's theorem (or equivalently, the corollary). More precisely, we propose the first purely combinatorial construction of locally decodable codes. Locally decodable codes are error-correcting codes that allow the recove

View details →
Negative / Null Result ReportOpen accessComputer Science

Testing operational phase concepts in quantum optics

J. Rehacek, Z. Hradil, M. Dusek et al. · 1999 · arXiv

An experimental comparison of several operational phase concepts is presented. In particular, it is shown that statistically motivated evaluation of experimental data may lead to a significant improvement in phase fitting upon the conventional Noh, Fouge'res and Mandel procedure. The analysis is extended to the asymptotic limit of large intensities, where a strong evidence in favor of multi--dimensional estimation procedures has been found.

View details →
Negative / Null Result ReportOpen accessComputer Science

Preventing Over-Smoothing for Hypergraph Neural Networks

Guanzi Chen, Jiying Zhang, Xi Xiao et al. · 2022 · arXiv

In recent years, hypergraph learning has attracted great attention due to its capacity in representing complex and high-order relationships. However, current neural network approaches designed for hypergraphs are mostly shallow, thus limiting their ability to extract information from high-order neighbors. In this paper, we show both theoretically and empirically, that the performance of hypergraph neural networks does not improve as the number of layers increases, which is known as the over-smoothing problem. To avoid this issue, we develop a new deep hypergraph convolutional network called De

View details →
Negative / Null Result ReportOpen accessComputer Science

Reliability of Large Language Model Generated Clinical Reasoning in Assisted Reproductive Technology: Blinded Comparative Evaluation Study

Dou Liu, Ying Long, Sophia Zuoqiu et al. · 2025 · arXiv

Creating high-quality clinical Chains-of-Thought (CoTs) is crucial for explainable medical Artificial Intelligence (AI) while constrained by data scarcity. Although Large Language Models (LLMs) can synthesize medical data, their clinical reliability remains unverified. This study evaluates the reliability of LLM-generated CoTs and investigates prompting strategies to enhance their quality. In a blinded comparative study, senior clinicians in Assisted Reproductive Technology (ART) evaluated CoTs generated via three distinct strategies: Zero-shot, Random Few-shot (using shallow examples), and Se

View details →
Negative / Null Result ReportOpen accessComputer Science

Neutron structure function and inclusive DIS from H-3 and He-3 targets at large Bjorken-x

M. M. Sargsian, S. Simula, M. I. Strikman · 2002 · arXiv

A detailed study of inclusive deep inelastic scattering from mirror A = 3 nuclei at large values of Bjorken-x is presented. The main purpose is to estimate the theoretical uncertainties on the extraction of F2n from such measurements. Within the convolution approach we confirm the cancellation of nuclear effects at the level of ~1 % for x < 0.75 in overall agreement with previous findings. However, within models in which modifications of the bound nucleon structure functions are accounted for to describe the EMC effect in nuclei, we find that the nuclear effects may be canceled at a level of ~

View details →
Negative / Null Result ReportOpen accessComputer Science

Handwritten Text Recognition from Crowdsourced Annotations

Solène Tarride, Tristan Faine, Mélodie Boillet et al. · 2023 · arXiv

In this paper, we explore different ways of training a model for handwritten text recognition when multiple imperfect or noisy transcriptions are available. We consider various training configurations, such as selecting a single transcription, retaining all transcriptions, or computing an aggregated transcription from all available annotations. In addition, we evaluate the impact of quality-based data selection, where samples with low agreement are removed from the training set. Our experiments are carried out on municipal registers of the city of Belfort (France) written between 1790 and 1946

View details →
Negative / Null Result ReportOpen accessComputer Science

Mean Estimation Under Heterogeneous Privacy: Some Privacy Can Be Free

Syomantak Chaudhuri, Thomas A. Courtade · 2023 · arXiv

Differential Privacy (DP) is a well-established framework to quantify privacy loss incurred by any algorithm. Traditional DP formulations impose a uniform privacy requirement for all users, which is often inconsistent with real-world scenarios in which users dictate their privacy preferences individually. This work considers the problem of mean estimation under heterogeneous DP constraints, where each user can impose their own distinct privacy level. The algorithm we propose is shown to be minimax optimal when there are two groups of users with distinct privacy levels. Our results elicit an in

View details →
Negative / Null Result ReportOpen accessComputer Science

Validity of common modelling approximations for precessing binary black holes with higher-order modes

Antoni Ramos-Buades, Patricia Schmidt, Geraint Pratten et al. · 2020 · arXiv

The current paradigm for constructing waveforms from precessing compact binaries is to first construct a waveform in a non-inertial, co-precessing binary source frame followed by a time-dependent rotation to map back to the physical, inertial frame. A key insight in the construction of these models is that the co-precessing waveform can be effectively mapped to some equivalent aligned spin waveform. Secondly, the time-dependent rotation implicitly introduces $m$-mode mixing, necessitating an accurate description of higher-order modes in the co-precessing frame. We assess the efficacy of this m

View details →
Negative / Null Result ReportOpen accessComputer Science

Single- and Multi-Task Architectures for Tool Presence Detection Challenge at M2CAI 2016

Andru P. Twinanda, Didier Mutter, Jacques Marescaux et al. · 2016 · arXiv

The tool presence detection challenge at M2CAI 2016 consists of identifying the presence/absence of seven surgical tools in the images of cholecystectomy videos. Here, we propose to use deep architectures that are based on our previous work where we presented several architectures to perform multiple recognition tasks on laparoscopic videos. In this technical report, we present the tool presence detection results using two architectures: (1) a single-task architecture designed to perform solely the tool presence detection task and (2) a multi-task architecture designed to perform jointly phase

View details →
Negative / Null Result ReportOpen accessComputer Science

From Co-Design to Metacognitive Laziness: Evaluating Generative AI in Vocational Education

Amir Yunus, Peng Rend Gay, Oon Teng Lee · 2025 · arXiv

This study examines the development and deployment of a Generative AI proof-of-concept (POC) designed to support lecturers in a vocational education setting in Singapore. Employing a user-centred, mixed-methods design process, we co-developed an AI chatbot with lecturers to address recurring instructional challenges during exam preparation, specifically managing repetitive questions and scaling feedback delivery. The POC achieved its primary operational goals: lecturers reported streamlined workflows, reduced cognitive load, and observed improved student confidence in navigating course content

View details →
Negative / Null Result ReportOpen accessComputer Science

Refined Continuous Control of DDPG Actors via Parametrised Activation

Mohammed Hossny, Julie Iskander, Mohammed Attia et al. · 2020 · arXiv

In this paper, we propose enhancing actor-critic reinforcement learning agents by parameterising the final actor layer which produces the actions in order to accommodate the behaviour discrepancy of different actuators, under different load conditions during interaction with the environment. We propose branching the action producing layer in the actor to learn the tuning parameter controlling the activation layer (e.g. Tanh and Sigmoid). The learned parameters are then used to create tailored activation functions for each actuator. We ran experiments on three OpenAI Gym environments, i.e. Pend

View details →
Negative / Null Result ReportOpen accessComputer Science

Augmenting Immersive Telepresence Experience with a Virtual Body

Nikunj Arora, Markku Suomalainen, Matti Pouke et al. · 2022 · arXiv

We propose augmenting immersive telepresence by adding a virtual body, representing the user's own arm motions, as realized through a head-mounted display and a 360-degree camera. Previous research has shown the effectiveness of having a virtual body in simulated environments; however, research on whether seeing one's own virtual arms increases presence or preference for the user in an immersive telepresence setup is limited. We conducted a study where a host introduced a research lab while participants wore a head-mounted display which allowed them to be telepresent at the host's physical loc

View details →
Negative / Null Result ReportOpen accessComputer Science

MaScQA: A Question Answering Dataset for Investigating Materials Science Knowledge of Large Language Models

Mohd Zaki, Jayadeva, Mausam et al. · 2023 · arXiv

Information extraction and textual comprehension from materials literature are vital for developing an exhaustive knowledge base that enables accelerated materials discovery. Language models have demonstrated their capability to answer domain-specific questions and retrieve information from knowledge bases. However, there are no benchmark datasets in the materials domain that can evaluate the understanding of the key concepts by these language models. In this work, we curate a dataset of 650 challenging questions from the materials domain that require the knowledge and skills of a materials st

View details →
Negative / Null Result ReportOpen accessComputer Science

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models

Jason Z Wang · 2026 · arXiv

We introduce MIRROR, a benchmark comprising eight experiments across four metacognitive levels that evaluates whether large language models can use self-knowledge to make better decisions. We evaluate 16 models from 8 labs across approximately 250,000 evaluation instances using five independent behavioral measurement channels. Core experiments are run across the full model roster; experiments with specialized infrastructure requirements report explicitly marked model subsets. We find two phenomena with direct implications for agentic deployment: (1) compositional self-prediction fails universa

View details →
Negative / Null Result ReportOpen accessComputer Science

Significant Improvements over the State of the Art? A Case Study of the MS MARCO Document Ranking Leaderboard

Jimmy Lin, Daniel Campos, Nick Craswell et al. · 2021 · arXiv

Leaderboards are a ubiquitous part of modern research in applied machine learning. By design, they sort entries into some linear order, where the top-scoring entry is recognized as the "state of the art" (SOTA). Due to the rapid progress being made in information retrieval today, particularly with neural models, the top entry in a leaderboard is replaced with some regularity. These are touted as improvements in the state of the art. Such pronouncements, however, are almost never qualified with significance testing. In the context of the MS MARCO document ranking leaderboard, we pose a specific

View details →
Negative / Null Result ReportOpen accessComputer Science

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study

Xiaolong Jin, Xuandong Zhao, Wenbo Guo et al. · 2026 · arXiv

Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in large language model reasoning, but relies on ground-truth supervision that is costly or infeasible, especially in coding tasks. Recent work addresses this by deriving rewards from a model's own signals, such as majority voting or confidence-based scores, achieving notable success on mathematical reasoning benchmarks. However, code generation poses distinct challenges: programs are structurally complex, semantically equivalent solutions may differ syntactically, and verification typically requires executio

View details →
Negative / Null Result ReportOpen accessComputer Science

Association Is Not Similarity: Learning Corpus-Specific Associations for Multi-Hop Retrieval

Jason Dury · 2026 · arXiv

Dense retrieval systems rank passages by embedding similarity to a query, but multi-hop questions require passages that are associatively related through shared reasoning chains. We introduce Association-Augmented Retrieval (AAR), a lightweight transductive reranking method that trains a small MLP (4.2M parameters) to learn associative relationships between passages in embedding space using contrastive learning on co-occurrence annotations. At inference time, AAR reranks an initial dense retrieval candidate set using bi-directional association scoring. On HotpotQA, AAR improves passage Recall@

View details →
Negative / Null Result ReportOpen accessComputer Science

Benchmarking Language Model Creativity: A Case Study on Code Generation

Yining Lu, Dixuan Wang, Tianjian Li et al. · 2024 · arXiv

As LLMs become increasingly prevalent, it is interesting to consider how ``creative'' these models can be. From cognitive science, creativity consists of at least two key characteristics: \emph{convergent} thinking (purposefulness to achieve a given goal) and \emph{divergent} thinking (adaptability to explore new environments or constraints) \citep{runco2003critical}. In this work, we introduce a framework for quantifying LLM creativity that incorporates the two design ingredients: (1) We introduce DENIAL PROMPTING which pushes LLMs to develop more creative solutions to a given problem by incr

View details →
Negative / Null Result ReportOpen accessComputer Science

When Medical Imaging Met Self-Attention: A Love Story That Didn't Quite Work Out

Tristan Piater, Niklas Penzel, Gideon Stein et al. · 2024 · arXiv

A substantial body of research has focused on developing systems that assist medical professionals during labor-intensive early screening processes, many based on convolutional deep-learning architectures. Recently, multiple studies explored the application of so-called self-attention mechanisms in the vision domain. These studies often report empirical improvements over fully convolutional approaches on various datasets and tasks. To evaluate this trend for medical imaging, we extend two widely adopted convolutional architectures with different self-attention variants on two different medical

View details →
Negative / Null Result ReportOpen accessComputer Science

Hidden Meanings in Plain Sight: RebusBench for Evaluating Cognitive Visual Reasoning

Seyed Amir Kasaei, Arash Marioriyad, Mahbod Khaleti et al. · 2026 · arXiv

Large Vision-Language Models (LVLMs) have achieved remarkable proficiency in explicit visual recognition, effectively describing what is directly visible in an image. However, a critical cognitive gap emerges when the visual input serves only as a clue rather than the answer. We identify that current models struggle with the complex, multi-step reasoning required to solve problems where information is not explicitly depicted. Successfully solving a rebus puzzle requires a distinct cognitive workflow: the model must extract visual and textual attributes, retrieve linguistic prior knowledge (suc

View details →
Negative / Null Result ReportOpen accessComputer Science

Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?

Hyeong Kyu Choi, Xiaojin Zhu, Sharon Li · 2025 · arXiv

Multi-Agent Debate~(MAD) has emerged as a promising paradigm for improving the performance of large language models through collaborative reasoning. Despite recent advances, the key factors driving MAD's effectiveness remain unclear. In this work, we disentangle MAD into two key components--Majority Voting and inter-agent Debate--and assess their respective contributions. Through extensive experiments across seven NLP benchmarks, we find that Majority Voting alone accounts for most of the performance gains typically attributed to MAD. To explain this, we propose a theoretical framework that mo

View details →
Failed Experiment ReportOpen accessComputer Science

Nearly ETH-Tight Algorithms for Planar Steiner Tree with Terminals on Few Faces

Sándor Kisfaludi-Bak, Jesper Nederlof, Erik Jan van Leeuwen · 2018 · arXiv

The Planar Steiner Tree problem is one of the most fundamental NP-complete problems as it models many network design problems. Recall that an instance of this problem consists of a graph with edge weights, and a subset of vertices (often called terminals); the goal is to find a subtree of the graph of minimum total weight that connects all terminals. A seminal paper by Erickson et al. [Math. Oper. Res., 1987] considers instances where the underlying graph is planar and all terminals can be covered by the boundary of $k$ faces. Erickson et al. show that the problem can be solved by an algorithm

View details →
Negative / Null Result ReportOpen accessComputer Science

The Unlearnability Phenomenon in RLVR for Language Models

Yulin Chen, He He, Chen Zhao · 2026 · arXiv

Reinforcement Learning with Verifiable Reward (RLVR) has proven effective in improving Large Language Model's (LLM) reasoning ability. However, the learning dynamics of RLVR remain underexplored. In this paper, we reveal a counterintuitive phenomenon: among hard examples that the model initially struggles with, a substantial subset remains unlearnable even when correct rollouts are present. To understand the phenomenon, we first demonstrate that existing optimization and sampling techniques fail to resolve unlearnability. With cross-example gradient analysis, we show that unlearnable examples

View details →