e-ISSN: Pending

Browse the failure-mode index

744 real negative results, null findings, and replication failures in Computer Science. Search the index →

WASTE indexes published research — it does not host or republish full papers. Each entry is a metadata record compiled from open scholarly databases; the abstract is shown in full only where the paper is openly licensed, otherwise a short excerpt under fair use. Classifications are automated and approximate.

Negative / Null Result ReportOpen accessComputer Science

LLMs Corrupt Your Documents When You Delegate

Philippe Laban, Tobias Schnabel, Jennifer Neville · 2026 · arXiv

Large Language Models (LLMs) are poised to disrupt knowledge work, with the emergence of delegated work as a new interaction paradigm (e.g., vibe coding). Delegation requires trust - the expectation that the LLM will faithfully execute the task without introducing errors into documents. We introduce DELEGATE-52 to study the readiness of AI systems in delegated workflows. DELEGATE-52 simulates long delegated workflows that require in-depth document editing across 52 professional domains, such as coding, crystallography, and music notation. Our large-scale experiment with 19 LLMs reveals that cu

View details →
Negative / Null Result ReportOpen accessComputer Science

Bounds on Nonlocality and Random Access Codes from Extended Information Causality Principle

Prabhav Jain, Nikolai Miklin, Mariami Gachechiladze · 2026 · arXiv

Information Causality was introduced as a physical principle for constraining the set of nonlocal correlations. In recent work, we proposed an extension of Information Causality that allows correlations among Alice's inputs. This extended principle yields tighter constraints than the original formulation and recovers part of the quantum boundary in certain Bell scenarios. In this work, we further investigate the implications of extended Information Causality and apply it to scenarios beyond binary inputs and outputs. We derive a family of quantum Bell inequalities that strengthen previously kn

View details →
Negative / Null Result ReportOpen accessComputer Science

Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models

Jirui Qi, Raquel Fernández, Arianna Bisazza · 2023 · arXiv

Multilingual large-scale Pretrained Language Models (PLMs) have been shown to store considerable amounts of factual knowledge, but large variations are observed across languages. With the ultimate goal of ensuring that users with different language backgrounds obtain consistent feedback from the same model, we study the cross-lingual consistency (CLC) of factual knowledge in various multilingual PLMs. To this end, we propose a Ranking-based Consistency (RankC) metric to evaluate knowledge consistency across languages independently from accuracy. Using this metric, we conduct an in-depth analys

View details →
Negative / Null Result ReportOpen accessComputer Science

Collaborative Distributed Hypothesis Testing

Gil Katz, Pablo Piantanida, Merouane Debbah · 2016 · arXiv

A collaborative distributed binary decision problem is considered. Two statisticians are required to declare the correct probability measure of two jointly distributed memoryless process, denoted by $X^n=(X_1,\dots,X_n)$ and $Y^n=(Y_1,\dots,Y_n)$, out of two possible probability measures on finite alphabets, namely $P_{XY}$ and $P_{\bar{X}\bar{Y}}$. The marginal samples given by $X^n$ and $Y^n$ are assumed to be available at different locations. The statisticians are allowed to exchange limited amount of data over multiple rounds of interactions, which differs from previous work that deals mai

View details →
Negative / Null Result ReportOpen accessComputer Science

Optimal EEG Electrode Set for Emotion Recognition From Brain Signals: An Empirical Quest

Rumman Ahmed Prodhan, Sumya Akter, Tanmoy Sarkar Pias et al. · 2023 · arXiv

The human brain is a complex organ, still completely undiscovered, that controls almost all the parts of the body. Apart from survival, the human brain stimulates emotions. Recent research indicates that brain signals can be very effective for emotion recognition. However, which parts of the brain exhibit most of the emotions is still under-explored. In this study, we empirically analyze the contribution of each part of the brain in exhibiting emotions. We use the DEAP dataset to find the most optimal electrode set which eventually leads to the effective brain part associated with emotions. We

View details →
Negative / Null Result ReportOpen accessComputer Science

Improved WKB analysis of cosmological perturbations

Roberto Casadio, Fabio Finelli, Mattia Luzzi et al. · 2004 · arXiv

Improved Wentzel-Kramers-Brillouin (WKB)-type approximations are presented in order to study cosmological perturbations beyond the lowest order. Our methods are based on functions which approximate the true perturbation modes over the complete range of the independent (Langer) variable, from sub-horizon to super-horizon scales, and include the region near the turning point. We employ both a perturbative Green's function technique and an adiabatic (or ``semiclassical'') expansion (for a linear turning point) in order to compute higher order corrections. Improved general expressions for the WKB

View details →
Negative / Null Result ReportOpen accessComputer Science

VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding

Abdul Waheed, Zhen Wu, Dareen Alharthi et al. · 2025 · arXiv

Precisely evaluating video understanding models remains challenging: commonly used metrics such as BLEU, ROUGE, and BERTScore fail to capture the fineness of human judgment, while obtaining such judgments through manual evaluation is costly. Recent work has explored using large language models (LLMs) or multimodal LLMs (MLLMs) as evaluators, but their extension to video understanding remains relatively unexplored. In this work, we introduce VideoJudge, a 3B and 7B-sized MLLM judge specialized to evaluate outputs from video understanding models (\textit{i.e.}, text responses conditioned on vide

View details →
Negative / Null Result ReportOpen accessComputer Science

Unitarity Problems in 3$D$ Gravity Theories

Gokhan Alkac, Luca Basanisi, Ercan Kilicarslan et al. · 2017 · arXiv

We revisit the problem of the bulk-boundary unitarity clash in 2 + 1 dimensional gravity theories, which has been an obstacle in providing a viable dual two-dimensional conformal field theory for bulk gravity in anti-de Sitter (AdS) spacetime. Chiral gravity, which is a particular limit of cosmological topologically massive gravity (TMG), suffers from pertur- bative log-modes with negative energies inducing a non-unitary logarithmic boundary field theory. We show here that any f(R) extension of TMG does not improve the situation. We also study the perturbative modes in the metric formulation o

View details →
Negative / Null Result ReportOpen accessComputer Science

The Curious Case of Representational Alignment: Unravelling Visio-Linguistic Tasks in Emergent Communication

Tom Kouwenhoven, Max Peeperkorn, Bram van Dijk et al. · 2024 · arXiv

Natural language has the universal properties of being compositional and grounded in reality. The emergence of linguistic properties is often investigated through simulations of emergent communication in referential games. However, these experiments have yielded mixed results compared to similar experiments addressing linguistic properties of human language. Here we address representational alignment as a potential contributing factor to these results. Specifically, we assess the representational alignment between agent image representations and between agent representations and input images.

View details →
Negative / Null Result ReportOpen accessComputer Science

Laser and cavity cooling of a mechanical resonator with a Nitrogen-Vacancy center in diamond

Luigi Giannelli, Ralf Betzholz, Laura Kreiner et al. · 2016 · arXiv

We theoretically analyse the cooling dynamics of a high-Q mode of a mechanical resonator, when the structure is also an optical cavity and is coupled with a NV center. The NV center is driven by a laser and interacts with the cavity photon field and with the strain field of the mechanical oscillator, while radiation pressure couples mechanical resonator and cavity field. Starting from the full master equation we derive the rate equation for the mechanical resonator's motion, whose coefficients depend on the system parameters and on the noise sources. We then determine the cooling regime, the c

View details →
Negative / Null Result ReportOpen accessComputer Science

MedBayes-Lite: A Clinical Uncertainty Governance Layer for Risk-Aware Medical Decision Support

Elias Hossain, Md Mehedi Hasan Nipu, Maleeha Sheikh et al. · 2025 · arXiv

Clinical language models often assign high confidence to incorrect predictions, particularly in high-severity and out-of-distribution cases. We present MedBayes-Lite, a retraining-free uncertainty governance layer for transformer-based clinical predictors. It combines Monte Carlo dropout, predictive calibration, and confidence-guided abstention to defer low-confidence predictions for human review, adding no trainable parameters. Evaluated on MedMCQA and MedQA-USMLE, MedBayes-Lite reduces expected calibration error by 0.23 to 0.33 and drives harmful overconfident errors (confident, incorrect, h

View details →
Negative / Null Result ReportOpen accessComputer Science

Endpoint Logarithms in the NLO Mueller-Navelet Jet Vertex: Threshold Matching and BLM/MOM Prescription Sensitivity

Lei Wang · 2026 · arXiv

The endpoint region $ζ\to1$ of the NLO forward jet vertex has not been systematically separated from BFKL energy-scale terms in Mueller-Navelet phenomenology. Starting from the small-cone NLO vertex, we isolate the quark and gluon plus distributions and construct a BFKL-aware threshold matching scheme that preserves exact NLO accuracy. The conservative Scheme-II exponent resums only the ordinary endpoint logarithms and leaves the $χ(n,γ)\ln\bar N$ term in the fixed-order coefficient, avoiding an uncontrolled tower of mixed endpoint-BFKL logarithms. In fixed-baseline CMS tests, this matched ver

View details →
Negative / Null Result ReportOpen accessComputer Science

A Tight Upper Bound on the Second-Order Coding Rate of the Parallel Gaussian Channel with Feedback

Silas L. Fong, Vincent Y. F. Tan · 2017 · arXiv

This paper investigates the asymptotic expansion for the maximum rate of fixed-length codes over a parallel Gaussian channel with feedback under the following setting: A peak power constraint is imposed on every transmitted codeword, and the average error probability of decoding the transmitted message is non-vanishing as the blocklength increases. It is well known that the presence of feedback does not increase the first-order asymptotics of the channel, i.e., capacity, in the asymptotic expansion, and the closed-form expression of the capacity can be obtained by the well-known water-filling

View details →
Negative / Null Result ReportOpen accessComputer Science

Memes-as-Replies: Can Models Select Humorous Manga Panel Responses?

Ryosuke Kohita, Seiichiro Yoshioka · 2026 · arXiv

Memes are a popular element of modern web communication, used not only as static artifacts but also as interactive replies within conversations. While computational research has focused on analyzing the intrinsic properties of memes, the dynamic and contextual use of memes to create humor remains an understudied area of web science. To address this gap, we introduce the Meme Reply Selection task and present MaMe-Re (Manga Meme Reply Benchmark), a benchmark of 100,000 human-annotated pairs (500,000 total annotations from 2,325 unique annotators) consisting of openly licensed Japanese manga pane

View details →
Negative / Null Result ReportOpen accessComputer Science

Analyzing and Adapting Large Language Models for Few-Shot Multilingual NLU: Are We There Yet?

Evgeniia Razumovskaia, Ivan Vulić, Anna Korhonen · 2024 · arXiv

Supervised fine-tuning (SFT), supervised instruction tuning (SIT) and in-context learning (ICL) are three alternative, de facto standard approaches to few-shot learning. ICL has gained popularity recently with the advent of LLMs due to its simplicity and sample efficiency. Prior research has conducted only limited investigation into how these approaches work for multilingual few-shot learning, and the focus so far has been mostly on their performance. In this work, we present an extensive and systematic comparison of the three approaches, testing them on 6 high- and low-resource languages, thr

View details →
Negative / Null Result ReportOpen accessComputer Science

Point-in-Time Financial RAG with Frozen LLMs and Market-Feedback Adaptive Retrieval

Zijie Zhao, Roy E. Welsch · 2026 · arXiv

Financial retrieval-augmented generation (RAG) systems typically rank evidence by textual relevance, but in financial markets evidence utility depends on event type, forecast horizon, and market context. We study news-triggered event-impact prediction as a point-in-time financial RAG problem. For each company-news anchor, the system retrieves financial news and SEC filing passages, appends a pre-decision market-context card, and predicts multi-horizon residual-return signals. Our method keeps the LLM frozen and adapts retrieval through an external Bayesian source memory updated from matured re

View details →
Negative / Null Result ReportOpen accessComputer Science

Performance Analysis of Cooperative Communications at Road Intersections Using Stochastic Geometry Tools

Baha Eddine Youcef Belmekki, Abdelkrim Hamza, Benoît Escrig · 2018 · arXiv

Vehicular safety communications (VSCs) are known to provide relevant contributions to avoid congestions and prevent road accidents, and more particularly at road intersections since these areas are more prone to accidents. In this context, one of the main impairments that affect the performance of VSCs are interference. In this paper, we develop a tractable framework to model cooperative transmissions in presence of interference for VSCs at intersections. We use tools from stochastic geometry, and model interferer vehicles locations as a Poisson point process. First, we calculate the outage pr

View details →
Negative / Null Result ReportOpen accessComputer Science

Simulating Eating Disorder Patients with LLMs: Evaluating Psychological Persona Stability in Multi-Turn Conversations

Jennifer Haase, Jana Gonnermann-Müller, See Heng Yim et al. · 2026 · arXiv

Large language model (LLM)-based simulations of clinical patients are increasingly used for research and training, yet their validity requires persona stability: coherent maintenance of an assigned psychological profile across and within conversations. We evaluate this prerequisite using eating disorder personas grounded in five published case vignettes, a dual-assessment framework (self-report + independent observer ratings), and validated psychometric instruments (EDE-Q) with known ground-truth scores. Across six LLMs and two experiments (between-conversation stability (Exp. I) and within-co

View details →
Negative / Null Result ReportOpen accessComputer Science

My Body is a Cage: the Role of Morphology in Graph-Based Incompatible Control

Vitaly Kurin, Maximilian Igl, Tim Rocktäschel et al. · 2020 · arXiv

Multitask Reinforcement Learning is a promising way to obtain models with better performance, generalisation, data efficiency, and robustness. Most existing work is limited to compatible settings, where the state and action space dimensions are the same across tasks. Graph Neural Networks (GNN) are one way to address incompatible environments, because they can process graphs of arbitrary size. They also allow practitioners to inject biases encoded in the structure of the input graph. Existing work in graph-based continuous control uses the physical morphology of the agent to construct the inpu

View details →
Negative / Null Result ReportOpen accessComputer Science

Sparse Graph Learning with Spectrum Prior for Deep Graph Convolutional Networks

Jin Zeng, Yang Liu, Gene Cheung et al. · 2022 · arXiv

A graph convolutional network (GCN) employs a graph filtering kernel tailored for data with irregular structures. However, simply stacking more GCN layers does not improve performance; instead, the output converges to an uninformative low-dimensional subspace, where the convergence rate is characterized by the graph spectrum -- this is the known over-smoothing problem in GCN. In this paper, we propose a sparse graph learning algorithm incorporating a new spectrum prior to compute a graph topology that circumvents over-smoothing while preserving pairwise correlations inherent in data. Specifica

View details →
Negative / Null Result ReportOpen accessComputer Science

A pipeline for fair comparison of graph neural networks in node classification tasks

Wentao Zhao, Dalin Zhou, Xinguo Qiu et al. · 2020 · arXiv

Graph neural networks (GNNs) have been investigated for potential applicability in multiple fields that employ graph data. However, there are no standard training settings to ensure fair comparisons among new methods, including different model architectures and data augmentation techniques. We introduce a standard, reproducible benchmark to which the same training settings can be applied for node classification. For this benchmark, we constructed 9 datasets, including both small- and medium-scale datasets from different fields, and 7 different models. We design a k-fold model assessment strate

View details →
Negative / Null Result ReportOpen accessComputer Science

Incompatibility Clustering as a Defense Against Backdoor Poisoning Attacks

Charles Jin, Melinda Sun, Martin Rinard · 2021 · arXiv

We propose a novel clustering mechanism based on an incompatibility property between subsets of data that emerges during model training. This mechanism partitions the dataset into subsets that generalize only to themselves, i.e., training on one subset does not improve performance on the other subsets. Leveraging the interaction between the dataset and the training process, our clustering mechanism partitions datasets into clusters that are defined by--and therefore meaningful to--the objective of the training process. We apply our clustering mechanism to defend against data poisoning attacks,

View details →
Negative / Null Result ReportOpen accessComputer Science

On Relativistic $f$-Divergences

Alexia Jolicoeur-Martineau · 2019 · arXiv

This paper provides a more rigorous look at Relativistic Generative Adversarial Networks (RGANs). We prove that the objective function of the discriminator is a statistical divergence for any concave function $f$ with minimal properties ($f(0)=0$, $f'(0) \neq 0$, $\sup_x f(x)>0$). We also devise a few variants of relativistic $f$-divergences. Wasserstein GAN was originally justified by the idea that the Wasserstein distance (WD) is most sensible because it is weak (i.e., it induces a weak topology). We show that the WD is weaker than $f$-divergences which are weaker than relativistic $f$-diver

View details →
Negative / Null Result ReportOpen accessComputer Science

Analysis of Variational Sparse Autoencoders

Zachary Baker, Yuxiao Li · 2025 · arXiv

Sparse Autoencoders (SAEs) have emerged as a promising approach for interpreting neural network representations by learning sparse, human-interpretable features from dense activations. We investigate whether incorporating variational methods into SAE architectures can improve feature organization and interpretability. We introduce the Variational Sparse Autoencoder (vSAE), which replaces deterministic ReLU gating with stochastic sampling from learned Gaussian posteriors and incorporates KL divergence regularization toward a standard normal prior. Our hypothesis is that this probabilistic sampl

View details →
Negative / Null Result ReportOpen accessComputer Science

Reproducing DragDiffusion: Interactive Point-Based Editing with Diffusion Models

Ali Subhan, Ashir Raza · 2026 · arXiv

DragDiffusion is a diffusion-based method for interactive point-based image editing that enables users to manipulate images by directly dragging selected points. The method claims that accurate spatial control can be achieved by optimizing a single diffusion latent at an intermediate timestep, together with identity-preserving fine-tuning and spatial regularization. This work presents a reproducibility study of DragDiffusion using the authors' released implementation and the DragBench benchmark. We reproduce the main ablation studies on diffusion timestep selection, LoRA-based fine-tuning, mas

View details →
Replication FailureOpen accessComputer Science

Chains That See, Answers That Don't: A Multi-Aspect Evaluation Recipe for Forced Chain-of-Thought on Video-MME

Zhichao Fan, Yanhang Li, Zexin Zhuang · 2026 · arXiv

Forced chain-of-thought (CoT) is widely assumed to make vision-language models more reliable on video question answering. We propose a small three-probe evaluation recipe to test that assumption: paired accuracy across direct, CoT, answer-first, and no-video conditions; a counterfactual video-swap diagnostic over the CoT chains; and a four-rung visual-degradation ladder. Each probe is reported under both a strict and a permissive regex scorer, with multiplicity correction over a manuscript-declared primary family. Applied to Qwen2.5-VL on Video-MME subsets, the recipe returns a two-part findin

View details →
Replication FailureOpen accessComputer Science

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate

Xiqi Hao, Zengqing Wu, Yu-Xuan Qiu et al. · 2026 · arXiv

Multi-agent debate (MAD) is a promising strategy for improving LLM reasoning, but when agents converge on a shared answer, it is unclear whether that convergence reflects genuine deliberation or social compliance. We show that the conventional answer flip rate conflates three distinct mechanisms: spontaneous instability, stance-induced conformity, and reasoning-induced persuasion. Our three-source decomposition framework isolates each through controlled counterfactual conditions. In the primary MMLU-Pro setting, 37% of agent-question observations change under self-reflection alone, while robus

View details →
Replication FailureOpen accessComputer Science

Contrastive Predictive Coding Done Right for Mutual Information Estimation

J. Jon Ryu, Pavan Yeddanapudi, Xiangxiang Xu et al. · 2025 · arXiv

The InfoNCE objective, originally introduced for contrastive representation learning, has become a popular choice for mutual information (MI) estimation, despite its indirect connection to MI. In this paper, we demonstrate why InfoNCE should not be regarded as a valid MI estimator, and we introduce a simple modification, which we refer to as InfoNCE-anchor, for accurate MI estimation. Our modification introduces an auxiliary anchor class, enabling consistent density ratio estimation and yielding a plug-in MI estimator with significantly reduced bias. Beyond this, we generalize our framework us

View details →
Negative / Null Result ReportOpen accessComputer Science

Temporal Adaptation of BERT and Performance on Downstream Document Classification: Insights from Social Media

Paul Röttger, Janet B. Pierrehumbert · 2021 · arXiv

Language use differs between domains and even within a domain, language use changes over time. For pre-trained language models like BERT, domain adaptation through continued pre-training has been shown to improve performance on in-domain downstream tasks. In this article, we investigate whether temporal adaptation can bring additional benefits. For this purpose, we introduce a corpus of social media comments sampled over three years. It contains unlabelled data for adaptation and evaluation on an upstream masked language modelling task as well as labelled data for fine-tuning and evaluation on

View details →
Negative / Null Result ReportOpen accessComputer Science

Spiking Two-Stream Methods with Unsupervised STDP-based Learning for Action Recognition

Mireille El-Assal, Pierre Tirilly, Ioan Marius Bilasco · 2023 · arXiv

Video analysis is a computer vision task that is useful for many applications like surveillance, human-machine interaction, and autonomous vehicles. Deep Convolutional Neural Networks (CNNs) are currently the state-of-the-art methods for video analysis. However they have high computational costs, and need a large amount of labeled data for training. In this paper, we use Convolutional Spiking Neural Networks (CSNNs) trained with the unsupervised Spike Timing-Dependent Plasticity (STDP) learning rule for action classification. These networks represent the information using asynchronous low-ener

View details →