e-ISSN: Pending

Browse the failure-mode index

694 real negative results, null findings, and replication failures in Computer Science · Negative / Null Result Report. Search the index →

WASTE indexes published research — it does not host or republish full papers. Each entry is a metadata record compiled from open scholarly databases; the abstract is shown in full only where the paper is openly licensed, otherwise a short excerpt under fair use. Classifications are automated and approximate.

Negative / Null Result ReportOpen accessComputer Science

Exploring Major Transitions in the Evolution of Biological Cognition With Artificial Neural Networks

Konstantinos Voudouris, Andrew Barron, Marta Halina et al. · 2025 · arXiv

Transitional accounts of evolution emphasise a few changes that shape what is evolvable, with dramatic consequences for derived lineages. More recently it has been proposed that cognition might also have evolved via a series of major transitions that manipulate the structure of biological neural networks, fundamentally changing the flow of information. We used idealised models of information flow, artificial neural networks (ANNs), to evaluate whether changes in information flow in a network can yield a transitional change in cognitive performance. We compared networks with feed-forward, recur

View details →
Negative / Null Result ReportOpen accessComputer Science

When LLMs fall short in Deductive Coding: Model Comparison and Human AI Collaboration Workflow Design

Zijian Li, Luzhen Tang, Mengyu Xia et al. · 2025 · arXiv

With generative artificial intelligence driving the growth of dialogic data in education, automated coding is a promising direction for learning analytics to improve efficiency. This surge highlights the need to understand the nuances of student-AI interactions, especially those rare yet crucial. However, automated coding may struggle to capture these rare codes due to imbalanced data, while human coding remains time-consuming and labour-intensive. The current study examined the potential of large language models (LLMs) to approximate or replace humans in deductive, theory-driven coding, while

View details →
Negative / Null Result ReportOpen accessComputer Science

Active Learning of Molecular Data for Task-Specific Objectives

Kunal Ghosh, Milica Todorović, Aki Vehtari et al. · 2024 · arXiv

Active learning (AL) has shown promise for being a particularly data-efficient machine learning approach. Yet, its performance depends on the application and it is not clear when AL practitioners can expect computational savings. Here, we carry out a systematic AL performance assessment for three diverse molecular datasets and two common scientific tasks: compiling compact, informative datasets and targeted molecular searches. We implemented AL with Gaussian processes (GP) and used the many-body tensor as molecular representation. For the first task, we tested different data acquisition strate

View details →
Negative / Null Result ReportOpen accessComputer Science

Algorithmic Tradeoffs, Applied NLP, and the State-of-the-Art Fallacy

AJ Alvero, Ruohong Dong, Klint Kanopka et al. · 2025 · arXiv

Computational sociology is growing in popularity, yet the analytic tools employed differ widely in power, transparency, and interpretability. In computer science, methods gain popularity after surpassing benchmarks of predictive accuracy, becoming the "state of the art." Computer scientists favor novelty and innovation for different reasons, but prioritizing technical prestige over methodological fit could unintentionally limit the scope of sociological inquiry. To illustrate, we focus on computational text analysis and revisit a prior study of college admissions essays, comparing analyses wit

View details →
Negative / Null Result ReportOpen accessComputer Science

Teleportation of Hybrid Entangled States with Continuous-Variable Entanglement

Mingjian He, Robert Malaney · 2022 · arXiv

Hybrid entanglement between discrete-variable (DV) and continuous-variable (CV) quantum systems is an essential resource for heterogeneous quantum networks. Our previous work showed that in lossy channels the teleportation of DV qubits, via CV-entangled states, can be significantly improved by a new protocol defined by a modified Bell state measurement at the sender. This work explores whether a new, similarly modified, CV-based teleportation protocol can lead to improvement in the transfer of hybrid entangled states. To set the scene, we first determine the performance of such a modified prot

View details →
Negative / Null Result ReportOpen accessComputer Science

Cross-lingual robustness of LLM-brain alignment and its computational roots

Ni Yang, Rui He, Philipp Homan et al. · 2026 · arXiv

Large language models (LLMs) reliably predict neural activity during language comprehension and transformer depth has been interpreted as mirroring hierarchical cortical organization. However, it remains unclear whether such alignment extends to subcortical regions, overlaps spatially across languages, and what the computational roots of such alignment are. Here, we used a multilingual, whole-brain encoding framework to examine brain-LLM alignment across three typologically distinct languages: Mandarin, English, and French during naturalistic story listening. Our results show that across langu

View details →
Negative / Null Result ReportOpen accessComputer Science

Reducing Biases towards Minoritized Populations in Medical Curricular Content via Artificial Intelligence for Fairer Health Outcomes

Chiman Salavati, Shannon Song, Willmar Sosa Diaz et al. · 2024 · arXiv

Biased information (recently termed bisinformation) continues to be taught in medical curricula, often long after having been debunked. In this paper, we introduce BRICC, a firstin-class initiative that seeks to mitigate medical bisinformation using machine learning to systematically identify and flag text with potential biases, for subsequent review in an expert-in-the-loop fashion, thus greatly accelerating an otherwise labor-intensive process. A gold-standard BRICC dataset was developed throughout several years, and contains over 12K pages of instructional materials. Medical experts meticul

View details →
Negative / Null Result ReportOpen accessComputer Science

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Mirac Suzgun, Nathan Scales, Nathanael Schärli et al. · 2022 · arXiv

BIG-Bench (Srivastava et al., 2022) is a diverse evaluation suite that focuses on tasks believed to be beyond the capabilities of current language models. Language models have already made good progress on this benchmark, with the best model in the BIG-Bench paper outperforming average reported human-rater results on 65% of the BIG-Bench tasks via few-shot prompting. But on what tasks do language models fall short of average human-rater performance, and are those tasks actually unsolvable by current language models? In this work, we focus on a suite of 23 challenging BIG-Bench tasks which we c

View details →
Negative / Null Result ReportOpen accessComputer Science

Stochastic dynamics of storm surge with stable noise

Joshua Frankie Rayo, Vena Pearl Boñgolan · 2020 · arXiv

The Advanced Circulation (ADCIRC) and Simulating Nearshore Waves (SWAN) coupled model is modified to include a stochastic term in the shallow water equations that represents random external forces from debris carried by surge and short-term local scale atmospheric fluctuations. We added $α$-stable noise, uncorrelated in space and time, in the forcing terms of the coupled model. Inputs to the model are unstructured computational mesh derived from topography and bathymetry, land cover classification, tidal potential constituents and atmospheric forcing. The model simulated surge height of around

View details →
Negative / Null Result ReportOpen accessComputer Science

TSA-WF: Exploring the Effectiveness of Time Series Analysis for Website Fingerprinting

Michael Wrana, Uzma Maroof, Diogo Barradas · 2025 · arXiv

Website fingerprinting (WF) is a technique that allows an eavesdropper to determine the website a target user is accessing by inspecting the metadata associated with the packets she exchanges via some encrypted tunnel, e.g., Tor. Recent WF attacks built using machine learning (and deep learning) process and summarize trace metadata during their feature extraction phases. This methodology leads to predictions that lack information about the instant at which a given website is detected within a (potentially large) network trace comprised of multiple sequential website accesses -- a setting known

View details →
Negative / Null Result ReportOpen accessComputer Science

Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models

Rylan Schaeffer, Joshua Kazdan, Yegor Denisov-Blanch · 2025 · arXiv

Sampling from language models impacts the quality and diversity of outputs, affecting both research and real-world applications. Recently, Nguyen et al. 2024's "Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs" introduced a new sampler called min-p, claiming it achieves superior quality and diversity over established samplers such as basic, top-k, and top-p sampling. The significance of these claims was underscored by the paper's recognition as the 18th highest-scoring submission to ICLR 2025 and selection for an Oral presentation. This paper conducts a comprehensive r

View details →
Negative / Null Result ReportOpen accessComputer Science

Effect of Static vs. Conversational AI-Generated Messages on Colorectal Cancer Screening Intent: a Randomized Controlled Trial

Neil K. R. Sehgal, Manuel Tonneau, Andy Tan et al. · 2025 · arXiv

Large language model (LLM) chatbots show increasing promise in persuasive communication. Yet their real-world utility remains uncertain, particularly in clinical settings where sustained conversations are difficult to scale. In a pre-registered randomized controlled trial, we enrolled 915 U.S. adults (ages 45-75) who had never completed colorectal cancer (CRC) screening. Participants were randomized to: (1) no message control, (2) expert-written patient materials, (3) single AI-generated message, or (4) a motivational interviewing chatbot. All participants were required to remain in their assi

View details →
Negative / Null Result ReportOpen accessComputer Science

CogniPlay: a work-in-progress Human-like model for General Game Playing

Aloïs Rautureau, Éric Piette · 2025 · arXiv

While AI systems have equaled or surpassed human performance in a wide variety of games such as Chess, Go, or Dota 2, describing these systems as truly "human-like" remains far-fetched. Despite their success, they fail to replicate the pattern-based, intuitive decision-making processes observed in human cognition. This paper presents an overview of findings from cognitive psychology and previous efforts to model human-like behavior in artificial agents, discusses their applicability to General Game Playing (GGP) and introduces our work-in-progress model based on these observations: CogniPlay.

View details →
Negative / Null Result ReportOpen accessComputer Science

S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-bit Neural Networks via Guided Distribution Calibration

Zhiqiang Shen, Zechun Liu, Jie Qin et al. · 2021 · arXiv

Previous studies dominantly target at self-supervised learning on real-valued networks and have achieved many promising results. However, on the more challenging binary neural networks (BNNs), this task has not yet been fully explored in the community. In this paper, we focus on this more difficult scenario: learning networks where both weights and activations are binary, meanwhile, without any human annotated labels. We observe that the commonly used contrastive objective is not satisfying on BNNs for competitive accuracy, since the backbone network contains relatively limited capacity and re

View details →
Negative / Null Result ReportOpen accessComputer Science

Peer Grading in a Course on Algorithms and Data Structures: Machine Learning Algorithms do not Improve over Simple Baselines

Mehdi S. M. Sajjadi, Morteza Alamgir, Ulrike von Luxburg · 2015 · arXiv

Peer grading is the process of students reviewing each others' work, such as homework submissions, and has lately become a popular mechanism used in massive open online courses (MOOCs). Intrigued by this idea, we used it in a course on algorithms and data structures at the University of Hamburg. Throughout the whole semester, students repeatedly handed in submissions to exercises, which were then evaluated both by teaching assistants and by a peer grading mechanism, yielding a large dataset of teacher and peer grades. We applied different statistical and machine learning methods to aggregate t

View details →
Negative / Null Result ReportOpen accessComputer Science

Towards precise baryogenesis in the 2HDM$+a$

T. Gent, S. Huber, K. Mimasu et al. · 2025 · arXiv

We perform a detailed investigation of the viable baryogenesis parameter space of a non-minimal Higgs sector consisting of two Higgs doublets and a singlet pseudoscalar (2HDM$+a$). In such a model, an early Universe period of transient CP violation may occur, driven by a nonvanishing vacuum expectation value of the CP-odd scalar $a$. This naturally avoids the stringent electric dipole moment experimental constraints on beyond-the-Standard-Model sources of CP violation. We provide a state-of-art computation of the baryon asymmetry, providing several important improvements over existing baryogen

View details →
Negative / Null Result ReportOpen accessComputer Science

Addressing Climate Action Misperceptions with Generative AI

Miriam Remshard, Yara Kyrychenko, Sander van der Linden et al. · 2026 · arXiv

Mitigating climate change requires behaviour change. However, even climate-concerned individuals often hold misperceptions about which actions most reduce carbon emissions. We recruited 1201 climate-concerned individuals to examine whether discussing climate actions with a large language model (LLM) equipped with climate knowledge and prompted to provide personalised responses would foster more accurate perceptions of the impacts of climate actions and increase willingness to adopt feasible, high-impact behaviours. We compared this to having participants run a web search, have a conversation w

View details →
Negative / Null Result ReportOpen accessComputer Science

On an Improvement over Rényi's Equivocation Bound

Nandakishore Santhi, Alexander Vardy · 2006 · arXiv

We consider the problem of estimating the probability of error in multi-hypothesis testing when MAP criterion is used. This probability, which is also known as the Bayes risk is an important measure in many communication and information theory problems. In general, the exact Bayes risk can be difficult to obtain. Many upper and lower bounds are known in literature. One such upper bound is the equivocation bound due to Rényi which is of great philosophical interest because it connects the Bayes risk to conditional entropy. Here we give a simple derivation for an improved equivocation bound. We

View details →
Negative / Null Result ReportOpen accessComputer Science

Learning to Identify Patients at Risk of Uncontrolled Hypertension Using Electronic Health Records Data

Ramin Mohammadi, Sarthak Jain, Stephen Agboola et al. · 2019 · arXiv

Hypertension is a major risk factor for stroke, cardiovascular disease, and end-stage renal disease, and its prevalence is expected to rise dramatically. Effective hypertension management is thus critical. A particular priority is decreasing the incidence of uncontrolled hypertension. Early identification of patients at risk for uncontrolled hypertension would allow targeted use of personalized, proactive treatments. We develop machine learning models (logistic regression and recurrent neural networks) to stratify patients with respect to the risk of exhibiting uncontrolled hypertension within

View details →
Negative / Null Result ReportOpen accessComputer Science

An Improvement Over Threads Communications on Multi-Core Processors

Reza Fotohi, Mehdi Effatparvar, Fateme Sarkohaki et al. · 2019 · arXiv

Multicore is an integrated circuit chip that uses two or more computational engines (cores) places in a single processor. This new approach is used to split the computational work of a threaded application and spread it over multiple execution cores, so that the computer system can benefits from a better performance and better responsiveness of the system. A thread is a unit of execution inside a process that is created and maintained to execute a set of actions/ instructions. Threads can be implemented differently from an operating system to another, but the operating system is in most cases

View details →
Negative / Null Result ReportOpen accessComputer Science

Trust Over Fear: How Motivation Framing in System Prompts Affects AI Agent Debugging Depth

Wu Ji · 2026 · arXiv

System prompts for AI coding agents increasingly employ motivational framing -- from neutral task descriptions to fear-driven threats -- yet no controlled study has examined whether such framing affects agent behavior. We present two studies investigating how trust-based versus fear-based motivation framing in system prompts influences AI agent debugging performance. In Study 1, we conducted a controlled manual experiment comparing a trust-framed methodology (NoPUA) against an unframed baseline across 9 debugging scenarios using Claude Sonnet 4. Trust-framed agents found 59% more hidden issues

View details →
Negative / Null Result ReportOpen accessComputer Science

Coherent-Classical Estimation versus Purely-Classical Estimation for Linear Quantum Systems

Shibdas Roy, Ian R. Petersen, Elanor H. Huntington · 2014 · arXiv

We consider a coherent-classical estimation scheme for a class of linear quantum systems. It comprises an estimator that is a mixed quantum-classical system without involving coherent feedback. The estimator yields a classical estimate of a variable for the quantum plant. We demonstrate that for a passive plant that can be characterized by annihilation operators only, such coherent-classical estimation provides no improvement over purely-classical estimation. An example is also given which shows that if the plant is not assumed to be an annihilation operator only quantum system, it is possible

View details →
Negative / Null Result ReportOpen accessComputer Science

Benchmarking Pretrained Molecular Embedding Models For Molecular Representation Learning

Mateusz Praski, Jakub Adamczyk, Wojciech Czech · 2025 · arXiv

Pretrained neural networks have attracted significant interest in chemistry and small molecule drug design. Embeddings from these models are widely used for molecular property prediction, virtual screening, and small data learning in molecular chemistry. This study presents the most extensive comparison of such models to date, evaluating 25 models across 25 datasets. Under a fair comparison framework, we assess models spanning various modalities, architectures, and pretraining strategies. Using a dedicated hierarchical Bayesian statistical testing model, we arrive at a surprising result: nearl

View details →
Negative / Null Result ReportOpen accessComputer Science

Covering models of the asymmetric quantum Rabi model: $η$-shifted non-commutative harmonic oscillators

Cid Reyes-Bustos, Masato Wakayama · 2022 · arXiv

The non-commutative harmonic oscillator (NCHO) is a matrix valued differential operator originally introduced as a generalization of the quantum harmonic oscillator having a weaker $\mathfrak{sl}_2(\mathbb{R})$-symmetry. The spectrum of the NCHO has remarkable properties, including the presence of number theoretical structures such as modular forms, elliptic curves and Eichler cohomology observed in the special values of the associated spectral zeta function. In addition, the Heun ODE picture of the eigenvalue problem of the NCHO reveals a connection with the quantum Rabi model (QRM), a fundam

View details →
Negative / Null Result ReportOpen accessComputer Science

One-at-a-time: A Meta-Learning Recommender-System for Recommendation-Algorithm Selection on Micro Level

Andrew Collins, Dominika Tkaczyk, Joeran Beel · 2018 · arXiv

The effectiveness of recommendation algorithms is typically assessed with evaluation metrics such as root mean square error, F1, or click through rates, calculated over entire datasets. The best algorithm is typically chosen based on these overall metrics. However, there is no single-best algorithm for all users, items, and contexts. Choosing a single algorithm based on overall evaluation results is not optimal. In this paper, we propose a meta-learning-based approach to recommendation, which aims to select the best algorithm for each user-item pair. We evaluate our approach using the MovieLen

View details →
Negative / Null Result ReportOpen accessComputer Science

The Capacity of MIMO Channels with Per-Antenna Power Constraint

Mai Vu · 2011 · arXiv

We establish the optimal input signaling and the capacity of MIMO channels under per-antenna power constraint. While admitting a linear eigenbeam structure, the optimal input is no longer diagonalizable by the channel right singular vectors as with sum power constraint. We formulate the capacity optimization as an SDP problem and solve in closed-form the optimal input covariance as a function of the dual variable. We then design an efficient algorithm to find this optimal input signaling for all channel sizes. The proposed algorithm allows for straightforward implementation in practical system

View details →
Negative / Null Result ReportOpen accessComputer Science

Loose LIPS Sink Ships: Asking Questions in Battleship with Language-Informed Program Sampling

Gabriel Grand, Valerio Pepe, Jacob Andreas et al. · 2024 · arXiv

Questions combine our mastery of language with our remarkable facility for reasoning about uncertainty. How do people navigate vast hypothesis spaces to pose informative questions given limited cognitive resources? We study these tradeoffs in a classic grounded question-asking task based on the board game Battleship. Our language-informed program sampling (LIPS) model uses large language models (LLMs) to generate natural language questions, translate them into symbolic programs, and evaluate their expected information gain. We find that with a surprisingly modest resource budget, this simple M

View details →
Negative / Null Result ReportOpen accessComputer Science

On the Effectiveness of Mode Exploration in Bayesian Model Averaging for Neural Networks

John T. Holodnak, Allan B. Wollaber · 2021 · arXiv

Multiple techniques for producing calibrated predictive probabilities using deep neural networks in supervised learning settings have emerged that leverage approaches to ensemble diverse solutions discovered during cyclic training or training from multiple random starting points (deep ensembles). However, only a limited amount of work has investigated the utility of exploring the local region around each diverse solution (posterior mode). Using three well-known deep architectures on the CIFAR-10 dataset, we evaluate several simple methods for exploring local regions of the weight space with re

View details →
Negative / Null Result ReportOpen accessComputer Science

Towards Single Exponential Time for Temporal and Spatial Reasoning: A Study via Redundancy and Dynamic Programming

Victor Lagerkvist, Johanna Groven, Leif Eriksson · 2026 · arXiv

The region connection calculus ($RCC$) and Allen's interval algebra ($IA$) are two well-known NP-hard spatial-temporal qualitative reasoning problems. They are solvable in $2^{O(n \log n)}$ time, where $n$ is the number of variables, and $IA$ is additionally known to be solvable in $o(n)^n$ time. However, no improvement over exhaustive search is known for $RCC$, and if they are also solvable in single exponential time $2^{O(n)}$ is unknown. We investigate multiple avenues towards reaching such bounds. First, we show that branching is insufficient since there are too many non-redundant constrai

View details →
Negative / Null Result ReportOpen accessComputer Science

Honey, I shrunk the scientist -- Evaluating 2D, 3D, and VR interfaces for navigating samples under the microscope

Jan Tiemann, Matthew McGinity, Ulrik Günther · 2026 · arXiv

In contemporary biology and medicine, 3D microscopy is one of the most widely-used techniques for imaging and manipulation of various kinds of samples. Navigating such a micrometer-sized, 3-dimensional sample under the microscope -- e.g. to find relevant imaging regions -- can pose a tedious challenge for the experimenter. In this paper, we examine whether 2D desktop, 3D desktop, or Virtual Reality (VR) interfaces provide the best user experience and performance for the exploration of 3D samples. We invited 12 skilled microscope operators to perform two different exploration tasks in 2D, 3D an

View details →