e-ISSN: Pending

Browse the failure-mode index

744 real negative results, null findings, and replication failures in Computer Science. Search the index →

WASTE indexes published research — it does not host or republish full papers. Each entry is a metadata record compiled from open scholarly databases; the abstract is shown in full only where the paper is openly licensed, otherwise a short excerpt under fair use. Classifications are automated and approximate.

Failed Experiment ReportOpen accessComputer Science

When Good Components Go Bad: Formally Secure Compilation Despite Dynamic Compromise

Carmine Abate, Arthur Azevedo de Amorim, Roberto Blanco et al. · 2018 · arXiv

We propose a new formal criterion for evaluating secure compilation schemes for unsafe languages, expressing end-to-end security guarantees for software components that may become compromised after encountering undefined behavior---for example, by accessing an array out of bounds. Our criterion is the first to model dynamic compromise in a system of mutually distrustful components with clearly specified privileges. It articulates how each component should be protected from all the others---in particular, from components that have encountered undefined behavior and become compromised. Each comp

View details →
Negative / Null Result ReportOpen accessComputer Science

Looking for a Handsome Carpenter! Debiasing GPT-3 Job Advertisements

Conrad Borchers, Dalia Sara Gala, Benjamin Gilburt et al. · 2022 · arXiv

The growing capability and availability of generative language models has enabled a wide range of new downstream tasks. Academic research has identified, quantified and mitigated biases present in language models but is rarely tailored to downstream tasks where wider impact on individuals and society can be felt. In this work, we leverage one popular generative language model, GPT-3, with the goal of writing unbiased and realistic job advertisements. We first assess the bias and realism of zero-shot generated advertisements and compare them to real-world advertisements. We then evaluate prompt

View details →
Negative / Null Result ReportOpen accessComputer Science

Style Over Substance: Evaluation Biases for Large Language Models

Minghao Wu, Alham Fikri Aji · 2023 · arXiv

As large language models (LLMs) continue to advance, accurately and comprehensively evaluating their performance becomes increasingly challenging. Ranking the relative performance of LLMs based on Elo ratings, according to human judgment, is gaining more popularity. However, the extent to which humans and LLMs are capable evaluators remains uncertain. This study investigates the behavior of crowd-sourced and expert annotators, as well as LLMs, when comparing outputs from different models. To achieve this, we curate a dataset of intentionally flawed machine-generated answers. Our findings revea

View details →
Negative / Null Result ReportOpen accessComputer Science

Classical simulation of bosonic linear-optical random circuits beyond linear light cone

Changhun Oh, Youngrong Lim, Bill Fefferman et al. · 2021 · arXiv

Sampling from probability distributions of quantum circuits is a fundamentally and practically important task which can be used to demonstrate quantum supremacy using noisy intermediate-scale quantum devices. In the present work, we examine classical simulability of sampling from the output photon-number distribution of linear-optical circuits composed of random beam splitters with equally distributed squeezed vacuum states and single-photon states input. We provide efficient classical algorithms to simulate linear-optical random circuits and show that the algorithms' error is exponentially sm

View details →
Negative / Null Result ReportOpen accessComputer Science

The ATLAS discovery potential for a heavy charged Higgs boson in gg->tbH^{+-} with H^{+-}->tb

K. A. Assamagan, N. Gollub · 2004 · arXiv

The feasibility of detecting a heavy charged Higgs boson, m(H^{+-})>m(t)+m(b), decaying in the H^{+-}->tb channel is studied with the fast simulation of the ATLAS detector. We study the gg->H^{+-}tb production process at the LHC which together with the aforementioned decay channel leads to four b-quarks in the final state. The whole production and decay chain reads gg->H^{+-}tb->t\bar{t}b\bar{b}->b\bar{b}b\bar{b}lν\bar{q}q'. Combinatorial background is a major difficulty in this multi-jet environment but can be overcome by employing multivariate techniques in the event reconstruction. Requirin

View details →
Negative / Null Result ReportOpen accessComputer Science

Charm-quark Yukawa Coupling in $h\rightarrow c\bar{c}γ$ at LHC

Tao Han, Benjamin Nachman, Xing Wang · 2018 · arXiv

It is extremely challenging to probe the charm-quark Yukawa coupling at hadron colliders primarily due to the large Standard Model (SM) background (including $h\to b\bar b$) and the lack of an effective trigger for the signal $h\to c\bar c$. We examine the feasibility of probing this coupling at the LHC via a Higgs radiative decay $h\rightarrow c\bar{c}γ$. The existence of an additional photon in the final state may help for the signal identification and background suppression. Adopting a refined triggering strategy and utilizing basic machine learning, we find that a coupling limit of about 8

View details →
Null DatasetOpen accessComputer Science

Data Transmissions in Blockchain enabled AGVs

Junhan Liu, Shile Liu, Chenming Xu et al. · 2023 · arXiv

Automated Guided Vehicles (AGVs) operate in synergy to execute specific tasks. These vehicles exchange information to ensure seamless collaboration, prevent collisions, and eliminate task redundancy. The advent of blockchain technology offers a promising avenue for establishing a secure and dependable communication infrastructure for AGVs. Nonetheless, it becomes imperative for AGVs to adopt efficient data transmission methodologies, especially when interacting with the dynamic nature of blockchain infrastructure where data undergoes frequent modifications. In the present study, we introduce a

View details →
Negative / Null Result ReportOpen accessComputer Science

NLP Techniques for Water Quality Analysis in Social Media Content

Muhammad Asif Ayub, Khubaib Ahmad, Kashif Ahmad et al. · 2021 · arXiv

This paper presents our contributions to the MediaEval 2021 task namely "WaterMM: Water Quality in Social Multimedia". The task aims at analyzing social media posts relevant to water quality with particular focus on the aspects like watercolor, smell, taste, and related illnesses. To this aim, a multimodal dataset containing both textual and visual information along with meta-data is provided. Considering the quality and quantity of available content, we mainly focus on textual information by employing three different models individually and jointly in a late-fusion manner. These models includ

View details →
Negative / Null Result ReportOpen accessComputer Science

Diving Deep into Context-Aware Neural Machine Translation

Jingjing Huo, Christian Herold, Yingbo Gao et al. · 2020 · arXiv

Context-aware neural machine translation (NMT) is a promising direction to improve the translation quality by making use of the additional context, e.g., document-level translation, or having meta-information. Although there exist various architectures and analyses, the effectiveness of different context-aware NMT models is not well explored yet. This paper analyzes the performance of document-level NMT models on four diverse domains with a varied amount of parallel document-level bilingual data. We conduct a comprehensive set of experiments to investigate the impact of document-level NMT. We

View details →
Negative / Null Result ReportOpen accessComputer Science

RECAP: Regression Evaluation for Continual Adaptation of Prompts

Harsh Deshpande, Kushal Chawla, Sangwoo Cho et al. · 2026 · arXiv

Production agentic systems routinely face evolving constraints and must comply from the very next interaction. Scenarios like a tool-call notification changing a compliance threshold or a policy update adding disclosure requirements fit this criteria, having close to no room for errors in production. This proactive adaptation setting is common in deployment, but absent from current benchmarks, which assume either static constraint sets or reactive protocols with evaluation feedback. We introduce RECAP, a benchmark that measures continual-learning phenomena (forgetting, regression, forward tran

View details →
Negative / Null Result ReportOpen accessComputer Science

An Eye on Clinical BERT: Investigating Language Model Generalization for Diabetic Eye Disease Phenotyping

Keith Harrigian, Tina Tang, Anthony Gonzales et al. · 2023 · arXiv

Diabetic eye disease is a major cause of blindness worldwide. The ability to monitor relevant clinical trajectories and detect lapses in care is critical to managing the disease and preventing blindness. Alas, much of the information necessary to support these goals is found only in the free text of the electronic medical record. To fill this information gap, we introduce a system for extracting evidence from clinical text of 19 clinical concepts related to diabetic eye disease and inferring relevant attributes for each. In developing this ophthalmology phenotyping system, we are also afforded

View details →
Negative / Null Result ReportOpen accessComputer Science

Acceleration of particles and shells by Reissner-Nordström naked singularities

Mandar Patil, Pankaj S. Joshi, Masashi Kimura et al. · 2011 · arXiv

We explore the Reissner-Nordström naked singularities with a charge $Q$ larger than its mass $M$ from the perspective of the particle acceleration. We first consider a collision between two test particles following the radial geodesics in the Reissner-Nordström naked singular geometry. An initially radially ingoing particle turns back due to the repulsive effect of gravity in the vicinity of naked singularity. Such a particle then collides with an another radially ingoing particle. We show that the center of mass energy of collision taking place at $r \approx M$ is unbound, in the limit where

View details →
Negative / Null Result ReportOpen accessComputer Science

Metabook: A Mobile-to-Headset Pipeline for 3D Story Book Creation in Augmented Reality

Yibo Wang, Yuanyuan Mao, Lik-Hang Lee et al. · 2024 · arXiv

The AR 3D book has shown significant potential in enhancing students' learning outcomes. However, the creation process of 3D books requires a significant investment of time, effort, and specialized skills. Thus, in this paper, we first conduct a three-day workshop investigating how AI can support the automated creation of 3D books. Informed by the design insights derived from the workshop, we developed Metabook, a system that enables even novice users to create 3D books from text automatically. To our knowledge, Metabook is the first system to offer end-to-end 3D book generation. A follow-up s

View details →
Negative / Null Result ReportOpen accessComputer Science

EdgeFlow: Edge-Map Augmented VLM-Based Flowchart Processing for Industrial Requirements Engineering

Zhifei Dou, Shabnam Hassani, Ou Wei · 2026 · arXiv

Flowcharts are widely used in industrial requirements, but usually remain embedded as static images. Vision Language Models (VLMs) show promise in the conversion of these flowcharts into machine-readable models for RE activities, yet, when directly applied to flowchart conversion, they often fail on topology-critical visual details. To address this, we propose EdgeFlow that augments a VLM's original input with a deterministically extracted Canny edge map-acting as a structural prior-to improve flowchart-to-Mermaid conversion, without requiring annotated training data or domain-specific model f

View details →
Negative / Null Result ReportOpen accessComputer Science

Measuring the Effect of Discourse Relations on Blog Summarization

Shamima Mithun, Leila Kosseim · 2017 · arXiv

The work presented in this paper attempts to evaluate and quantify the use of discourse relations in the context of blog summarization and compare their use to more traditional and factual texts. Specifically, we measured the usefulness of 6 discourse relations - namely comparison, contingency, illustration, attribution, topic-opinion, and attributive for the task of text summarization from blogs. We have evaluated the effect of each relation using the TAC 2008 opinion summarization dataset and compared them with the results with the DUC 2007 dataset. The results show that in both textual genr

View details →
Negative / Null Result ReportOpen accessComputer Science

Tagset Reduction Without Information Loss

Thorsten Brants · 1995 · arXiv

A technique for reducing a tagset used for n-gram part-of-speech disambiguation is introduced and evaluated in an experiment. The technique ensures that all information that is provided by the original tagset can be restored from the reduced one. This is crucial, since we are interested in the linguistically motivated tags for part-of-speech disambiguation. The reduced tagset needs fewer parameters for its statistical model and allows more accurate parameter estimation. Additionally, there is a slight but not significant improvement of tagging accuracy.

View details →
Negative / Null Result ReportOpen accessComputer Science

Reliability and User-Plane Latency Analysis of mmWave Massive MIMO for Grant-Free URLLC Applications

Joao V. C. Evangelista, Georges Kaddoum, Zeeshan Sattar · 2021 · arXiv

5G cellular networks are designed to support a new range of applications not supported by previous standards. Among these, ultra-reliable low-latency communication (URLLC) applications are arguably the most challenging. URLLC service requires the user equipment (UE) to be able to transmit its data under strict latency constraints with high reliability. To address these requirements, new technologies, such as mini-slots, semi-persistent scheduling and grant-free access were introduced in 5G standards. In this work, we formulate a spatiotemporal mathematical model to evaluate the user-plane late

View details →
Negative / Null Result ReportOpen accessComputer Science

Intersecting non-SUSY branes and closed string tachyon condensation

J. X. Lu, Shibaji Roy, Zhao-Long Wang et al. · 2007 · arXiv

Following \cite{Bai:2006vv} we here consider the supergravity solutions representing the charged non-supersymmetric $p$-brane (for $1\leq p \leq 6$) intersecting with chargeless non-supersymmetric 1-brane and 0-brane of type II string theories. We show how these solutions nicely interpolate between black $p$-branes and the Kaluza-Klein "bubble of nothing" (BON) by continuously varying some parameters characterizing the solutions from one set of values to another. By performing a time symmetric general bubble initial data analysis, we show that the interpolation implies a possible transition fr

View details →
Failed Experiment ReportOpen accessComputer Science

Nearly ETH-Tight Algorithms for Planar Steiner Tree with Terminals on Few Faces

Sándor Kisfaludi-Bak, Jesper Nederlof, Erik Jan van Leeuwen · 2018 · arXiv

The Planar Steiner Tree problem is one of the most fundamental NP-complete problems as it models many network design problems. Recall that an instance of this problem consists of a graph with edge weights, and a subset of vertices (often called terminals); the goal is to find a subtree of the graph of minimum total weight that connects all terminals. A seminal paper by Erickson et al. [Math. Oper. Res., 1987] considers instances where the underlying graph is planar and all terminals can be covered by the boundary of $k$ faces. Erickson et al. show that the problem can be solved by an algorithm

View details →
Negative / Null Result ReportOpen accessComputer Science

Hidden Meanings in Plain Sight: RebusBench for Evaluating Cognitive Visual Reasoning

Seyed Amir Kasaei, Arash Marioriyad, Mahbod Khaleti et al. · 2026 · arXiv

Large Vision-Language Models (LVLMs) have achieved remarkable proficiency in explicit visual recognition, effectively describing what is directly visible in an image. However, a critical cognitive gap emerges when the visual input serves only as a clue rather than the answer. We identify that current models struggle with the complex, multi-step reasoning required to solve problems where information is not explicitly depicted. Successfully solving a rebus puzzle requires a distinct cognitive workflow: the model must extract visual and textual attributes, retrieve linguistic prior knowledge (suc

View details →
Negative / Null Result ReportOpen accessComputer Science

When Medical Imaging Met Self-Attention: A Love Story That Didn't Quite Work Out

Tristan Piater, Niklas Penzel, Gideon Stein et al. · 2024 · arXiv

A substantial body of research has focused on developing systems that assist medical professionals during labor-intensive early screening processes, many based on convolutional deep-learning architectures. Recently, multiple studies explored the application of so-called self-attention mechanisms in the vision domain. These studies often report empirical improvements over fully convolutional approaches on various datasets and tasks. To evaluate this trend for medical imaging, we extend two widely adopted convolutional architectures with different self-attention variants on two different medical

View details →
Negative / Null Result ReportOpen accessComputer Science

Benchmarking Language Model Creativity: A Case Study on Code Generation

Yining Lu, Dixuan Wang, Tianjian Li et al. · 2024 · arXiv

As LLMs become increasingly prevalent, it is interesting to consider how ``creative'' these models can be. From cognitive science, creativity consists of at least two key characteristics: \emph{convergent} thinking (purposefulness to achieve a given goal) and \emph{divergent} thinking (adaptability to explore new environments or constraints) \citep{runco2003critical}. In this work, we introduce a framework for quantifying LLM creativity that incorporates the two design ingredients: (1) We introduce DENIAL PROMPTING which pushes LLMs to develop more creative solutions to a given problem by incr

View details →
Negative / Null Result ReportOpen accessComputer Science

Association Is Not Similarity: Learning Corpus-Specific Associations for Multi-Hop Retrieval

Jason Dury · 2026 · arXiv

Dense retrieval systems rank passages by embedding similarity to a query, but multi-hop questions require passages that are associatively related through shared reasoning chains. We introduce Association-Augmented Retrieval (AAR), a lightweight transductive reranking method that trains a small MLP (4.2M parameters) to learn associative relationships between passages in embedding space using contrastive learning on co-occurrence annotations. At inference time, AAR reranks an initial dense retrieval candidate set using bi-directional association scoring. On HotpotQA, AAR improves passage Recall@

View details →
Negative / Null Result ReportOpen accessComputer Science

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study

Xiaolong Jin, Xuandong Zhao, Wenbo Guo et al. · 2026 · arXiv

Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in large language model reasoning, but relies on ground-truth supervision that is costly or infeasible, especially in coding tasks. Recent work addresses this by deriving rewards from a model's own signals, such as majority voting or confidence-based scores, achieving notable success on mathematical reasoning benchmarks. However, code generation poses distinct challenges: programs are structurally complex, semantically equivalent solutions may differ syntactically, and verification typically requires executio

View details →
Negative / Null Result ReportOpen accessComputer Science

Significant Improvements over the State of the Art? A Case Study of the MS MARCO Document Ranking Leaderboard

Jimmy Lin, Daniel Campos, Nick Craswell et al. · 2021 · arXiv

Leaderboards are a ubiquitous part of modern research in applied machine learning. By design, they sort entries into some linear order, where the top-scoring entry is recognized as the "state of the art" (SOTA). Due to the rapid progress being made in information retrieval today, particularly with neural models, the top entry in a leaderboard is replaced with some regularity. These are touted as improvements in the state of the art. Such pronouncements, however, are almost never qualified with significance testing. In the context of the MS MARCO document ranking leaderboard, we pose a specific

View details →
Negative / Null Result ReportOpen accessComputer Science

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models

Jason Z Wang · 2026 · arXiv

We introduce MIRROR, a benchmark comprising eight experiments across four metacognitive levels that evaluates whether large language models can use self-knowledge to make better decisions. We evaluate 16 models from 8 labs across approximately 250,000 evaluation instances using five independent behavioral measurement channels. Core experiments are run across the full model roster; experiments with specialized infrastructure requirements report explicitly marked model subsets. We find two phenomena with direct implications for agentic deployment: (1) compositional self-prediction fails universa

View details →
Negative / Null Result ReportOpen accessComputer Science

MaScQA: A Question Answering Dataset for Investigating Materials Science Knowledge of Large Language Models

Mohd Zaki, Jayadeva, Mausam et al. · 2023 · arXiv

Information extraction and textual comprehension from materials literature are vital for developing an exhaustive knowledge base that enables accelerated materials discovery. Language models have demonstrated their capability to answer domain-specific questions and retrieve information from knowledge bases. However, there are no benchmark datasets in the materials domain that can evaluate the understanding of the key concepts by these language models. In this work, we curate a dataset of 650 challenging questions from the materials domain that require the knowledge and skills of a materials st

View details →
Negative / Null Result ReportOpen accessComputer Science

Refined Continuous Control of DDPG Actors via Parametrised Activation

Mohammed Hossny, Julie Iskander, Mohammed Attia et al. · 2020 · arXiv

In this paper, we propose enhancing actor-critic reinforcement learning agents by parameterising the final actor layer which produces the actions in order to accommodate the behaviour discrepancy of different actuators, under different load conditions during interaction with the environment. We propose branching the action producing layer in the actor to learn the tuning parameter controlling the activation layer (e.g. Tanh and Sigmoid). The learned parameters are then used to create tailored activation functions for each actuator. We ran experiments on three OpenAI Gym environments, i.e. Pend

View details →
Negative / Null Result ReportOpen accessComputer Science

From Co-Design to Metacognitive Laziness: Evaluating Generative AI in Vocational Education

Amir Yunus, Peng Rend Gay, Oon Teng Lee · 2025 · arXiv

This study examines the development and deployment of a Generative AI proof-of-concept (POC) designed to support lecturers in a vocational education setting in Singapore. Employing a user-centred, mixed-methods design process, we co-developed an AI chatbot with lecturers to address recurring instructional challenges during exam preparation, specifically managing repetitive questions and scaling feedback delivery. The POC achieved its primary operational goals: lecturers reported streamlined workflows, reduced cognitive load, and observed improved student confidence in navigating course content

View details →
Negative / Null Result ReportOpen accessComputer Science

Single- and Multi-Task Architectures for Tool Presence Detection Challenge at M2CAI 2016

Andru P. Twinanda, Didier Mutter, Jacques Marescaux et al. · 2016 · arXiv

The tool presence detection challenge at M2CAI 2016 consists of identifying the presence/absence of seven surgical tools in the images of cholecystectomy videos. Here, we propose to use deep architectures that are based on our previous work where we presented several architectures to perform multiple recognition tasks on laparoscopic videos. In this technical report, we present the tool presence detection results using two architectures: (1) a single-task architecture designed to perform solely the tool presence detection task and (2) a multi-task architecture designed to perform jointly phase

View details →