Negative / Null Result ReportOpen accessComputer Science
Yunquan Dong, Zhengchuan Chen, Shanyun Liu et al. · 2018 · arXiv
We consider an M/M/1 update-and-decide system where Poisson distributed decisions are made based on the received updates. We propose to characterize the freshness of the received updates at decision epochs with Age upon Decisions (AuD). Under the first-come-first-served policy (FCFS), the closed form average AuD is derived. We show that the average AuD of the system is determined by the arrival rate and the service rate, and is independent of the decision rate. Thus, merely increasing the decision rate does not improve the timeliness of decisions. Nevertheless, increasing the arrival rate and
View details →Negative / Null Result ReportOpen accessComputer Science
G. Chiribella, G. M. D'Ariano, P. Perinotti · 2008 · arXiv
A sequential network of quantum operations is efficiently described by its quantum comb, a non-negative operator with suitable normalization constraints. Here we analyze the case of networks enjoying symmetry with respect to the action of a given group of physical transformations, introducing the notion of covariant combs and testers, and proving the basic structure theorems for these objects. As an application, we discuss the optimal alignment of reference frames (without pre-established common references) with multiple rounds of quantum communication, showing that i) allowing an arbitrary am
View details →Negative / Null Result ReportOpen accessComputer Science
Paul Schneider, Amalie Schramm · 2025 · arXiv
Structured deliberation has been found to improve the performance of human forecasters. This study investigates whether a similar intervention, i.e. allowing LLMs to review each other's forecasts before updating, can improve accuracy in large language models (GPT-5, Claude Sonnet 4.5, Gemini Pro 2.5). Using 202 resolved binary questions from the Metaculus Q2 2025 AI Forecasting Tournament, accuracy was assessed across four scenarios: (1) diverse models with distributed information, (2) diverse models with shared information, (3) homogeneous models with distributed information, and (4) homogene
View details →Negative / Null Result ReportOpen accessComputer Science
Pierre Epron, Adrien Coulet, Mehwish Alam · 2026 · arXiv
Despite their strong linguistic capabilities, Large Language Models (LLMs) are computationally demanding and require substantial resources for fine-tuning, which is unadapted to privacy and budget constraints of many healthcare settings. To address this, we present an experimental analysis focused on Biomedical Named Entity Recognition using lightweight LLMs, we evaluate the impact of different output formats on model performance. The results reveal that lightweight LLMs can achieve competitive performance compared to the larger models, highlighting their potential as lightweight yet effective
View details →Negative / Null Result ReportOpen accessComputer Science
Masahito Hayashi · 2025 · arXiv
We consider estimation of unknown unitary operation when the set of possible unitary operations is given by a projective unitary representation of a compact group. We show that neither indefinite causal order strategy nor adaptive strategy improves the performance of this estimation when error function satisfies group covariance. That is, the optimal parallel strategy gives the optimal performance even under indefinite causal order strategy and adaptive strategy. To study this problem, we newly introduce the concept of generalized positive operator valued measure (GPOVM), and its convariance c
View details →Negative / Null Result ReportOpen accessComputer Science
Amir Homayounirad, Enrico Liscio, Tong Wang et al. · 2025 · arXiv
Aggregating multiple annotations into a single ground truth label may hide valuable insights into annotator disagreement, particularly in tasks where subjectivity plays a crucial role. In this work, we explore methods for identifying subjectivity in recognizing the human values that motivate arguments. We evaluate two main approaches: inferring subjectivity through value prediction vs. directly identifying subjectivity. Our experiments show that direct subjectivity identification significantly improves the model performance of flagging subjective arguments. Furthermore, combining contrastive l
View details →Negative / Null Result ReportOpen accessComputer Science
Xirui Li, Ming Li, Tianyi Zhou · 2026 · arXiv
Reinforcement learning (RL) with verifiable rewards has become a standard post-training stage for boosting visual reasoning in vision-language models, yet it remains unclear what capabilities RL actually improves compared with supervised fine-tuning as cold-start initialization (IN). End-to-end benchmark gains conflate multiple factors, making it difficult to attribute improvements to specific skills. To bridge the gap, we propose a Frankenstein-style analysis framework including: (i) functional localization via causal probing; (ii) update characterization via parameter comparison; and (iii) t
View details →Negative / Null Result ReportOpen accessComputer Science
Matthew Aitchison · 2019 · arXiv
Although reinforcement learning has made great strides recently, a continuing limitation is that it requires an extremely high number of interactions with the environment. In this paper, we explore the effectiveness of reusing experience from the experience replay buffer in the Deep Q-Learning algorithm. We test the effectiveness of applying learning update steps multiple times per environmental step in the VizDoom environment and show first, this requires a change in the learning rate, and second that it does not improve the performance of the agent. Furthermore, we show that updating less fr
View details →Negative / Null Result ReportOpen accessComputer Science
John S. Van Dyke, Zackary White, Gregory Quiroz · 2024 · arXiv
Zero-noise extrapolation (ZNE), a technique to estimate quantum circuit expectation values through noise scaling and extrapolation, is well-studied in the context of quantum computing. We examine the applicability of ZNE to the field of quantum sensing. Focusing on the problem of DC magnetometry using the Ramsey protocol, we show that the sensitivity (in the sense of the minimum detectable signal) does not improve upon using ZNE in the slope detection scheme. On the other hand, signals of sufficiently large magnitude can be estimated more accurately. Our results are robust across various noise
View details →Negative / Null Result ReportOpen accessComputer Science
Baruch Lubinsky, Bekir Genc, Tshilidzi Marwala · 2008 · arXiv
Neural networks are powerful tools for classification and regression in static environments. This paper describes a technique for creating an ensemble of neural networks that adapts dynamically to changing conditions. The model separates the input space into four regions and each network is given a weight in each region based on its performance on samples from that region. The ensemble adapts dynamically by constantly adjusting these weights based on the current performance of the networks. The data set used is a collection of financial indicators with the goal of predicting the future platinu
View details →Negative / Null Result ReportOpen accessComputer Science
Damian Ruck, R. Alexander Bentley, Alberto Acerbi et al. · 2017 · arXiv
Here we test Neutral models against the evolution of English word frequency and vocabulary at the population scale, as recorded in annual word frequencies from three centuries of English language books. Against these data, we test both static and dynamic predictions of two neutral models, including the relation between corpus size and vocabulary size, frequency distributions, and turnover within those frequency distributions. Although a commonly used Neutral model fails to replicate all these emergent properties at once, we find that modified two-stage Neutral model does replicate the static a
View details →Negative / Null Result ReportOpen accessComputer Science
Zichao Wang, Alexa Siu · 2026 · arXiv
Large language models (LLMs) have shown strong performance on standardized social science instruments, but their value for product discovery remains unclear. We investigate whether interview-informed generative agents can simulate user responses in concept testing scenarios. Using in-depth workflow interviews with knowledge workers, we created personalized agents and compared their evaluations of novel AI concepts against the same participants' responses. Our results show that agents are distribution-calibrated but identity-imprecise: they fail to replicate the specific individual they are gro
View details →Negative / Null Result ReportOpen accessComputer Science
Roland S. Zimmermann, Thomas Klein, Wieland Brendel · 2023 · arXiv
In light of the recent widespread adoption of AI systems, understanding the internal information processing of neural networks has become increasingly critical. Most recently, machine vision has seen remarkable progress by scaling neural networks to unprecedented levels in dataset and model size. We here ask whether this extraordinary increase in scale also positively impacts the field of mechanistic interpretability. In other words, has our understanding of the inner workings of scaled neural networks improved as well? We use a psychophysical paradigm to quantify one form of mechanistic inter
View details →Negative / Null Result ReportOpen accessComputer Science
A. Hietanen, K. Kajantie, M. Laine et al. · 2008 · arXiv
We update Monte Carlo simulations of the three-dimensional SU(3) + adjoint Higgs theory, by extrapolating carefully to the infinite volume and continuum limits, in order to estimate the contribution of the infrared modes to the pressure of hot QCD. The sum of infrared contributions beyond the known 4-loop order turns out to be a smooth function, of a reasonable magnitude and specific sign. Unfortunately, adding this function to the known 4-loop terms does not improve the match to four-dimensional lattice data, in spite of the fact that other quantities, such as correlation lengths, spatial str
View details →Negative / Null Result ReportOpen accessComputer Science
Thomas Vogel, Chinh Tran, Lars Grunske · 2019 · arXiv
In search-based software engineering we often use popular heuristics with default configurations, which typically lead to suboptimal results, or we perform experiments to identify configurations on a trial-and-error basis, which may lead to better results for a specific problem. To obtain better results while avoiding trial-and-error experiments, a fitness landscape analysis is helpful in understanding the search problem, and making an informed decision about the heuristics. In this paper, we investigate the search problem of test suite generation for mobile applications (apps) using SAPIENZ w
View details →Negative / Null Result ReportOpen accessComputer Science
Shintaro Sakai, Jisun An, Migyeong Kang et al. · 2025 · arXiv
Prior clinical psychology research shows that Western individuals with depression tend to report psychological symptoms, while Eastern individuals report somatic ones. We test whether Large Language Models (LLMs), which are increasingly used in mental health, reproduce these cultural patterns by prompting them with Western or Eastern personas. Results show that LLMs largely fail to replicate the patterns when prompted in English, though prompting in major Eastern languages (i.e., Chinese, Japanese, and Hindi) improves alignment in several configurations. Our analysis pinpoints two key reasons
View details →Negative / Null Result ReportOpen accessComputer Science
Shufan Wang, Yixiao Song, Andrew Drozdov et al. · 2023 · arXiv
In this paper, we study the generation quality of interpolation-based retrieval-augmented language models (LMs). These methods, best exemplified by the KNN-LM, interpolate the LM's predicted distribution of the next word with a distribution formed from the most relevant retrievals for a given prefix. While the KNN-LM and related methods yield impressive decreases in perplexity, we discover that they do not exhibit corresponding improvements in open-ended generation quality, as measured by both automatic evaluation metrics (e.g., MAUVE) and human evaluations. Digging deeper, we find that interp
View details →Negative / Null Result ReportOpen accessComputer Science
Ryoutaro Watanabe · 2017 · arXiv
We study possible new physics (NP) effects on $B_c \to J/ψτ\barν$, which has been recently measured at LHCb as the ratio of $R_{J/ψ} = \mathcal B(B_c \to J/ψτ\barν)/\mathcal B(B_c \to J/ψμ\barν)$. Combining it with the long-standing $R_{D^{(*)}}$ measurements, in which the discrepancy with the prediction of the standard model is present, we find possible solutions to the anomaly by several NP types. Then, we see that adding the $R_{J/ψ}$ measurement does not improve NP fit to data, but the NP scenarios still give better $χ^2$ than the SM. We also investigate indirect NP constraints from the li
View details →Negative / Null Result ReportOpen accessComputer Science
Niek Beckers, Edwin van Asseldonk, Herman van der Kooij · 2020 · arXiv
Haptic interaction between two humans, for example, a physiotherapist assisting a patient regaining the ability to grasp a cup, likely facilitates motor skill acquisition. Haptic human-human interaction has been shown to enhance individual performance improvement in a tracking task with a visuomotor rotation perturbation. These results are remarkable given that haptically assisting or guiding an individual rarely benefits their individual improvement when the assistance is removed. We, therefore, replicated a study that reported that haptic interaction between humans was beneficial for individ
View details →Negative / Null Result ReportOpen accessComputer Science
Chen Cai, Yusu Wang · 2020 · arXiv
Graph Neural Networks (GNNs) have achieved a lot of success on graph-structured data. However, it is observed that the performance of graph neural networks does not improve as the number of layers increases. This effect, known as over-smoothing, has been analyzed mostly in linear cases. In this paper, we build upon previous results \cite{oono2019graph} to further analyze the over-smoothing effect in the general graph neural network architecture. We show when the weight matrix satisfies the conditions determined by the spectrum of augmented normalized Laplacian, the Dirichlet energy of embeddin
View details →Negative / Null Result ReportOpen accessComputer Science
Matthew Rueben, Frank J. Bernieri, Cindy M. Grimm et al. · 2019 · arXiv
Privacy-sensitive robotics is an emerging area of HRI research. Judgments about privacy would seem to be context-dependent, but none of the promising work on contextual "frames" has focused on privacy concerns. This work studies the impact of contextual "frames" on local users' privacy judgments in a home telepresence setting. Our methodology consists of using an online questionnaire to collect responses to animated videos of a telepresence robot after framing people with an introductory paragraph. The results of four studies indicate a large effect of manipulating the robot operator's identit
View details →Negative / Null Result ReportOpen accessComputer Science
Christian Mollière, Iker Cumplido, Marco Zeulner et al. · 2025 · arXiv
The rapid growth of data from satellite-based Earth observation (EO) systems poses significant challenges in data transmission and storage. We evaluate the potential of task-specific learned compression algorithms in this context to reduce data volumes while retaining crucial information. In detail, we compare traditional compression (JPEG 2000) versus a learned compression approach (Discretized Mixed Gaussian Likelihood) on three EO segmentation tasks: Fire, cloud, and building detection. Learned compression notably outperforms JPEG 2000 for large-scale, multi-channel optical imagery in both
View details →Negative / Null Result ReportOpen accessComputer Science
Runtian Zhai, Chen Dan, Zico Kolter et al. · 2022 · arXiv
Empirical risk minimization (ERM) is known in practice to be non-robust to distributional shift where the training and the test distributions are different. A suite of approaches, such as importance weighting, and variants of distributionally robust optimization (DRO), have been proposed to solve this problem. But a line of recent work has empirically shown that these approaches do not significantly improve over ERM in real applications with distribution shift. The goal of this work is to obtain a comprehensive theoretical understanding of this intriguing phenomenon. We first posit the class o
View details →Negative / Null Result ReportOpen accessComputer Science
Dhananjay Srivastava · 2023 · arXiv
Clinical conversation summarization has become an important application of Natural language Processing. In this work, we intend to analyze summarization model ensembling approaches, that can be utilized to improve the overall accuracy of the generated medical report called chart note. The work starts with a single summarization model creating the baseline. Then leads to an ensemble of summarization models trained on a separate section of the chart note. This leads to the final approach of passing the generated results to another summarization model in a multi-layer/stage fashion for better coh
View details →Negative / Null Result ReportOpen accessComputer Science
Roy Abel, Idan Benami, Yoram Louzoun · 2019 · arXiv
In colored graphs, node classes are often associated with either their neighbors class or with information not incorporated in the graph associated with each node. We here propose that node classes are also associated with topological features of the nodes. We use this association to improve Graph machine learning in general and specifically, Graph Convolutional Networks (GCN). First, we show that even in the absence of any external information on nodes, a good accuracy can be obtained on the prediction of the node class using either topological features, or using the neighbors class as an inp
View details →Negative / Null Result ReportOpen accessComputer Science
Guanxu Chen, Dongrui Liu, Jing Shao · 2026 · arXiv
Large Language Models (LLMs) often exhibit a gap between their internal knowledge and their explicit linguistic outputs. In this report, we empirically investigate whether Looped Transformers (LTs)--architectures that increase computational depth by iterating shared layers--can bridge this gap by utilizing their iterative nature as a form of introspection. Our experiments reveal that while increasing loop iterations narrows the gap, it is partly driven by a degradation of their internal knowledge carried by representations. Moreover, another empirical analysis suggests that current LTs' abilit
View details →Negative / Null Result ReportOpen accessComputer Science
Alex Ayoub, Samuel Robertson, Dawen Liang et al. · 2025 · arXiv
Matrix factorization is a widely used approach for top-N recommendation and collaborative filtering. When implemented on implicit feedback data (such as clicks), a common heuristic is to upweight the observed interactions. This strategy has been shown to improve performance for certain algorithms. In this paper, we conduct a systematic study of various weighting schemes and matrix factorization algorithms. Somewhat surprisingly, we find that training with unweighted data can perform comparably to, and sometimes outperform, training with weighted data, especially for large models. This observat
View details →Negative / Null Result ReportOpen accessComputer Science
Paul K. Mandal · 2025 · arXiv
In this paper, I investigate the effectiveness of dataset cartography for extractive question answering on the SQuAD dataset. I begin by analyzing annotation artifacts in SQuAD and evaluate the impact of two adversarial datasets, AddSent and AddOneSent, on an ELECTRA-small model. Using training dynamics, I partition SQuAD into easy-to-learn, ambiguous, and hard-to-learn subsets. I then compare the performance of models trained on these subsets to those trained on randomly selected samples of equal size. Results show that training on cartography-based subsets does not improve generalization to
View details →Negative / Null Result ReportOpen accessComputer Science
Ab Mosca, Alvitta Ottley, Remco Chang · 2021 · arXiv
Interaction enables users to navigate large amounts of data effectively, supports cognitive processing, and increases data representation methods. However, there have been few attempts to empirically demonstrate whether adding interaction to a static visualization improves its function beyond popular beliefs. In this paper, we address this gap. We use a classic Bayesian reasoning task as a testbed for evaluating whether allowing users to interact with a static visualization can improve their reasoning. Through two crowdsourced studies, we show that adding interaction to a static Bayesian reaso
View details →Negative / Null Result ReportOpen accessComputer Science
Richard G. Clegg · 2006 · arXiv
The aim of this paper is to use a very simple queuing model to compare a number of models from the literature which have been used to replicate the statistical nature of internet traffic and, in particular, the long-range dependence of this traffic. The four models all have the form of discrete time Markov-modulated processes (two other models are introduced for comparison purposes). While it is often stated that long-range dependence has a critical effect on queuing performance, it appears that the models used here do not well replicated the queuing performance of real internet traffic. In pa
View details →