Negative / Null Result ReportOpen accessComputer Science
Philippe Laban, Tobias Schnabel, Jennifer Neville · 2026 · arXiv
Large Language Models (LLMs) are poised to disrupt knowledge work, with the emergence of delegated work as a new interaction paradigm (e.g., vibe coding). Delegation requires trust - the expectation that the LLM will faithfully execute the task without introducing errors into documents. We introduce DELEGATE-52 to study the readiness of AI systems in delegated workflows. DELEGATE-52 simulates long delegated workflows that require in-depth document editing across 52 professional domains, such as coding, crystallography, and music notation. Our large-scale experiment with 19 LLMs reveals that cu
View details →Negative / Null Result ReportOpen accessComputer Science
Prabhav Jain, Nikolai Miklin, Mariami Gachechiladze · 2026 · arXiv
Information Causality was introduced as a physical principle for constraining the set of nonlocal correlations. In recent work, we proposed an extension of Information Causality that allows correlations among Alice's inputs. This extended principle yields tighter constraints than the original formulation and recovers part of the quantum boundary in certain Bell scenarios. In this work, we further investigate the implications of extended Information Causality and apply it to scenarios beyond binary inputs and outputs. We derive a family of quantum Bell inequalities that strengthen previously kn
View details →Negative / Null Result ReportOpen accessComputer Science
Jirui Qi, Raquel Fernández, Arianna Bisazza · 2023 · arXiv
Multilingual large-scale Pretrained Language Models (PLMs) have been shown to store considerable amounts of factual knowledge, but large variations are observed across languages. With the ultimate goal of ensuring that users with different language backgrounds obtain consistent feedback from the same model, we study the cross-lingual consistency (CLC) of factual knowledge in various multilingual PLMs. To this end, we propose a Ranking-based Consistency (RankC) metric to evaluate knowledge consistency across languages independently from accuracy. Using this metric, we conduct an in-depth analys
View details →Negative / Null Result ReportOpen accessComputer Science
Gil Katz, Pablo Piantanida, Merouane Debbah · 2016 · arXiv
A collaborative distributed binary decision problem is considered. Two statisticians are required to declare the correct probability measure of two jointly distributed memoryless process, denoted by $X^n=(X_1,\dots,X_n)$ and $Y^n=(Y_1,\dots,Y_n)$, out of two possible probability measures on finite alphabets, namely $P_{XY}$ and $P_{\bar{X}\bar{Y}}$. The marginal samples given by $X^n$ and $Y^n$ are assumed to be available at different locations. The statisticians are allowed to exchange limited amount of data over multiple rounds of interactions, which differs from previous work that deals mai
View details →Negative / Null Result ReportOpen accessComputer Science
Rumman Ahmed Prodhan, Sumya Akter, Tanmoy Sarkar Pias et al. · 2023 · arXiv
The human brain is a complex organ, still completely undiscovered, that controls almost all the parts of the body. Apart from survival, the human brain stimulates emotions. Recent research indicates that brain signals can be very effective for emotion recognition. However, which parts of the brain exhibit most of the emotions is still under-explored. In this study, we empirically analyze the contribution of each part of the brain in exhibiting emotions. We use the DEAP dataset to find the most optimal electrode set which eventually leads to the effective brain part associated with emotions. We
View details →Negative / Null Result ReportOpen accessComputer Science
Roberto Casadio, Fabio Finelli, Mattia Luzzi et al. · 2004 · arXiv
Improved Wentzel-Kramers-Brillouin (WKB)-type approximations are presented in order to study cosmological perturbations beyond the lowest order. Our methods are based on functions which approximate the true perturbation modes over the complete range of the independent (Langer) variable, from sub-horizon to super-horizon scales, and include the region near the turning point. We employ both a perturbative Green's function technique and an adiabatic (or ``semiclassical'') expansion (for a linear turning point) in order to compute higher order corrections. Improved general expressions for the WKB
View details →Negative / Null Result ReportOpen accessComputer Science
Abdul Waheed, Zhen Wu, Dareen Alharthi et al. · 2025 · arXiv
Precisely evaluating video understanding models remains challenging: commonly used metrics such as BLEU, ROUGE, and BERTScore fail to capture the fineness of human judgment, while obtaining such judgments through manual evaluation is costly. Recent work has explored using large language models (LLMs) or multimodal LLMs (MLLMs) as evaluators, but their extension to video understanding remains relatively unexplored. In this work, we introduce VideoJudge, a 3B and 7B-sized MLLM judge specialized to evaluate outputs from video understanding models (\textit{i.e.}, text responses conditioned on vide
View details →Negative / Null Result ReportOpen accessComputer Science
Gokhan Alkac, Luca Basanisi, Ercan Kilicarslan et al. · 2017 · arXiv
We revisit the problem of the bulk-boundary unitarity clash in 2 + 1 dimensional gravity theories, which has been an obstacle in providing a viable dual two-dimensional conformal field theory for bulk gravity in anti-de Sitter (AdS) spacetime. Chiral gravity, which is a particular limit of cosmological topologically massive gravity (TMG), suffers from pertur- bative log-modes with negative energies inducing a non-unitary logarithmic boundary field theory. We show here that any f(R) extension of TMG does not improve the situation. We also study the perturbative modes in the metric formulation o
View details →Negative / Null Result ReportOpen accessComputer Science
Tom Kouwenhoven, Max Peeperkorn, Bram van Dijk et al. · 2024 · arXiv
Natural language has the universal properties of being compositional and grounded in reality. The emergence of linguistic properties is often investigated through simulations of emergent communication in referential games. However, these experiments have yielded mixed results compared to similar experiments addressing linguistic properties of human language. Here we address representational alignment as a potential contributing factor to these results. Specifically, we assess the representational alignment between agent image representations and between agent representations and input images.
View details →Negative / Null Result ReportOpen accessComputer Science
Luigi Giannelli, Ralf Betzholz, Laura Kreiner et al. · 2016 · arXiv
We theoretically analyse the cooling dynamics of a high-Q mode of a mechanical resonator, when the structure is also an optical cavity and is coupled with a NV center. The NV center is driven by a laser and interacts with the cavity photon field and with the strain field of the mechanical oscillator, while radiation pressure couples mechanical resonator and cavity field. Starting from the full master equation we derive the rate equation for the mechanical resonator's motion, whose coefficients depend on the system parameters and on the noise sources. We then determine the cooling regime, the c
View details →Negative / Null Result ReportOpen accessComputer Science
Elias Hossain, Md Mehedi Hasan Nipu, Maleeha Sheikh et al. · 2025 · arXiv
Clinical language models often assign high confidence to incorrect predictions, particularly in high-severity and out-of-distribution cases. We present MedBayes-Lite, a retraining-free uncertainty governance layer for transformer-based clinical predictors. It combines Monte Carlo dropout, predictive calibration, and confidence-guided abstention to defer low-confidence predictions for human review, adding no trainable parameters. Evaluated on MedMCQA and MedQA-USMLE, MedBayes-Lite reduces expected calibration error by 0.23 to 0.33 and drives harmful overconfident errors (confident, incorrect, h
View details →Negative / Null Result ReportOpen accessComputer Science
Lei Wang · 2026 · arXiv
The endpoint region $ζ\to1$ of the NLO forward jet vertex has not been systematically separated from BFKL energy-scale terms in Mueller-Navelet phenomenology. Starting from the small-cone NLO vertex, we isolate the quark and gluon plus distributions and construct a BFKL-aware threshold matching scheme that preserves exact NLO accuracy. The conservative Scheme-II exponent resums only the ordinary endpoint logarithms and leaves the $χ(n,γ)\ln\bar N$ term in the fixed-order coefficient, avoiding an uncontrolled tower of mixed endpoint-BFKL logarithms. In fixed-baseline CMS tests, this matched ver
View details →Negative / Null Result ReportOpen accessComputer Science
Silas L. Fong, Vincent Y. F. Tan · 2017 · arXiv
This paper investigates the asymptotic expansion for the maximum rate of fixed-length codes over a parallel Gaussian channel with feedback under the following setting: A peak power constraint is imposed on every transmitted codeword, and the average error probability of decoding the transmitted message is non-vanishing as the blocklength increases. It is well known that the presence of feedback does not increase the first-order asymptotics of the channel, i.e., capacity, in the asymptotic expansion, and the closed-form expression of the capacity can be obtained by the well-known water-filling
View details →Negative / Null Result ReportOpen accessComputer Science
Ryosuke Kohita, Seiichiro Yoshioka · 2026 · arXiv
Memes are a popular element of modern web communication, used not only as static artifacts but also as interactive replies within conversations. While computational research has focused on analyzing the intrinsic properties of memes, the dynamic and contextual use of memes to create humor remains an understudied area of web science. To address this gap, we introduce the Meme Reply Selection task and present MaMe-Re (Manga Meme Reply Benchmark), a benchmark of 100,000 human-annotated pairs (500,000 total annotations from 2,325 unique annotators) consisting of openly licensed Japanese manga pane
View details →Negative / Null Result ReportOpen accessComputer Science
Evgeniia Razumovskaia, Ivan Vulić, Anna Korhonen · 2024 · arXiv
Supervised fine-tuning (SFT), supervised instruction tuning (SIT) and in-context learning (ICL) are three alternative, de facto standard approaches to few-shot learning. ICL has gained popularity recently with the advent of LLMs due to its simplicity and sample efficiency. Prior research has conducted only limited investigation into how these approaches work for multilingual few-shot learning, and the focus so far has been mostly on their performance. In this work, we present an extensive and systematic comparison of the three approaches, testing them on 6 high- and low-resource languages, thr
View details →Negative / Null Result ReportOpen accessComputer Science
Zijie Zhao, Roy E. Welsch · 2026 · arXiv
Financial retrieval-augmented generation (RAG) systems typically rank evidence by textual relevance, but in financial markets evidence utility depends on event type, forecast horizon, and market context. We study news-triggered event-impact prediction as a point-in-time financial RAG problem. For each company-news anchor, the system retrieves financial news and SEC filing passages, appends a pre-decision market-context card, and predicts multi-horizon residual-return signals. Our method keeps the LLM frozen and adapts retrieval through an external Bayesian source memory updated from matured re
View details →Negative / Null Result ReportOpen accessComputer Science
Baha Eddine Youcef Belmekki, Abdelkrim Hamza, Benoît Escrig · 2018 · arXiv
Vehicular safety communications (VSCs) are known to provide relevant contributions to avoid congestions and prevent road accidents, and more particularly at road intersections since these areas are more prone to accidents. In this context, one of the main impairments that affect the performance of VSCs are interference. In this paper, we develop a tractable framework to model cooperative transmissions in presence of interference for VSCs at intersections. We use tools from stochastic geometry, and model interferer vehicles locations as a Poisson point process. First, we calculate the outage pr
View details →Negative / Null Result ReportOpen accessComputer Science
Jennifer Haase, Jana Gonnermann-Müller, See Heng Yim et al. · 2026 · arXiv
Large language model (LLM)-based simulations of clinical patients are increasingly used for research and training, yet their validity requires persona stability: coherent maintenance of an assigned psychological profile across and within conversations. We evaluate this prerequisite using eating disorder personas grounded in five published case vignettes, a dual-assessment framework (self-report + independent observer ratings), and validated psychometric instruments (EDE-Q) with known ground-truth scores. Across six LLMs and two experiments (between-conversation stability (Exp. I) and within-co
View details →Negative / Null Result ReportOpen accessComputer Science
Vitaly Kurin, Maximilian Igl, Tim Rocktäschel et al. · 2020 · arXiv
Multitask Reinforcement Learning is a promising way to obtain models with better performance, generalisation, data efficiency, and robustness. Most existing work is limited to compatible settings, where the state and action space dimensions are the same across tasks. Graph Neural Networks (GNN) are one way to address incompatible environments, because they can process graphs of arbitrary size. They also allow practitioners to inject biases encoded in the structure of the input graph. Existing work in graph-based continuous control uses the physical morphology of the agent to construct the inpu
View details →Negative / Null Result ReportOpen accessComputer Science
Jin Zeng, Yang Liu, Gene Cheung et al. · 2022 · arXiv
A graph convolutional network (GCN) employs a graph filtering kernel tailored for data with irregular structures. However, simply stacking more GCN layers does not improve performance; instead, the output converges to an uninformative low-dimensional subspace, where the convergence rate is characterized by the graph spectrum -- this is the known over-smoothing problem in GCN. In this paper, we propose a sparse graph learning algorithm incorporating a new spectrum prior to compute a graph topology that circumvents over-smoothing while preserving pairwise correlations inherent in data. Specifica
View details →Negative / Null Result ReportOpen accessComputer Science
Wentao Zhao, Dalin Zhou, Xinguo Qiu et al. · 2020 · arXiv
Graph neural networks (GNNs) have been investigated for potential applicability in multiple fields that employ graph data. However, there are no standard training settings to ensure fair comparisons among new methods, including different model architectures and data augmentation techniques. We introduce a standard, reproducible benchmark to which the same training settings can be applied for node classification. For this benchmark, we constructed 9 datasets, including both small- and medium-scale datasets from different fields, and 7 different models. We design a k-fold model assessment strate
View details →Negative / Null Result ReportOpen accessComputer Science
Charles Jin, Melinda Sun, Martin Rinard · 2021 · arXiv
We propose a novel clustering mechanism based on an incompatibility property between subsets of data that emerges during model training. This mechanism partitions the dataset into subsets that generalize only to themselves, i.e., training on one subset does not improve performance on the other subsets. Leveraging the interaction between the dataset and the training process, our clustering mechanism partitions datasets into clusters that are defined by--and therefore meaningful to--the objective of the training process. We apply our clustering mechanism to defend against data poisoning attacks,
View details →Negative / Null Result ReportOpen accessComputer Science
Alexia Jolicoeur-Martineau · 2019 · arXiv
This paper provides a more rigorous look at Relativistic Generative Adversarial Networks (RGANs). We prove that the objective function of the discriminator is a statistical divergence for any concave function $f$ with minimal properties ($f(0)=0$, $f'(0) \neq 0$, $\sup_x f(x)>0$). We also devise a few variants of relativistic $f$-divergences. Wasserstein GAN was originally justified by the idea that the Wasserstein distance (WD) is most sensible because it is weak (i.e., it induces a weak topology). We show that the WD is weaker than $f$-divergences which are weaker than relativistic $f$-diver
View details →Negative / Null Result ReportOpen accessComputer Science
Zachary Baker, Yuxiao Li · 2025 · arXiv
Sparse Autoencoders (SAEs) have emerged as a promising approach for interpreting neural network representations by learning sparse, human-interpretable features from dense activations. We investigate whether incorporating variational methods into SAE architectures can improve feature organization and interpretability. We introduce the Variational Sparse Autoencoder (vSAE), which replaces deterministic ReLU gating with stochastic sampling from learned Gaussian posteriors and incorporates KL divergence regularization toward a standard normal prior. Our hypothesis is that this probabilistic sampl
View details →Negative / Null Result ReportOpen accessComputer Science
Ali Subhan, Ashir Raza · 2026 · arXiv
DragDiffusion is a diffusion-based method for interactive point-based image editing that enables users to manipulate images by directly dragging selected points. The method claims that accurate spatial control can be achieved by optimizing a single diffusion latent at an intermediate timestep, together with identity-preserving fine-tuning and spatial regularization. This work presents a reproducibility study of DragDiffusion using the authors' released implementation and the DragBench benchmark. We reproduce the main ablation studies on diffusion timestep selection, LoRA-based fine-tuning, mas
View details →Replication FailureOpen accessComputer Science
Zhichao Fan, Yanhang Li, Zexin Zhuang · 2026 · arXiv
Forced chain-of-thought (CoT) is widely assumed to make vision-language models more reliable on video question answering. We propose a small three-probe evaluation recipe to test that assumption: paired accuracy across direct, CoT, answer-first, and no-video conditions; a counterfactual video-swap diagnostic over the CoT chains; and a four-rung visual-degradation ladder. Each probe is reported under both a strict and a permissive regex scorer, with multiplicity correction over a manuscript-declared primary family. Applied to Qwen2.5-VL on Video-MME subsets, the recipe returns a two-part findin
View details →Replication FailureOpen accessComputer Science
Xiqi Hao, Zengqing Wu, Yu-Xuan Qiu et al. · 2026 · arXiv
Multi-agent debate (MAD) is a promising strategy for improving LLM reasoning, but when agents converge on a shared answer, it is unclear whether that convergence reflects genuine deliberation or social compliance. We show that the conventional answer flip rate conflates three distinct mechanisms: spontaneous instability, stance-induced conformity, and reasoning-induced persuasion. Our three-source decomposition framework isolates each through controlled counterfactual conditions. In the primary MMLU-Pro setting, 37% of agent-question observations change under self-reflection alone, while robus
View details →Replication FailureOpen accessComputer Science
J. Jon Ryu, Pavan Yeddanapudi, Xiangxiang Xu et al. · 2025 · arXiv
The InfoNCE objective, originally introduced for contrastive representation learning, has become a popular choice for mutual information (MI) estimation, despite its indirect connection to MI. In this paper, we demonstrate why InfoNCE should not be regarded as a valid MI estimator, and we introduce a simple modification, which we refer to as InfoNCE-anchor, for accurate MI estimation. Our modification introduces an auxiliary anchor class, enabling consistent density ratio estimation and yielding a plug-in MI estimator with significantly reduced bias. Beyond this, we generalize our framework us
View details →Negative / Null Result ReportOpen accessComputer Science
Paul Röttger, Janet B. Pierrehumbert · 2021 · arXiv
Language use differs between domains and even within a domain, language use changes over time. For pre-trained language models like BERT, domain adaptation through continued pre-training has been shown to improve performance on in-domain downstream tasks. In this article, we investigate whether temporal adaptation can bring additional benefits. For this purpose, we introduce a corpus of social media comments sampled over three years. It contains unlabelled data for adaptation and evaluation on an upstream masked language modelling task as well as labelled data for fine-tuning and evaluation on
View details →Negative / Null Result ReportOpen accessComputer Science
Mireille El-Assal, Pierre Tirilly, Ioan Marius Bilasco · 2023 · arXiv
Video analysis is a computer vision task that is useful for many applications like surveillance, human-machine interaction, and autonomous vehicles. Deep Convolutional Neural Networks (CNNs) are currently the state-of-the-art methods for video analysis. However they have high computational costs, and need a large amount of labeled data for training. In this paper, we use Convolutional Spiking Neural Networks (CSNNs) trained with the unsupervised Spike Timing-Dependent Plasticity (STDP) learning rule for action classification. These networks represent the information using asynchronous low-ener
View details →