Negative / Null Result ReportOpen accessComputer Science
Kaiyuan Li, Jing-Cheng Pang, Yang Yu · 2026 · arXiv
Reinforcement learning from verifiable rewards (RLVR) stimulates the thinking processes of large language models (LLMs), substantially enhancing their reasoning abilities on verifiable tasks. It is often assumed that similar gains should transfer to general question answering (GQA), but this assumption has not been thoroughly validated. To assess whether RLVR automatically improves LLM performance on GQA, we propose a Cross-Generation evaluation framework that measures the quality of intermediate reasoning by feeding the generated thinking context into LLMs of varying capabilities. Our evaluat
Negative / Null Result ReportOpen accessComputer Science
Taeho Kang, Yiyu Chen, Christian Wallraven · 2026 · arXiv
In this paper, we conduct a detailed investigation on the effect of independent component (IC)-based noise rejection methods in neural network classifier-based decoding of electroencephalography (EEG) data in different task datasets. We apply a pipeline matrix of two popular different independent component (IC) decomposition methods (Infomax and Adaptive Mixture Independent Component Analysis (AMICA)) with three different component rejection strategies (none, ICLabel, and multiple artifact rejection algorithm [MARA]) on three different EEG datasets (motor imagery, long-term memory formation, a
Negative / Null Result ReportOpen accessComputer Science
Vasilis Niarchos, Constantinos Papageorgakis, Alexander G. Stapleton et al. · 2026 · arXiv
As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical question emerges: How does the interaction between researchers and agents affect the results? We study this using SCALAR (Structured Critic--Actor Loop for AI Reasoning), an Actor--Critic--Judge pipeline applied to quantum field theory and string theory problems. The Actor proposes solutions, the Critic provides iterative feedback, and an independent Judge evaluates the transcript against reference solutions. We vary the Actor persona, the Critic fee
Negative / Null Result ReportOpen accessComputer Science
Kazutoshi Shinoda, Saku Sugawara, Akiko Aizawa · 2022 · arXiv
Question answering (QA) models are shown to be insensitive to large perturbations to inputs; that is, they make correct and confident predictions even when given largely perturbed inputs from which humans can not correctly derive answers. In addition, QA models fail to generalize to other domains and adversarial test sets, while humans maintain high accuracy. Based on these observations, we assume that QA models do not use intended features necessary for human reading but rely on spurious features, causing the lack of generalization ability. Therefore, we attempt to answer the question: If the
Negative / Null Result ReportOpen accessComputer Science
Masahito Hayashi · 2008 · arXiv
The optimal exponential error rate for adaptive discrimination of two channels is discussed. In this problem, adaptive choice of input signal is allowed. This problem is discussed in various settings. It is proved that adaptive choice does not improve the exponential error rate in these settings. These results are applied to quantum state discrimination.
Negative / Null Result ReportOpen accessComputer Science
Ruhollah Majdoddin · 2014 · arXiv
AFP Algorithm is a learning algorithm for Horn formulas. We show that it does not improve the complexity of AFP Algorithm, if after each negative counterexample more that just one refinements are performed. Moreover, a canonical normal form for Horn formulas is presented, and it is proved that the output formula of AFP Algorithm is in this normal form.
Negative / Null Result ReportOpen accessComputer Science
Lennart Luettgau, Vanessa Cheung, Magda Dubois et al. · 2025 · arXiv
People increasingly seek personal advice from large language models (LLMs), yet whether humans follow their advice, and its consequences for their well-being, remains unknown. In a longitudinal randomised controlled trial with a representative UK sample (N = 6,474), we found that up to 79% of participants who had a 20-minute discussion with one of three AI chatbots (GPT-4o, LLama-3.3-70B, Gemini 3 Pro) about health, careers or relationships subsequently reported following its advice. Advice-following remained above 60% even for high-stakes recommendations, suggesting that users only weakly cal
Negative / Null Result ReportOpen accessComputer Science
Aaron B. Wagner, Nirmal V. Shende, Yücel Altuğ · 2019 · arXiv
We introduce a novel mechanism, called timid/bold coding, by which feedback can be used to improve coding performance. For a certain class of DMCs, called compound-dispersion channels, we show that timid/bold coding allows for an improved second-order coding rate compared with coding without feedback. For DMCs that are not compound dispersion, we show that feedback does not improve the second-order coding rate. Thus we completely determine the class of DMCs for which feedback improves the second-order coding rate. An upper bound on the second-order coding rate is provided for compound-dispersi
Negative / Null Result ReportOpen accessAgricultural and Biological Sciences
Denis Boyer, Peter D. Walsh · 2010 · arXiv
Thanks to recent technological advances, it is now possible to track with an unprecedented precision and for long periods of time the movement patterns of many living organisms in their habitat. The increasing amount of data available on single trajectories offers the possibility of understanding how animals move and of testing basic movement models. Random walks have long represented the main description for micro-organisms and have also been useful to understand the foraging behaviour of large animals. Nevertheless, most vertebrates, in particular humans and other primates, rely on sophistic
Negative / Null Result ReportOpen accessComputer Science
Saikat Mondal, Chanchal K. Roy · 2025 · arXiv
High-quality answers in technical Q&A platforms like Stack Overflow (SO) are crucial as they directly influence software development practices. Poor-quality answers can introduce inefficiencies, bugs, and security vulnerabilities, and thus increase maintenance costs and technical debt in production software. To improve content quality, SO allows collaborative editing, where users revise answers to enhance clarity, correctness, and formatting. Several studies have examined rejected edits and identified the causes of rejection. However, prior research has not systematically assessed whether acce
Negative / Null Result ReportOpen accessComputer Science
Hongyu Zhou, Yinan Zhang, Aixin Sun et al. · 2025 · arXiv
Multimodal recommendation systems are increasingly popular for their potential to improve performance by integrating diverse data types. However, the actual benefits of this integration remain unclear, raising questions about when and how it truly enhances recommendations. In this paper, we propose a structured evaluation framework to systematically assess multimodal recommendations across four dimensions: Comparative Efficiency, Recommendation Tasks, Recommendation Stages, and Multimodal Data Integration. We benchmark a set of reproducible multimodal models against strong traditional baseline
Negative / Null Result ReportOpen accessPhysics
Martin Amft, Natalia V. Skorodumova · 2011 · arXiv
The present density functional theory study addresses the question whether the presence of H2O influences the catalytic activity of small gold clusters, Au1-4/MgO(100), towards the oxidation of carbon monoxide. To this end, we studied the (co-)adsorption of H2O and CO/O2 on these gold clusters. The ground state structures in the presence of all three molecular species, that we found, are Au1O2/MgO and Au2-4CO/MgO with H2O adsorbed on the surface in the proximity of the clusters-molecule complex. In this configuration the catalytic activity of Au1-4/MgO is indifferent to the presence of H2O. We
Negative / Null Result ReportOpen accessComputer Science
Mark A. Rubin, Sumanth Kaushik · 2006 · arXiv
The signal-to-noise ratio for heterodyne laser radar with a coherent target-return beam and a squeezed local-oscillator beam is lower than that obtained using a coherent local oscillator, regardless of the method employed to combine the beams at the detector.
Negative / Null Result ReportOpen accessComputer Science
Paul Schneider, Amalie Schramm · 2025 · arXiv
Structured deliberation has been found to improve the performance of human forecasters. This study investigates whether a similar intervention, i.e. allowing LLMs to review each other's forecasts before updating, can improve accuracy in large language models (GPT-5, Claude Sonnet 4.5, Gemini Pro 2.5). Using 202 resolved binary questions from the Metaculus Q2 2025 AI Forecasting Tournament, accuracy was assessed across four scenarios: (1) diverse models with distributed information, (2) diverse models with shared information, (3) homogeneous models with distributed information, and (4) homogene
Negative / Null Result ReportOpen accessPhysics
Chao Shen, Yu Liu, Tianquan Tang et al. · 2023 · arXiv
The term `sub-wavelength' is commonly used to describe innovative sound-absorbing structures usually labeled as `metamaterials'. Such structures, however, inherently do not bring groundbreaking advancements. This study addresses the limitations imposed by the thickness criterion of Yang et al. by introducing the concept of equivalent mass-spring-damping parameters within the resonator framework. This innovative approach introduces an index of `half-absorption bandwidth' to effectively overcome the thickness restriction. Four practical cases are then presented to correct prevalent misleading co
Negative / Null Result ReportOpen accessComputer Science
Masahito Hayashi · 2025 · arXiv
We consider estimation of unknown unitary operation when the set of possible unitary operations is given by a projective unitary representation of a compact group. We show that neither indefinite causal order strategy nor adaptive strategy improves the performance of this estimation when error function satisfies group covariance. That is, the optimal parallel strategy gives the optimal performance even under indefinite causal order strategy and adaptive strategy. To study this problem, we newly introduce the concept of generalized positive operator valued measure (GPOVM), and its convariance c
Negative / Null Result ReportOpen accessComputer Science
Xirui Li, Ming Li, Tianyi Zhou · 2026 · arXiv
Reinforcement learning (RL) with verifiable rewards has become a standard post-training stage for boosting visual reasoning in vision-language models, yet it remains unclear what capabilities RL actually improves compared with supervised fine-tuning as cold-start initialization (IN). End-to-end benchmark gains conflate multiple factors, making it difficult to attribute improvements to specific skills. To bridge the gap, we propose a Frankenstein-style analysis framework including: (i) functional localization via causal probing; (ii) update characterization via parameter comparison; and (iii) t
Negative / Null Result ReportOpen accessComputer Science
Roland S. Zimmermann, Thomas Klein, Wieland Brendel · 2023 · arXiv
In light of the recent widespread adoption of AI systems, understanding the internal information processing of neural networks has become increasingly critical. Most recently, machine vision has seen remarkable progress by scaling neural networks to unprecedented levels in dataset and model size. We here ask whether this extraordinary increase in scale also positively impacts the field of mechanistic interpretability. In other words, has our understanding of the inner workings of scaled neural networks improved as well? We use a psychophysical paradigm to quantify one form of mechanistic inter
Negative / Null Result Report
M Erickson, N Tomassoni, R Sun et al. · 2026 · Europace
Abstract Background/Introduction Bachmann bundle pacing (BBP) has emerged as a favorable alternative to traditional right atrial appendage (RAA) pacing, offering benefits such as reduced atrial fibrillation and improved diastolic function.…
View details →DOI: 10.1093/europace/euag105.733 Negative / Null Result ReportOpen accessPhysics
B. Megyeri, A. Lampis, G. Harvie et al. · 2017 · arXiv
We discuss the prospects for enhancing absorption and scattering of light from a weakly coupled atom in a high-finesse optical cavity by adding a medium with large, positive group index of refraction. The slow-light effect is known to narrow the cavity transmission spectrum and increase the photon lifetime, but the quality factor of the cavity may not be increased in a metrologically useful sense. Specifically, detection of the weakly coupled atom through either cavity ringdown measurements or the Purcell effect fails to improve with the addition of material slow light. A single-atom model of
Negative / Null Result Report
Muhammad Dawood Mian, Saadullah Jan Khan, Rehana Rani et al. · 2026 · Access Microbiology
Rapid and reliable identification of bacteria is essential in clinical and environmental microbiology. Gram staining remains a widely used method for preliminary classification; however, it may require additional steps and can be difficult…
View details →DOI: 10.1099/acmi.0.000965.v3 Negative / Null Result Report
Mary Cushman, Suzanne E Judd, Virginia J Howard et al. · 2011 · Circulation
Background. The AHA 2020 Goal includes improving cardiovascular health, defined using a metric consisting of 7 health factors, Life's Simple 7. A central concept of the goal is that small improvements in behavior / lifestyle factors at the…
View details →DOI: 10.1161/circ.124.suppl_21.a17917 Negative / Null Result ReportOpen accessComputer Science
Thomas Vogel, Chinh Tran, Lars Grunske · 2019 · arXiv
In search-based software engineering we often use popular heuristics with default configurations, which typically lead to suboptimal results, or we perform experiments to identify configurations on a trial-and-error basis, which may lead to better results for a specific problem. To obtain better results while avoiding trial-and-error experiments, a fitness landscape analysis is helpful in understanding the search problem, and making an informed decision about the heuristics. In this paper, we investigate the search problem of test suite generation for mobile applications (apps) using SAPIENZ w
Negative / Null Result ReportOpen accessComputer Science
Shufan Wang, Yixiao Song, Andrew Drozdov et al. · 2023 · arXiv
In this paper, we study the generation quality of interpolation-based retrieval-augmented language models (LMs). These methods, best exemplified by the KNN-LM, interpolate the LM's predicted distribution of the next word with a distribution formed from the most relevant retrievals for a given prefix. While the KNN-LM and related methods yield impressive decreases in perplexity, we discover that they do not exhibit corresponding improvements in open-ended generation quality, as measured by both automatic evaluation metrics (e.g., MAUVE) and human evaluations. Digging deeper, we find that interp
Negative / Null Result Report
Ainun Oktavia Sari, Rahayu Sulistyowati, Ita Prihantika · 2020 · Administrativa: Jurnal Birokrasi, Kebijakan dan Pelayanan Publik
The Conditional Cash Transfer (CCT) is a conditional social cash transfer program that provides assistance to Very Poor Households (RTSM) appointed as participants in the Conditional Cash Transfer program which is related to improving the…
View details →DOI: 10.23960/administrativa.v2i3.51 Negative / Null Result ReportOpen accessComputer Science
Niek Beckers, Edwin van Asseldonk, Herman van der Kooij · 2020 · arXiv
Haptic interaction between two humans, for example, a physiotherapist assisting a patient regaining the ability to grasp a cup, likely facilitates motor skill acquisition. Haptic human-human interaction has been shown to enhance individual performance improvement in a tracking task with a visuomotor rotation perturbation. These results are remarkable given that haptically assisting or guiding an individual rarely benefits their individual improvement when the assistance is removed. We, therefore, replicated a study that reported that haptic interaction between humans was beneficial for individ
Negative / Null Result ReportOpen accessComputer Science
Runtian Zhai, Chen Dan, Zico Kolter et al. · 2022 · arXiv
Empirical risk minimization (ERM) is known in practice to be non-robust to distributional shift where the training and the test distributions are different. A suite of approaches, such as importance weighting, and variants of distributionally robust optimization (DRO), have been proposed to solve this problem. But a line of recent work has empirically shown that these approaches do not significantly improve over ERM in real applications with distribution shift. The goal of this work is to obtain a comprehensive theoretical understanding of this intriguing phenomenon. We first posit the class o
Negative / Null Result ReportOpen accessMathematics
Mark Rubin · 2020 · arXiv
Preregistration entails researchers registering their planned research hypotheses, methods, and analyses in a time-stamped document before they undertake their data collection and analyses. This document is then made available with the published research report to allow readers to identify discrepancies between what the researchers originally planned to do and what they actually ended up doing. This historical transparency is supposed to facilitate judgments about the credibility of the research findings. The present article provides a critical review of 17 of the reasons behind this argument.
Negative / Null Result Report
Rebecca L Rivera, Melissa K Maulding, Angela R Abbott et al. · 2016 · The FASEB Journal
Objective To determine the relationship of rural and urban county household status to long‐term food security among households with children in Indiana after a Supplemental Nutrition Assistance Program‐Education (SNAP‐Ed) intervention.…
View details →DOI: 10.1096/fasebj.30.1_supplement.674.26 Negative / Null Result ReportOpen accessComputer Science
Alex Ayoub, Samuel Robertson, Dawen Liang et al. · 2025 · arXiv
Matrix factorization is a widely used approach for top-N recommendation and collaborative filtering. When implemented on implicit feedback data (such as clicks), a common heuristic is to upweight the observed interactions. This strategy has been shown to improve performance for certain algorithms. In this paper, we conduct a systematic study of various weighting schemes and matrix factorization algorithms. Somewhat surprisingly, we find that training with unweighted data can perform comparably to, and sometimes outperform, training with weighted data, especially for large models. This observat