Negative / Null Result ReportOpen accessComputer Science
Marco Cognetta, Tatsuya Hiraoka, Naoaki Okazaki et al. · 2024 · arXiv
We explore threshold vocabulary trimming in Byte-Pair Encoding subword tokenization, a postprocessing step that replaces rare subwords with their component subwords. The technique is available in popular tokenization libraries but has not been subjected to rigorous scientific scrutiny. While the removal of rare subwords is suggested as best practice in machine translation implementations, both as a means to reduce model size and for improving model performance through robustness, our experiments indicate that, across a large space of hyperparameter settings, vocabulary trimming fails to improv
Negative / Null Result ReportOpen accessPhysics
Chao Shen, Yu Liu, Tianquan Tang et al. · 2023 · arXiv
The term `sub-wavelength' is commonly used to describe innovative sound-absorbing structures usually labeled as `metamaterials'. Such structures, however, inherently do not bring groundbreaking advancements. This study addresses the limitations imposed by the thickness criterion of Yang et al. by introducing the concept of equivalent mass-spring-damping parameters within the resonator framework. This innovative approach introduces an index of `half-absorption bandwidth' to effectively overcome the thickness restriction. Four practical cases are then presented to correct prevalent misleading co
Negative / Null Result ReportOpen accessComputer Science
Song Tae-Eun · 2026 · arXiv
Cross-Context Review (CCR) improves LLM verification by separating production and review into independent sessions. A natural extension is multi-turn review: letting the reviewer ask follow-up questions, receive author responses, and review again. We call this Dynamic Cross-Context Review (D-CCR). In a controlled experiment with 30 artifacts and 150 injected errors, we tested four D-CCR variants against the single-pass CCR baseline. Single-pass CCR (F1 = 0.376) significantly outperformed all multi-turn variants, including D-CCR-2b with question-and-answer exchange (F1 = 0.303, $p < 0.001$, $d
Negative / Null Result ReportOpen accessComputer Science
Rylan Malarchick · 2026 · arXiv
We present an analysis of quantum circuit fidelity across the full compilation stack, from high-level gate optimization through pulse-level control. We connect a C++ circuit optimizer to a per-gate Lindblad master-equation fidelity model whose decoherence channels are cross-validated against qiskit-dynamics and whose absolute predictions are benchmarked against execution on real hardware. Across a campaign of 4,452 experiment runs over 371 benchmark circuits, gate cancellation provides the dominant improvement ($d = 1.66$, 72% of circuits improved), while circuit size and pulse duration are th
Negative / Null Result ReportMedicine
Okada, Tominaga, Iwakura et al. · 2026 · Resuscitation
Out-of-hospital cardiac arrest (OHCA) is a public health concern. Whether conventional cardiopulmonary resuscitation (CPR) with rescue breathing may improve outcomes compared with compression-only or no CPR in suffocation-related OHCA is…
View details →DOI: 10.1016/j.resuscitation.2026.111189 Negative / Null Result ReportOpen accessComputer Science
Lennart Meincke, Ethan Mollick, Lilach Mollick et al. · 2025 · arXiv
This is the third in a series of short reports that seek to help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. In this report, we investigate two commonly held prompting beliefs: a) offering to tip the AI model and b) threatening the AI model. Tipping was a commonly shared tactic for improving AI performance and threats have been endorsed by Google Founder Sergey Brin (All-In, May 2025, 8:20) who observed that 'models tend to do better if you threaten them,' a claim we subject to empirical testing here. We evaluate model p
Negative / Null Result ReportOpen accessComputer Science
Dhruvi Khandelwal, Anurag Basistha, Ayushi Jolotia et al. · 2026 · arXiv
Deep learning proxies for Alternating Current Optimal Power Flow (ACOPF) lack systematic methods for determining architectural size. This paper conducts a constructive thought experiment to answer a fundamental inquiry: how wide must a neural network be to almost accurately approximate the ACOPF manifold? We introduce a Loss-Guided Neural Densification (LG-ND) algorithm that incrementally discovers necessary capacity by expanding only when the current deep neural network topology fails to improve further. Empirical results across various IEEE systems show that LG-ND achieves performance parity
Negative / Null Result ReportOpen accessComputer Science
Runzhe Zhan, Xuebo Liu, Derek F. Wong et al. · 2021 · arXiv
Meta-learning has been sufficiently validated to be beneficial for low-resource neural machine translation (NMT). However, we find that meta-trained NMT fails to improve the translation performance of the domain unseen at the meta-training stage. In this paper, we aim to alleviate this issue by proposing a novel meta-curriculum learning for domain adaptation in NMT. During meta-training, the NMT first learns the similar curricula from each domain to avoid falling into a bad local optimum early, and finally learns the curricula of individualities to improve the model robustness for learning dom
Negative / Null Result ReportOpen accessComputer Science
Masahito Hayashi · 2025 · arXiv
We consider estimation of unknown unitary operation when the set of possible unitary operations is given by a projective unitary representation of a compact group. We show that neither indefinite causal order strategy nor adaptive strategy improves the performance of this estimation when error function satisfies group covariance. That is, the optimal parallel strategy gives the optimal performance even under indefinite causal order strategy and adaptive strategy. To study this problem, we newly introduce the concept of generalized positive operator valued measure (GPOVM), and its convariance c
Negative / Null Result ReportOpen accessComputer Science
Xirui Li, Ming Li, Tianyi Zhou · 2026 · arXiv
Reinforcement learning (RL) with verifiable rewards has become a standard post-training stage for boosting visual reasoning in vision-language models, yet it remains unclear what capabilities RL actually improves compared with supervised fine-tuning as cold-start initialization (IN). End-to-end benchmark gains conflate multiple factors, making it difficult to attribute improvements to specific skills. To bridge the gap, we propose a Frankenstein-style analysis framework including: (i) functional localization via causal probing; (ii) update characterization via parameter comparison; and (iii) t
Negative / Null Result ReportOpen accessEconomics, Econometrics and Finance
Mohammad Hassan Shakil, Arne Johan Pollestad, Khine Kyaw et al. · 2025 · arXiv
With European Union initiatives mandating gender quotas on corporate boards, a key question arises: Is greater board gender diversity (BGD) associated with better emissions performance (EP)? To answer this question, we examine the influence of BGD on EP across a sample of European firms from 2016 to 2022. Using panel regressions, advanced machine learning algorithms, and explainable AI, we reveal a non-linear relationship. Specifically, EP improves with BGD up to an optimal level of approximately 35 %, beyond which further increases in BGD yield no additional improvement in EP. A minimum BGD t
Negative / Null Result ReportOpen accessComputer Science
Hanbing Liu, Haoyang Li, Xiaokang Zhang et al. · 2025 · arXiv
Direct Preference Optimization (DPO) has proven effective in complex reasoning tasks like math word problems and code generation. However, when applied to Text-to-SQL datasets, it often fails to improve performance and can even degrade it. Our investigation reveals the root cause: unlike math and code tasks, which naturally integrate Chain-of-Thought (CoT) reasoning with DPO, Text-to-SQL datasets typically include only final answers (gold SQL queries) without detailed CoT solutions. By augmenting Text-to-SQL datasets with synthetic CoT solutions, we achieve, for the first time, consistent and
Negative / Null Result Report
Rebecca L Rivera, Melissa K Maulding, Angela R Abbott et al. · 2016 · The FASEB Journal
Objective To determine the relationship of rural and urban county household status to long‐term food security among households with children in Indiana after a Supplemental Nutrition Assistance Program‐Education (SNAP‐Ed) intervention.…
View details →DOI: 10.1096/fasebj.30.1_supplement.674.26 Negative / Null Result ReportOpen accessComputer Science
Roland S. Zimmermann, Thomas Klein, Wieland Brendel · 2023 · arXiv
In light of the recent widespread adoption of AI systems, understanding the internal information processing of neural networks has become increasingly critical. Most recently, machine vision has seen remarkable progress by scaling neural networks to unprecedented levels in dataset and model size. We here ask whether this extraordinary increase in scale also positively impacts the field of mechanistic interpretability. In other words, has our understanding of the inner workings of scaled neural networks improved as well? We use a psychophysical paradigm to quantify one form of mechanistic inter
Negative / Null Result Report
M Erickson, N Tomassoni, R Sun et al. · 2026 · Europace
Abstract Background/Introduction Bachmann bundle pacing (BBP) has emerged as a favorable alternative to traditional right atrial appendage (RAA) pacing, offering benefits such as reduced atrial fibrillation and improved diastolic function.…
View details →DOI: 10.1093/europace/euag105.733 Negative / Null Result ReportOpen accessComputer Science
Fırat Öncel, Matthias Bethge, Beyza Ermis et al. · 2024 · arXiv
In the last decade, the generalization and adaptation abilities of deep learning models were typically evaluated on fixed training and test distributions. Contrary to traditional deep learning, large language models (LLMs) are (i) even more overparameterized, (ii) trained on unlabeled text corpora curated from the Internet with minimal human intervention, and (iii) trained in an online fashion. These stark contrasts prevent researchers from transferring lessons learned on model generalization and adaptation in deep learning contexts to LLMs. To this end, our short paper introduces empirical ob
Negative / Null Result ReportOpen accessPhysics
B. Megyeri, A. Lampis, G. Harvie et al. · 2017 · arXiv
We discuss the prospects for enhancing absorption and scattering of light from a weakly coupled atom in a high-finesse optical cavity by adding a medium with large, positive group index of refraction. The slow-light effect is known to narrow the cavity transmission spectrum and increase the photon lifetime, but the quality factor of the cavity may not be increased in a metrologically useful sense. Specifically, detection of the weakly coupled atom through either cavity ringdown measurements or the Purcell effect fails to improve with the addition of material slow light. A single-atom model of
Negative / Null Result ReportOpen accessComputer Science
Chunliang Li, Tianze Cao, Sanyuan Zhao · 2026 · arXiv
Visual Autoregressive (VAR) modeling inefficiently applies a fixed computational depth to each position when generating high-resolution images. While existing methods accelerate inference by pruning tokens using frequency maps, their binary hard-pruning approach is fundamentally limited and fails to improve quality even with better frequency estimation. Observing that VAR models possess significant depth redundancy, we propose a paradigm shift from pruning entire tokens to adaptively allocating per-token computational depth. To this end, we introduce DepthVAR, a training-free framework that dy
Negative / Null Result Report
Muhammad Dawood Mian, Saadullah Jan Khan, Rehana Rani et al. · 2026 · Access Microbiology
Rapid and reliable identification of bacteria is essential in clinical and environmental microbiology. Gram staining remains a widely used method for preliminary classification; however, it may require additional steps and can be difficult…
View details →DOI: 10.1099/acmi.0.000965.v3 Negative / Null Result ReportOpen accessComputer Science
Ruth Cohen, Lu Feng, Ayala Bloch et al. · 2026 · arXiv
While natural-language explanations from large language models (LLMs) are widely adopted to improve transparency and trust, their impact on objective human-AI team performance remains poorly understood. We identify a Persuasion Paradox: fluent explanations systematically increase user confidence and reliance on AI without reliably improving, and in some cases undermining, task accuracy. Across three controlled human-subject studies spanning abstract visual reasoning (RAVEN matrices) and deductive logical reasoning (LSAT problems), we disentangle the effects of AI predictions and explanations u
Negative / Null Result Report
Mary Cushman, Suzanne E Judd, Virginia J Howard et al. · 2011 · Circulation
Background. The AHA 2020 Goal includes improving cardiovascular health, defined using a metric consisting of 7 health factors, Life's Simple 7. A central concept of the goal is that small improvements in behavior / lifestyle factors at the…
View details →DOI: 10.1161/circ.124.suppl_21.a17917 Negative / Null Result ReportOpen accessComputer Science
OFM Riaz Rahman Aranya, Kevin Desai · 2026 · arXiv
Synthetic data, an appealing alternative to extensive expert-annotated data for medical image segmentation, consistently fails to improve segmentation performance despite its visual realism. The reason being that synthetic and real medical images exist in different semantic feature spaces, creating a domain gap that current semi-supervised learning methods cannot bridge. We propose SRA-Seg, a framework explicitly designed to align synthetic and real feature distributions for medical image segmentation. SRA-Seg introduces a similarity-alignment (SA) loss using frozen DINOv2 embeddings to pull s
Negative / Null Result ReportOpen accessEngineering
Sreyan Ghosh, Mohammad Sadegh Rasooli, Michael Levit et al. · 2024 · arXiv
Generative Error Correction (GEC) has emerged as a powerful post-processing method to enhance the performance of Automatic Speech Recognition (ASR) systems. However, we show that GEC models struggle to generalize beyond the specific types of errors encountered during training, limiting their ability to correct new, unseen errors at test time, particularly in out-of-domain (OOD) scenarios. This phenomenon amplifies with named entities (NEs), where, in addition to insufficient contextual information or knowledge about the NEs, novel NEs keep emerging. To address these issues, we propose DARAG (D
Negative / Null Result ReportMedicine
Yang, Mueller, D'Andrea et al. · 2026 · Academic psychiatry : the journal of the American Association of Directors of Psychiatric Residency Training and the Association for Academic Psychiatry
Resident physicians experience high rates of depression, anxiety, burnout, and loneliness, yet few evidence-based interventions have been evaluated in this population. This randomized controlled pilot trial examined the feasibility,…
View details →DOI: 10.1007/s40596-026-02384-y Negative / Null Result ReportOpen accessComputer Science
Thomas Vogel, Chinh Tran, Lars Grunske · 2019 · arXiv
In search-based software engineering we often use popular heuristics with default configurations, which typically lead to suboptimal results, or we perform experiments to identify configurations on a trial-and-error basis, which may lead to better results for a specific problem. To obtain better results while avoiding trial-and-error experiments, a fitness landscape analysis is helpful in understanding the search problem, and making an informed decision about the heuristics. In this paper, we investigate the search problem of test suite generation for mobile applications (apps) using SAPIENZ w
Negative / Null Result ReportOpen accessComputer Science
Shufan Wang, Yixiao Song, Andrew Drozdov et al. · 2023 · arXiv
In this paper, we study the generation quality of interpolation-based retrieval-augmented language models (LMs). These methods, best exemplified by the KNN-LM, interpolate the LM's predicted distribution of the next word with a distribution formed from the most relevant retrievals for a given prefix. While the KNN-LM and related methods yield impressive decreases in perplexity, we discover that they do not exhibit corresponding improvements in open-ended generation quality, as measured by both automatic evaluation metrics (e.g., MAUVE) and human evaluations. Digging deeper, we find that interp
Negative / Null Result Report
Ainun Oktavia Sari, Rahayu Sulistyowati, Ita Prihantika · 2020 · Administrativa: Jurnal Birokrasi, Kebijakan dan Pelayanan Publik
The Conditional Cash Transfer (CCT) is a conditional social cash transfer program that provides assistance to Very Poor Households (RTSM) appointed as participants in the Conditional Cash Transfer program which is related to improving the…
View details →DOI: 10.23960/administrativa.v2i3.51 Negative / Null Result ReportOpen accessComputer Science
Zhiwei Jia, Xuanlin Li, Zhan Ling et al. · 2022 · arXiv
Generalization in deep reinforcement learning over unseen environment variations usually requires policy learning over a large set of diverse training variations. We empirically observe that an agent trained on many variations (a generalist) tends to learn faster at the beginning, yet its performance plateaus at a less optimal level for a long time. In contrast, an agent trained only on a few variations (a specialist) can often achieve high returns under a limited computational budget. To have the best of both worlds, we propose a novel generalist-specialist training framework. Specifically, w
Negative / Null Result ReportOpen accessComputer Science
Niek Beckers, Edwin van Asseldonk, Herman van der Kooij · 2020 · arXiv
Haptic interaction between two humans, for example, a physiotherapist assisting a patient regaining the ability to grasp a cup, likely facilitates motor skill acquisition. Haptic human-human interaction has been shown to enhance individual performance improvement in a tracking task with a visuomotor rotation perturbation. These results are remarkable given that haptically assisting or guiding an individual rarely benefits their individual improvement when the assistance is removed. We, therefore, replicated a study that reported that haptic interaction between humans was beneficial for individ
Negative / Null Result ReportOpen accessComputer Science
Runtian Zhai, Chen Dan, Zico Kolter et al. · 2022 · arXiv
Empirical risk minimization (ERM) is known in practice to be non-robust to distributional shift where the training and the test distributions are different. A suite of approaches, such as importance weighting, and variants of distributionally robust optimization (DRO), have been proposed to solve this problem. But a line of recent work has empirically shown that these approaches do not significantly improve over ERM in real applications with distribution shift. The goal of this work is to obtain a comprehensive theoretical understanding of this intriguing phenomenon. We first posit the class o