Negative / Null Result ReportOpen accessComputer Science
Àlex Miranda-Pascual, Javier Parra-Arnau, Thorsten Strufe · 2026 · arXiv
Sampling is renowned for its privacy amplification in differential privacy (DP), and is often assumed to improve the utility of a DP mechanism by allowing a noise reduction. In this paper, we further show that this last assumption is flawed: When measuring utility at equal privacy levels, sampling as preprocessing consistently yields penalties due to utility loss from omitting records over all canonical DP mechanisms -- Laplace, Gaussian, exponential, and report noisy max -- , as well as recent applications of sampling, such as clustering. Extending this analysis, we investigate suppression as
Negative / Null Result ReportOpen accessEngineering
Haoyang Li, Yuchen Hu, Chen Chen et al. · 2024 · arXiv
Deep neural network (DNN)-based speech enhancement (SE) usually uses conventional activation functions, which lack the expressiveness to capture complex multiscale structures needed for high-fidelity SE. Group-Rational KAN (GR-KAN), a variant of Kolmogorov-Arnold Networks (KAN), retains KAN's expressiveness while improving scalability on complex tasks. We adapt GR-KAN to existing DNN-based SE by replacing dense layers with GR-KAN layers in the time-frequency (T-F) domain MP-SENet and adapting GR-KAN's activations into the 1D CNN layers in the time-domain Demucs. Results on Voicebank-DEMAND sho
Negative / Null Result ReportOpen accessComputer Science
Piyawat Lertvittayakumjorn, David Kinney, Vinodkumar Prabhakaran et al. · 2025 · arXiv
Generative large language models (LLMs) have demonstrated gaps in diverse cultural awareness across the globe. We investigate the effect of retrieval augmented generation and search-grounding techniques on LLMs' ability to display familiarity with various national cultures. Specifically, we compare the performance of standard LLMs, LLMs augmented with retrievals from a bespoke knowledge base (i.e., KB grounding), and LLMs augmented with retrievals from a web search (i.e., search grounding) on multiple cultural awareness benchmarks. We find that search grounding significantly improves the LLM p
Negative / Null Result ReportOpen accessComputer Science
Fırat Öncel, Matthias Bethge, Beyza Ermis et al. · 2024 · arXiv
In the last decade, the generalization and adaptation abilities of deep learning models were typically evaluated on fixed training and test distributions. Contrary to traditional deep learning, large language models (LLMs) are (i) even more overparameterized, (ii) trained on unlabeled text corpora curated from the Internet with minimal human intervention, and (iii) trained in an online fashion. These stark contrasts prevent researchers from transferring lessons learned on model generalization and adaptation in deep learning contexts to LLMs. To this end, our short paper introduces empirical ob
Negative / Null Result ReportOpen accessComputer Science
Adam Byerly, Daniel Khashabi · 2024 · arXiv
Self-consistency (SC) improves the performance of large language models (LLMs) across various tasks and domains that involve short content. However, does this support its effectiveness for long-context problems? We challenge the assumption that SC's benefits generalize to long-context settings, where LLMs often struggle with position bias, the systematic over-reliance on specific context regions-which hinders their ability to utilize information effectively from all parts of their context. Through comprehensive experimentation with varying state-of-the-art models, tasks, and SC formulations, w
Negative / Null Result ReportOpen accessComputer Science
Chunliang Li, Tianze Cao, Sanyuan Zhao · 2026 · arXiv
Visual Autoregressive (VAR) modeling inefficiently applies a fixed computational depth to each position when generating high-resolution images. While existing methods accelerate inference by pruning tokens using frequency maps, their binary hard-pruning approach is fundamentally limited and fails to improve quality even with better frequency estimation. Observing that VAR models possess significant depth redundancy, we propose a paradigm shift from pruning entire tokens to adaptively allocating per-token computational depth. To this end, we introduce DepthVAR, a training-free framework that dy
Negative / Null Result ReportOpen accessComputer Science
Amanpreet Singh, Mike D'Arcy, Arman Cohan et al. · 2022 · arXiv
Learned representations of scientific documents can serve as valuable input features for downstream tasks without further fine-tuning. However, existing benchmarks for evaluating these representations fail to capture the diversity of relevant tasks. In response, we introduce SciRepEval, the first comprehensive benchmark for training and evaluating scientific document representations. It includes 24 challenging and realistic tasks, 8 of which are new, across four formats: classification, regression, ranking and search. We then use this benchmark to study and improve the generalization ability o
Negative / Null Result ReportOpen accessComputer Science
Ruth Cohen, Lu Feng, Ayala Bloch et al. · 2026 · arXiv
While natural-language explanations from large language models (LLMs) are widely adopted to improve transparency and trust, their impact on objective human-AI team performance remains poorly understood. We identify a Persuasion Paradox: fluent explanations systematically increase user confidence and reliance on AI without reliably improving, and in some cases undermining, task accuracy. Across three controlled human-subject studies spanning abstract visual reasoning (RAVEN matrices) and deductive logical reasoning (LSAT problems), we disentangle the effects of AI predictions and explanations u
Negative / Null Result ReportOpen accessComputer Science
Andrei Liviu Nicolicioiu, Mohammad Pezeshki, Aaron Courville · 2026 · arXiv
On-policy self-distillation achieves strong pass@1 accuracy by using a single model as both teacher and student, with the teacher conditioned on a correct demonstration to provide dense token-level feedback. We show that this could come at a hidden cost: rollout diversity decreases and pass@k curves flatten (i.e., generating more rollouts fails to improve accuracy). We trace this to compounding biases in the design of self-distillation with sampled demonstrations. The teacher scores each student rollout while conditioned on a sampled correct rollout, channeling its feedback through the model's
Negative / Null Result ReportOpen accessComputer Science
OFM Riaz Rahman Aranya, Kevin Desai · 2026 · arXiv
Synthetic data, an appealing alternative to extensive expert-annotated data for medical image segmentation, consistently fails to improve segmentation performance despite its visual realism. The reason being that synthetic and real medical images exist in different semantic feature spaces, creating a domain gap that current semi-supervised learning methods cannot bridge. We propose SRA-Seg, a framework explicitly designed to align synthetic and real feature distributions for medical image segmentation. SRA-Seg introduces a similarity-alignment (SA) loss using frozen DINOv2 embeddings to pull s
Negative / Null Result ReportOpen accessEngineering
Sreyan Ghosh, Mohammad Sadegh Rasooli, Michael Levit et al. · 2024 · arXiv
Generative Error Correction (GEC) has emerged as a powerful post-processing method to enhance the performance of Automatic Speech Recognition (ASR) systems. However, we show that GEC models struggle to generalize beyond the specific types of errors encountered during training, limiting their ability to correct new, unseen errors at test time, particularly in out-of-domain (OOD) scenarios. This phenomenon amplifies with named entities (NEs), where, in addition to insufficient contextual information or knowledge about the NEs, novel NEs keep emerging. To address these issues, we propose DARAG (D
Negative / Null Result ReportMedicine
Yang, Mueller, D'Andrea et al. · 2026 · Academic psychiatry : the journal of the American Association of Directors of Psychiatric Residency Training and the Association for Academic Psychiatry
Resident physicians experience high rates of depression, anxiety, burnout, and loneliness, yet few evidence-based interventions have been evaluated in this population. This randomized controlled pilot trial examined the feasibility,…
View details →DOI: 10.1007/s40596-026-02384-y Negative / Null Result ReportOpen accessComputer Science
Julian Coda-Forno, Zhuokai Zhao, Qiang Zhang et al. · 2025 · arXiv
Should LLM reasoning live in a separate module, or within a single model's forward pass and representational space? We study dual-architecture latent reasoning, where a fluent Base exchanges latent messages with a Coprocessor, and test two hypotheses aimed at improving latent communication over Liu et al. (2024): (H1) increase channel capacity; (H2) learn communication via joint finetuning. Under matched latent-token budgets on GPT-2 and Qwen-3, H2 is consistently strongest while H1 yields modest gains. A unified soft-embedding baseline, a single model with the same forward pass and shared rep
Negative / Null Result ReportOpen accessComputer Science
Romain Cosentino, Sarath Shekkizhar, Adam Earle et al. · 2026 · arXiv
Negotiation requires more than inferring what the other side wants: it requires using that information to make advantageous offers and counteroffers over multiple turns. We study whether large language model (LLM) agents do this in a controlled multi-attribute bargaining environment. We find that current LLM agents can model a counterparty's preferences, but do not reliably turn that knowledge into strategic bargaining. When given negotiating partner preference information, agents model it accurately and early in their reasoning traces, yet this does not reliably improve outcomes for the infor
Negative / Null Result ReportOpen accessComputer Science
Mike Zhang, Ali Basirat, Desmond Elliott · 2026 · arXiv
Prior work establishes that controlled contrastiveness between self-generated responses from large language models, set via reward scores, improves downstream preference tuning in English. We extend this method to multiple languages and evaluate two models across a total of 14 high and low-resource languages on a diverse set of tasks. Our central finding is that cross-lingual contrastive preference tuning on self-generations (CroCo) transfers without language-specific preference annotation. A reward model trained on English preferences (atop a multilingual base) produces useful within-language
Negative / Null Result ReportOpen accessComputer Science
Zhiwei Jia, Xuanlin Li, Zhan Ling et al. · 2022 · arXiv
Generalization in deep reinforcement learning over unseen environment variations usually requires policy learning over a large set of diverse training variations. We empirically observe that an agent trained on many variations (a generalist) tends to learn faster at the beginning, yet its performance plateaus at a less optimal level for a long time. In contrast, an agent trained only on a few variations (a specialist) can often achieve high returns under a limited computational budget. To have the best of both worlds, we propose a novel generalist-specialist training framework. Specifically, w
Negative / Null Result ReportOpen accessComputer Science
Lily H. Zhang, Rajesh Ranganath · 2023 · arXiv
Methods which utilize the outputs or feature representations of predictive models have emerged as promising approaches for out-of-distribution (OOD) detection of image inputs. However, these methods struggle to detect OOD inputs that share nuisance values (e.g. background) with in-distribution inputs. The detection of shared-nuisance out-of-distribution (SN-OOD) inputs is particularly relevant in real-world applications, as anomalies and in-distribution inputs tend to be captured in the same settings during deployment. In this work, we provide a possible explanation for SN-OOD detection failur
Negative / Null Result ReportOpen accessComputer Science
Zhaofeng Wu, Shiqi Wang, Boya Peng et al. · 2026 · arXiv
Modern language models demonstrate impressive coding capabilities in common programming languages (PLs), such as C++ and Python, but their performance in lower-resource PLs is often limited by training data availability. In principle, however, most programming skills are universal across PLs, so the capability acquired in one PL should transfer to others. In this work, we propose the task of zero-shot cross-programming-language transfer for code RL. We find that, for Llama-3.1, RL training for code generation in a source PL fails to improve, and sometimes even degrades, the performance on othe
Negative / Null Result ReportOpen accessMathematics
Ziling Ma, Ángel López Oriona, Hernando Ombao et al. · 2026 · arXiv
We study adaptive pooling under predictive heterogeneity in high-dimensional multivariate time series forecasting, where global models improve statistical efficiency but may fail to capture heterogeneous predictive structure, while naive specialization can induce negative transfer. We formulate adaptive pooling as a statistical decision problem and propose a validation-driven framework that determines when and how specialization should be applied. Rather than grouping series based on representation similarity, we define partitions through out-of-sample predictive performance, thereby aligning
Negative / Null Result ReportOpen accessMathematics
Rong Chen, Simone Giannerini, Greta Goracci et al. · 2025 · arXiv
We develop an estimation methodology for a factor model for high-dimensional matrix-valued time series, where common stochastic trends and common stationary factors can be present. We study, in particular, the estimation of (row and column) loading spaces, of the common stochastic trends and of the common stationary factors, and the row and column ranks thereof. In a set of (negative) preliminary results, we show that a projection-based technique fails to improve the rates of convergence compared to a "flattened" estimation technique which does not take into account the matrix nature of the da
Negative / Null Result ReportOpen accessComputer Science
Minwu Kim, Anubhav Shrestha, Safal Shrestha et al. · 2025 · arXiv
Recent studies have shown that reinforcement learning with verifiable rewards (RLVR) enhances overall accuracy (pass@1) but often fails to improve capability (pass@k) of LLMs in reasoning tasks, while distillation can improve both. In this paper, we investigate the mechanisms behind these phenomena. First, we demonstrate that RLVR struggles to improve capability as it focuses on improving the accuracy of the easier questions to the detriment of the accuracy of the most difficult questions. Second, we show that RLVR does not merely increase the success probability for the easier questions, but
Negative / Null Result ReportOpen accessComputer Science
Hongsin Lee, Hye Won Chung · 2026 · arXiv
Adversarial Distillation aims to enhance student robustness by guiding the student with a robust teacher's soft labels within the min-max adversarial training framework, yet its success is notoriously inconsistent: a more robust teacher often fails to improve, or even harms, the student's robust generalization. In this paper, we identify a key mechanism of this teacher dependency: the misalignment between the teacher's supervisory confidence and the student's representational limitations on a consistent subset of training data -- the Robustly Unlearnable Set. We present a theoretical framework
Negative / Null Result ReportOpen accessEngineering
Jiaqi Wu, Jingyi Yuan, Yang Weng et al. · 2025 · arXiv
Power system voltage regulation is crucial to maintain power quality while integrating intermittent renewable resources in distribution grids. However, the system model on the grid edge is often unknown, making it difficult to model physical equations for optimal control. Therefore, previous work proposes structured data-driven methods like input convex neural networks (ICNN) for "optimal" control without relying on a physical model. While ICNNs offer theoretical guarantees based on restrictive assumptions of non-negative neural network parameters, can one improve the approximation power with
Negative / Null Result ReportOpen accessEngineering
Kwon Byung-Ki, Oh Hyun-Bin, Kim Jun-Seong et al. · 2023 · arXiv
Video motion magnification amplifies invisible small motions to be perceptible, which provides humans with a spatially dense and holistic understanding of small motions in the scene of interest. This is based on the premise that magnifying small motions enhances the legibility of motions. In the real world, however, vibrating objects often possess convoluted systems that have complex natural frequencies, modes, and directions. Existing motion magnification often fails to improve legibility since the intricate motions still retain complex characteristics even after being magnified, which may di
Negative / Null Result Report
Ye-seul Park, Ja-eun Kwak, Ho-ryong Yoo et al. · 2024 · The Journal of Internal Korean Medicine
Background: Unilateral vocal cord paralysis can occur due to various causes. Among them, unilateral vocal cord paralysis that takes place after endotracheal intubation is often caused by damage to the left recurrent nerve branch during…
View details →DOI: 10.22246/jikm.2024.45.6.1309 Negative / Null Result ReportMedicine
Fuentes-Expósito, Frid, Muñoz-Mateu et al. · 2026 · JCO clinical cancer informatics
We evaluated whether offering access to a multicomponent mHealth app improves quality of life (QoL) and psychosocial outcomes among breast cancer survivors under pragmatic, nonprescriptive conditions. In this single-center, randomized,…
View details →DOI: 10.1200/cci-26-00025 Negative / Null Result ReportMedicine
Yang, Zhu, Chen et al. · 2026 · Frontiers in neuroscience
To observe the therapeutic effects of deep transcranial magnetic stimulation (dTMS) and repetitive transcranial magnetic stimulation (rTMS) on upper and lower limb motor dysfunction in patients with basal ganglia infarction, and to…
View details →DOI: 10.3389/fnins.2026.1870537 Negative / Null Result ReportMedicine
Alshaibani, Kamadjaja, Sitalaksmi et al. · 2026 · Journal of molecular histology
Tooth extraction is a common procedure often followed by alveolar bone resorption, which may compromise future implant placement, prosthetic rehabilitation, esthetics, and periodontal support. Hydroxyapatite-chitosan (HA-Chi) scaffolds…
View details →DOI: 10.1007/s10735-026-10874-4 Negative / Null Result ReportMedicine
Nomura, Kimata, Ito et al. · 2026 · Diagnostics (Basel, Switzerland)
Background/Objectives: To directly compare the capabilities of hybrid-type iterative reconstruction (IR) with the newly developed deep learning reconstruction (DLR) for the inner ear on high-definition CT (HDCT) obtained using the…
View details →DOI: 10.3390/diagnostics16121756 Negative / Null Result ReportMedicine
Momii, Hata, Fudaba et al. · 2026 · Scientific reports
Glioblastoma continues to have a poor prognosis, although recent advancements in multimodal treatments have gradually improved outcomes. However, treatment options have become increasingly complex and highly specialized. In rural areas,…
View details →DOI: 10.1038/s41598-026-48867-8