Negative / Null Result ReportMedicine
Duman Aydin, Küçükosman, Mohamed et al. · 2026 · BMC medical education
Although oxygen therapy (OT) is a fundamental and life-saving intervention in the management of hypoxemia, it may lead to serious complications when applied incorrectly or in an uncontrolled manner. The aim of this study is to evaluate the…
View details →DOI: 10.1186/s12909-026-09766-8 Failed Experiment ReportOpen accessComputer Science
Zixian Huang, Kaichen Yang, Xu Huang et al. · 2026 · arXiv
A widely adopted strategy for model enhancement is to use synthetic data generated by a stronger model for supervised fine-tuning (SFT). However, for emerging reasoning models like Qwen3-8B, this approach often fails to improve reasoning capabilities and can even lead to a substantial drop in performance. In this work, we identify substantial stylistic divergence between teacher generated data and the distribution of student as a major factor impacting SFT. To bridge this gap, we propose a Teacher-Student Cooperation Data Synthesis framework (TESSY), which interleaves teacher and student model
Negative / Null Result ReportOpen accessComputer Science
Piyawat Lertvittayakumjorn, David Kinney, Vinodkumar Prabhakaran et al. · 2025 · arXiv
Generative large language models (LLMs) have demonstrated gaps in diverse cultural awareness across the globe. We investigate the effect of retrieval augmented generation and search-grounding techniques on LLMs' ability to display familiarity with various national cultures. Specifically, we compare the performance of standard LLMs, LLMs augmented with retrievals from a bespoke knowledge base (i.e., KB grounding), and LLMs augmented with retrievals from a web search (i.e., search grounding) on multiple cultural awareness benchmarks. We find that search grounding significantly improves the LLM p
Negative / Null Result ReportOpen accessEconomics, Econometrics and Finance
Francis X. Diebold, Maximilian Goebel, Philippe Goulet Coulombe · 2022 · arXiv
We use "glide charts" (plots of sequences of root mean squared forecast errors as the target date is approached) to evaluate and compare fixed-target forecasts of Arctic sea ice. We first use them to evaluate the simple feature-engineered linear regression (FELR) forecasts of Diebold and Goebel (2021), and to compare FELR forecasts to naive pure-trend benchmark forecasts. Then we introduce a much more sophisticated feature-engineered machine learning (FEML) model, and we use glide charts to evaluate FEML forecasts and compare them to a FELR benchmark. Our substantive results include the freque
Negative / Null Result ReportOpen accessComputer Science
Fırat Öncel, Matthias Bethge, Beyza Ermis et al. · 2024 · arXiv
In the last decade, the generalization and adaptation abilities of deep learning models were typically evaluated on fixed training and test distributions. Contrary to traditional deep learning, large language models (LLMs) are (i) even more overparameterized, (ii) trained on unlabeled text corpora curated from the Internet with minimal human intervention, and (iii) trained in an online fashion. These stark contrasts prevent researchers from transferring lessons learned on model generalization and adaptation in deep learning contexts to LLMs. To this end, our short paper introduces empirical ob
Negative / Null Result ReportOpen accessComputer Science
Adam Byerly, Daniel Khashabi · 2024 · arXiv
Self-consistency (SC) improves the performance of large language models (LLMs) across various tasks and domains that involve short content. However, does this support its effectiveness for long-context problems? We challenge the assumption that SC's benefits generalize to long-context settings, where LLMs often struggle with position bias, the systematic over-reliance on specific context regions-which hinders their ability to utilize information effectively from all parts of their context. Through comprehensive experimentation with varying state-of-the-art models, tasks, and SC formulations, w
Negative / Null Result ReportOpen accessComputer Science
Chunliang Li, Tianze Cao, Sanyuan Zhao · 2026 · arXiv
Visual Autoregressive (VAR) modeling inefficiently applies a fixed computational depth to each position when generating high-resolution images. While existing methods accelerate inference by pruning tokens using frequency maps, their binary hard-pruning approach is fundamentally limited and fails to improve quality even with better frequency estimation. Observing that VAR models possess significant depth redundancy, we propose a paradigm shift from pruning entire tokens to adaptively allocating per-token computational depth. To this end, we introduce DepthVAR, a training-free framework that dy
Negative / Null Result ReportOpen accessComputer Science
Chiara Lanza, Roberto Pereira, Marco Miozzo et al. · 2026 · arXiv
Centralized training is the standard paradigm in deep learning, enabling models to learn from a unified dataset in a single location. In such setup, isotropic feature distributions naturally arise as a mean to support well-structured and generalizable representations. In contrast, continual learning operates on streaming and non-stationary data, and trains models incrementally, inherently facing the well-known plasticity-stability dilemma. In such settings, learning dynamics tends to yield increasingly anisotropic feature space. This arises a fundamental question: should isotropy be enforced t
Negative / Null Result ReportMedicine
Zhou, Cao, You et al. · 2026 · BMC oral health
Photogrammetry technique may provide a promising approach compared to conventional techniques for multiple implants. However, the accuracy of photogrammetric technique for implant-supported fixed complete dentures in clinical scenarios…
View details →DOI: 10.1186/s12903-026-09003-0 Negative / Null Result ReportOpen accessComputer Science
Amanpreet Singh, Mike D'Arcy, Arman Cohan et al. · 2022 · arXiv
Learned representations of scientific documents can serve as valuable input features for downstream tasks without further fine-tuning. However, existing benchmarks for evaluating these representations fail to capture the diversity of relevant tasks. In response, we introduce SciRepEval, the first comprehensive benchmark for training and evaluating scientific document representations. It includes 24 challenging and realistic tasks, 8 of which are new, across four formats: classification, regression, ranking and search. We then use this benchmark to study and improve the generalization ability o
Negative / Null Result ReportOpen accessComputer Science
Hanlin Xiao, Rainer Breitling, Eriko Takano et al. · 2026 · arXiv
Recent advances in general-purpose foundation models have stimulated the development of large biological sequence models. While natural language shows symbolic granularity (characters, words, sentences), biological sequences exhibit hierarchical granularity whose levels (nucleotides, amino acids, protein domains, genes) further encode biologically functional information. In this paper, we investigate the integration of cross-granularity knowledge from models through a case study of BiGCARP, a Pfam domain-level model for biosynthetic gene clusters, and ESM, an amino acid-level protein language
Negative / Null Result ReportOpen accessComputer Science
Ruth Cohen, Lu Feng, Ayala Bloch et al. · 2026 · arXiv
While natural-language explanations from large language models (LLMs) are widely adopted to improve transparency and trust, their impact on objective human-AI team performance remains poorly understood. We identify a Persuasion Paradox: fluent explanations systematically increase user confidence and reliance on AI without reliably improving, and in some cases undermining, task accuracy. Across three controlled human-subject studies spanning abstract visual reasoning (RAVEN matrices) and deductive logical reasoning (LSAT problems), we disentangle the effects of AI predictions and explanations u
Negative / Null Result ReportOpen accessComputer Science
Andrei Liviu Nicolicioiu, Mohammad Pezeshki, Aaron Courville · 2026 · arXiv
On-policy self-distillation achieves strong pass@1 accuracy by using a single model as both teacher and student, with the teacher conditioned on a correct demonstration to provide dense token-level feedback. We show that this could come at a hidden cost: rollout diversity decreases and pass@k curves flatten (i.e., generating more rollouts fails to improve accuracy). We trace this to compounding biases in the design of self-distillation with sampled demonstrations. The teacher scores each student rollout while conditioned on a sampled correct rollout, channeling its feedback through the model's
Negative / Null Result ReportOpen accessComputer Science
OFM Riaz Rahman Aranya, Kevin Desai · 2026 · arXiv
Synthetic data, an appealing alternative to extensive expert-annotated data for medical image segmentation, consistently fails to improve segmentation performance despite its visual realism. The reason being that synthetic and real medical images exist in different semantic feature spaces, creating a domain gap that current semi-supervised learning methods cannot bridge. We propose SRA-Seg, a framework explicitly designed to align synthetic and real feature distributions for medical image segmentation. SRA-Seg introduces a similarity-alignment (SA) loss using frozen DINOv2 embeddings to pull s
Failed Experiment ReportOpen accessComputer Science
Chen-Rong Liu, Chuang Li, Runxia Tao et al. · 2026 · arXiv
Conventional noise analysis in atomic-ensemble sensing assumes a continuous-medium approximation, thereby treating the atomic system as a deterministic dielectric. Here, we demonstrate that this assumption breaks down due to the discrete, particulate nature of the ensemble, giving rise to an intrinsic "atomic granularity noise" (AGN) that fundamentally competes with the optical measurement noise (OMN, typically photon shot noise). By introducing a discrete-atom statistical framework, we derive a unified noise-scaling law governed by a single dimensionless resource ratio, $\mathcal{R} = \bar{N}
Negative / Null Result ReportOpen accessEngineering
Sreyan Ghosh, Mohammad Sadegh Rasooli, Michael Levit et al. · 2024 · arXiv
Generative Error Correction (GEC) has emerged as a powerful post-processing method to enhance the performance of Automatic Speech Recognition (ASR) systems. However, we show that GEC models struggle to generalize beyond the specific types of errors encountered during training, limiting their ability to correct new, unseen errors at test time, particularly in out-of-domain (OOD) scenarios. This phenomenon amplifies with named entities (NEs), where, in addition to insufficient contextual information or knowledge about the NEs, novel NEs keep emerging. To address these issues, we propose DARAG (D
Negative / Null Result ReportMedicine
Yang, Mueller, D'Andrea et al. · 2026 · Academic psychiatry : the journal of the American Association of Directors of Psychiatric Residency Training and the Association for Academic Psychiatry
Resident physicians experience high rates of depression, anxiety, burnout, and loneliness, yet few evidence-based interventions have been evaluated in this population. This randomized controlled pilot trial examined the feasibility,…
View details →DOI: 10.1007/s40596-026-02384-y Negative / Null Result ReportOpen accessComputer Science
Guowei Liu, Hongming Li, Yaning Guo et al. · 2026 · arXiv
Deploying large-scale MoE models presents challenges in memory capacity and bandwidth for expert activation. While Attention-FFN Disaggregation (AFD) has emerged as a potential architecture to decouple compute and memory resources, its performance boundaries compared to standard large-scale Expert Parallelism (EP) remain underexplored. In this paper, we conduct a systematic analysis of AFD by extending the roofline model to the communication level, correlating interconnect bandwidth, arithmetic intensity, and Hardware FLOPS Utilization (HFU). Our analysis reveals a dead zone on standard cluste
Negative / Null Result ReportOpen accessComputer Science
Julian Coda-Forno, Zhuokai Zhao, Qiang Zhang et al. · 2025 · arXiv
Should LLM reasoning live in a separate module, or within a single model's forward pass and representational space? We study dual-architecture latent reasoning, where a fluent Base exchanges latent messages with a Coprocessor, and test two hypotheses aimed at improving latent communication over Liu et al. (2024): (H1) increase channel capacity; (H2) learn communication via joint finetuning. Under matched latent-token budgets on GPT-2 and Qwen-3, H2 is consistently strongest while H1 yields modest gains. A unified soft-embedding baseline, a single model with the same forward pass and shared rep
Negative / Null Result ReportOpen accessComputer Science
Sepanta Zeighami, Cyrus Shahabi · 2024 · arXiv
While extremely useful (e.g., for COVID-19 forecasting and policy-making, urban mobility analysis and marketing, and obtaining business insights), location data collected from mobile devices often contain data from a biased population subset, with some communities over or underrepresented in the collected datasets. As a result, aggregate statistics calculated from such datasets (as is done by various companies including Safegraph, Google, and Facebook), while ignoring the bias, leads to an inaccurate representation of population statistics. Such statistics will not only be generally inaccurate
Negative / Null Result ReportOpen accessComputer Science
Dong Xu, Jiantao Wu, Qihua Pan et al. · 2026 · arXiv
Drug-drug interaction (DDI) prediction is central to drug discovery and clinical development, particularly in the context of increasingly prevalent polypharmacy. Although existing computational methods achieve strong performance on standard benchmarks, they often fail to generalize to realistic deployment scenarios, where most candidate drug pairs involve previously unseen drugs and validated interactions are scarce. We demonstrate that proximity in the embedding spaces of prevailing molecule-centric DDI models does not reliably correspond to interaction labels, and that simply scaling up mode
Negative / Null Result ReportOpen accessComputer Science
Romain Cosentino, Sarath Shekkizhar, Adam Earle et al. · 2026 · arXiv
Negotiation requires more than inferring what the other side wants: it requires using that information to make advantageous offers and counteroffers over multiple turns. We study whether large language model (LLM) agents do this in a controlled multi-attribute bargaining environment. We find that current LLM agents can model a counterparty's preferences, but do not reliably turn that knowledge into strategic bargaining. When given negotiating partner preference information, agents model it accurately and early in their reasoning traces, yet this does not reliably improve outcomes for the infor
Negative / Null Result ReportOpen accessComputer Science
Mike Zhang, Ali Basirat, Desmond Elliott · 2026 · arXiv
Prior work establishes that controlled contrastiveness between self-generated responses from large language models, set via reward scores, improves downstream preference tuning in English. We extend this method to multiple languages and evaluate two models across a total of 14 high and low-resource languages on a diverse set of tasks. Our central finding is that cross-lingual contrastive preference tuning on self-generations (CroCo) transfers without language-specific preference annotation. A reward model trained on English preferences (atop a multilingual base) produces useful within-language
Negative / Null Result ReportOpen accessComputer Science
Aaron Baier-Reinio, Hans De Sterck · 2020 · arXiv
We use neural ordinary differential equations to formulate a variant of the Transformer that is depth-adaptive in the sense that an input-dependent number of time steps is taken by the ordinary differential equation solver. Our goal in proposing the N-ODE Transformer is to investigate whether its depth-adaptivity may aid in overcoming some specific known theoretical limitations of the Transformer in handling nonlocal effects. Specifically, we consider the simple problem of determining the parity of a binary sequence, for which the standard Transformer has known limitations that can only be ove
Negative / Null Result ReportMedicine
Escoffier, Hedhli, Campos-Juanatey et al. · 2026 · The French journal of urology
Artificial intelligence (AI) is increasingly used in surgery, but its role in reconstructive urology remains insufficiently studied. The aim of this work was to evaluate the theoretical knowledge of several AI platforms and compare their…
View details →DOI: 10.1016/j.fjurol.2026.103150 Negative / Null Result ReportOpen accessComputer Science
Rui Xing, Qi Chai, Jie Ma et al. · 2026 · arXiv
Hate speech online targets individuals or groups based on identity attributes and spreads rapidly, posing serious social risks. Memes, which combine images and text, have emerged as a nuanced vehicle for disseminating hate speech, often relying on cultural knowledge for interpretation. However, existing multimodal hate speech datasets suffer from coarse-grained labeling and a lack of integration with surrounding discourse, leading to imprecise and incomplete assessments. To bridge this gap, we propose an agentic annotation framework that coordinates seven specialized agents to generate hierarc
Negative / Null Result ReportOpen accessComputer Science
Zhiwei Jia, Xuanlin Li, Zhan Ling et al. · 2022 · arXiv
Generalization in deep reinforcement learning over unseen environment variations usually requires policy learning over a large set of diverse training variations. We empirically observe that an agent trained on many variations (a generalist) tends to learn faster at the beginning, yet its performance plateaus at a less optimal level for a long time. In contrast, an agent trained only on a few variations (a specialist) can often achieve high returns under a limited computational budget. To have the best of both worlds, we propose a novel generalist-specialist training framework. Specifically, w
Negative / Null Result ReportOpen accessComputer Science
Jiamin Xu, Jacqueline Maasch, Kyra Gan · 2026 · arXiv
Online reinforcement learning (RL) relies on the Markov property for guaranteed performance, but real-world applications often lack well-defined states given raw observed variables. While causal RL has attracted growing interest, existing work typically assumes Markovian states are provided and focuses on using causality to accelerate learning, leaving a fundamental gap: \emph{given a longitudinal causal graph over observed variables, how does one construct MDP states that provably satisfy the Markov property?} We address this by providing a procedure that constructs a provably minimal state r
Negative / Null Result ReportOpen accessComputer Science
Lily H. Zhang, Rajesh Ranganath · 2023 · arXiv
Methods which utilize the outputs or feature representations of predictive models have emerged as promising approaches for out-of-distribution (OOD) detection of image inputs. However, these methods struggle to detect OOD inputs that share nuisance values (e.g. background) with in-distribution inputs. The detection of shared-nuisance out-of-distribution (SN-OOD) inputs is particularly relevant in real-world applications, as anomalies and in-distribution inputs tend to be captured in the same settings during deployment. In this work, we provide a possible explanation for SN-OOD detection failur
Negative / Null Result ReportOpen accessComputer Science
Zhaofeng Wu, Shiqi Wang, Boya Peng et al. · 2026 · arXiv
Modern language models demonstrate impressive coding capabilities in common programming languages (PLs), such as C++ and Python, but their performance in lower-resource PLs is often limited by training data availability. In principle, however, most programming skills are universal across PLs, so the capability acquired in one PL should transfer to others. In this work, we propose the task of zero-shot cross-programming-language transfer for code RL. We find that, for Llama-3.1, RL training for code generation in a source PL fails to improve, and sometimes even degrades, the performance on othe