Failed Experiment ReportOpen accessComputer Science
Xinzhe Chen, Jianjiang Li · 2023 · arXiv
The escalating surge in data generation presents formidable challenges to information technology, necessitating advancements in storage, retrieval, and utilization. With the proliferation of artificial intelligence and big data, the "Data Age 2025" report forecasts an exponential increase in global data production. The escalating data volumes raise concerns about efficient data processing. The paper addresses the predicament of achieving a lower compression ratio while maintaining or surpassing the compression performance of state-of-the-art techniques. This paper introduces a lossy compressio
View details →Null DatasetOpen accessComputer Science
Junhan Liu, Shile Liu, Chenming Xu et al. · 2023 · arXiv
Automated Guided Vehicles (AGVs) operate in synergy to execute specific tasks. These vehicles exchange information to ensure seamless collaboration, prevent collisions, and eliminate task redundancy. The advent of blockchain technology offers a promising avenue for establishing a secure and dependable communication infrastructure for AGVs. Nonetheless, it becomes imperative for AGVs to adopt efficient data transmission methodologies, especially when interacting with the dynamic nature of blockchain infrastructure where data undergoes frequent modifications. In the present study, we introduce a
View details →Failed Experiment ReportOpen accessComputer Science
Liyun Zhang, Jiayi Guo · 2026 · arXiv
We document an empirical phenomenon in chain-of-thought and ReAct agents driven by ten large language models from seven architecture families: meaning-bearing perturbations (e.g., paraphrase, synonym) alter final answers more often than presentation perturbations (e.g., formatting, reordering) of comparable severity. Across 68 cells spanning GSM8K, MATH, and HotpotQA (1,530 originals and $\sim$11,150 variants), the inconsistency gap averages +19.69 pp after severity matching (paired $t=9.58$, $p<0.0001$), with 64/68 cells positive. The gap survives four severity-proxy audits and remains signif
View details →Negative / Null Result ReportOpen accessComputer Science
Rohit Agarwala, Kalyan Dey · 2025 · arXiv
We employ PYTHIA8 simulations to study forward-backward (FB) correlations in pp collisions at LHC energies, probing non-perturbative QCD dynamics via color reconnection (CR) and QCD radiation (ISR/FSR). Using \texttt{PYTHIA8} (v8.311) under ALICE/ATLAS kinematics, we analyze: extensive: FB multiplicity ($b_{\rm corr}^{\rm mult}$) and summed $p_{\rm T}$ ($b_{\rm corr}^{\sum p_{\rm T}}$) correlations; intensive: FB mean $p_{\rm T}$ ($b_{\rm corr}^{\overline p_{\rm T}}$) correlations; and \textit{strongly intensive:} $Σ_{\rm N_F N_B}$ quantity in symmetric pseudorapidity intervals, validated agai
View details →Failed Experiment ReportOpen accessComputer Science
Jayoo Hwang, Xiaowen Zhang, Vedant Padwal · 2026 · arXiv
Autonomous web navigation remains challenging for LLM agents, and the strongest generalist systems rely on proprietary reasoning models whose inference cost is prohibitive for the repetitive tasks where such agents would be most useful. We argue this gap stems not from insufficient model capability but from agent architectures that fail to replicate three human cognitive advantages: selective attention to relevant page regions, persistent memory of website structure, and procedural fluency with common interaction patterns. We introduce WebChallenger, a web agent framework that addresses each g
View details →Negative / Null Result ReportOpen accessComputer Science
Keith Harrigian, Tina Tang, Anthony Gonzales et al. · 2023 · arXiv
Diabetic eye disease is a major cause of blindness worldwide. The ability to monitor relevant clinical trajectories and detect lapses in care is critical to managing the disease and preventing blindness. Alas, much of the information necessary to support these goals is found only in the free text of the electronic medical record. To fill this information gap, we introduce a system for extracting evidence from clinical text of 19 clinical concepts related to diabetic eye disease and inferring relevant attributes for each. In developing this ophthalmology phenotyping system, we are also afforded
View details →Negative / Null Result ReportOpen accessComputer Science
Shaoxuan Zhou, Yafei Sun, Jing Zhang et al. · 2026 · arXiv
Short-video platforms like Douyin and Kwai have become central to adolescent digital life, but they also risk exposing teens to algorithmically amplified harmful content. Despite its societal importance, the scale, mechanisms, and real-world impact of this exposure remain poorly understood. Measuring it is challenging: recommendation feeds are personalized black boxes, harmful content employs sophisticated evasion tactics, and naive crawlers fail to replicate authentic teen behavior. To bridge this gap, we propose PHTV-Scout, the first large-scale, behaviorally grounded measurement framework f
View details →Negative / Null Result ReportOpen accessComputer Science
Ali Ebrahimpour-Boroojeny · 2025 · arXiv
We propose new methodologies for both unlearning random set of samples and class unlearning and show that they outperform existing methods. The main driver of our unlearning methods is the similarity of predictions to a retrained model on both the forget and remain samples. We introduce Adversarial Machine UNlearning (AMUN), which surpasses prior state-of-the-art methods for image classification based on SOTA MIA scores. AMUN lowers the model's confidence on forget samples by fine-tuning on their corresponding adversarial examples. Through theoretical analysis, we identify factors governing AM
View details →Negative / Null Result ReportOpen accessComputer Science
Nilesh Kumar Sahu, Snehil Gupta, Haroon R Lone · 2025 · arXiv
Social Anxiety Disorder (SAD) is a widespread mental health condition, yet its lack of objective markers hinders timely detection and intervention. While previous research has focused on behavioral and non-verbal markers of SAD in structured activities (e.g., speeches or interviews), these settings fail to replicate real-world, unstructured social interactions fully. Identifying non-verbal markers in naturalistic, unstaged environments is essential for developing ubiquitous and non-intrusive monitoring solutions. To address this gap, we present AnxietyFaceTrack, a study leveraging facial video
View details →Negative / Null Result ReportOpen accessComputer Science
Syed Kazmi, Berk Gorgulu, Mucahit Cevik et al. · 2023 · arXiv
Wind power forecasting helps with the planning for the power systems by contributing to having a higher level of certainty in decision-making. Due to the randomness inherent to meteorological events (e.g., wind speeds), making highly accurate long-term predictions for wind power can be extremely difficult. One approach to remedy this challenge is to utilize weather information from multiple points across a geographical grid to obtain a holistic view of the wind patterns, along with temporal information from the previous power outputs of the wind farms. Our proposed CNN-RNN architecture combine
View details →Negative / Null Result ReportOpen accessComputer Science
Emily L. Aiken, Andre T. Nguyen, Mauricio Santillana · 2019 · arXiv
We introduce the use of a Gated Recurrent Unit (GRU) for influenza prediction at the state- and city-level in the US, and experiment with the inclusion of real-time flu-related Internet search data. We find that a GRU has lower prediction error than current state-of-the-art methods for data-driven influenza prediction at time horizons of over two weeks. In contrast with other machine learning approaches, the inclusion of real-time Internet search data does not improve GRU predictions.
View details →Negative / Null Result ReportOpen accessComputer Science
Sungmin Cha, Kyunghyun Cho · 2024 · arXiv
Continual learning (CL) aims to train a model on a sequence of tasks (i.e., a CL scenario) while balancing the trade-off between plasticity (learning new tasks) and stability (retaining prior knowledge). The dominantly adopted conventional evaluation protocol for CL algorithms selects the best hyperparameters (e.g., learning rate, mini-batch size, regularization strengths, etc.) within a given scenario and then evaluates the algorithms using these hyperparameters in the same scenario. However, this protocol has significant shortcomings: it overestimates the CL capacity of algorithms and relies
View details →Negative / Null Result ReportOpen accessComputer Science
Jiarui Xie, Mutahar Safdar, Andrei Mircea et al. · 2024 · arXiv
Machine learning (ML)-based cyber-physical systems (CPSs) have been extensively developed to improve the print quality of additive manufacturing (AM). However, the reproducibility of these systems, as presented in published research, has not been thoroughly investigated due to a lack of formal evaluation methods. Reproducibility, a critical component of trustworthy artificial intelligence, is achieved when an independent team can replicate the findings or artifacts of a study using a different experimental setup and achieve comparable performance. In many publications, critical information nec
View details →Negative / Null Result ReportOpen accessComputer Science
Xun Liang, Huayi Lai, Hanyu Wang et al. · 2025 · arXiv
Large language models (LLMs) have gained significant traction in medical decision support systems, particularly in the context of medical question answering and role-playing simulations. A common practice, Prompt-Based Role Playing (PBRP), instructs models to adopt different clinical roles (e.g., medical students, residents, attending physicians) to simulate varied professional behaviors. However, the impact of such role prompts on model reasoning capabilities remains unclear. This study introduces the RP-Neuron-Activated Evaluation Framework(RPNA) to evaluate whether role prompts induce disti
View details →Negative / Null Result ReportOpen accessComputer Science
Sarah Ball, Simeon Allmendinger, Frauke Kreuter et al. · 2025 · arXiv
Generative AI (GenAI) is increasingly used in survey contexts to simulate human preferences. While many research endeavors evaluate the quality of synthetic GenAI data by comparing model-generated responses to gold-standard survey results, fundamental questions about the validity and reliability of using LLMs as substitutes for human respondents remain. Our study provides a technical analysis of how demographic attributes and prompt variations influence latent opinion mappings in large language models (LLMs) and evaluates their suitability for survey-based predictions. Using 14 different model
View details →Negative / Null Result ReportOpen accessComputer Science
Mert Albaba, Sammy Christen, Thomas Langarek et al. · 2024 · arXiv
Acquiring complex behaviors is essential for artificially intelligent agents, yet learning these behaviors in high-dimensional settings poses a significant challenge due to the vast search space. Traditional reinforcement learning (RL) requires extensive manual effort for reward function engineering. Inverse reinforcement learning (IRL) uncovers reward functions from expert demonstrations but relies on an iterative process that is often computationally expensive. Imitation learning (IL) provides a more efficient alternative by directly comparing an agent's actions to expert demonstrations; how
View details →Failed Experiment ReportOpen accessComputer Science
Tilman Plehn, Michael Spannowsky, Michihisa Takeuchi · 2011 · arXiv
In time for the first tests on LHC data we introduce a set of improvements and tests of purely kinematic top tagging algorithms. First, we show how different jet algorithms can be used for different transverse momentum regimes. Combining pruning and filtering in the reconstruction can enhance the signal over background ratio significantly, while larger jet radii only give minor improvements. Finally, bottom tagging can be added to the top tagger, but at least for the HEPTopTagger does not improve the kinematic selection algorithm.
View details →Methods Dead-EndOpen accessComputer Science
Aytaç Sekmen, Fatih Emre Gunes, Furkan Horoz et al. · 2026 · arXiv
Monocular Depth Estimation (MDE) is crucial for autonomous lunar rover navigation using electro-optical cameras. However, deploying terrestrial MDE networks to the Moon brings a severe domain gap due to harsh shadows, textureless regolith, and zero atmospheric scattering. Existing evaluations rely on analogs that fail to replicate these conditions and lack actual metric ground truth. To address this, we present LuMon, a comprehensive benchmarking framework to evaluate MDE methods for lunar exploration. We introduce novel datasets featuring high-quality stereo ground truth depth from the real C
View details →Negative / Null Result ReportOpen accessComputer Science
Xin Ding, Yongwei Wang, Zuheng Xu · 2023 · arXiv
Continuous Conditional Generative Adversarial Networks (CcGANs) enable generative modeling conditional on continuous scalar variables (termed regression labels). However, they can produce subpar fake images due to limited training data. Although Negative Data Augmentation (NDA) effectively enhances unconditional and class-conditional GANs by introducing anomalies into real training images, guiding the GANs away from low-quality outputs, its impact on CcGANs is limited, as it fails to replicate negative samples that may occur during the CcGAN sampling. We present a novel NDA approach called Dua
View details →Negative / Null Result ReportOpen accessComputer Science
Rohan Jha, Reno Kriz, Benjamin Van Durme · 2026 · arXiv
The XTR (conteXtual Token Retrieval) algorithm is a modification to ColBERT retrieval that avoids the costly step of fully gathering and reranking the candidates' embeddings by imputing their missing similarity scores from the initial token retrieval step. The original work proposes a modified training objective as necessary for effective XTR retrieval, arguing that standard ColBERT token scoring is unsuitable for imputation. In this paper, we replicate both the XTR retrieval algorithm and its modified training objective, and extend the evaluation to knowledge-distillation (KD) training and ef
View details →Negative / Null Result ReportOpen accessComputer Science
Angelo Di Porzio, Marco Coraggio · 2025 · arXiv
The deployment of autonomous virtual avatars (in extended reality) and robots in human group activities -- such as rehabilitation therapy, sports, and manufacturing -- is expected to increase as these technologies become more pervasive. Designing cognitive architectures and control strategies to drive these agents requires realistic models of human motion. Furthermore, recent research has shown that each person exhibits a unique velocity signature, highlighting how individual motor behaviors are both rich in variability and internally consistent. However, existing models only provide simplifie
View details →Negative / Null Result ReportOpen accessComputer Science
Tianjiao Cao, Jiahao Lyu, Weichao Zeng et al. · 2025 · arXiv
Scene text detection has seen the emergence of high-performing methods that excel on academic benchmarks. However, these detectors often fail to replicate such success in real-world scenarios. We uncover two key factors contributing to this discrepancy through extensive experiments. First, a \textit{Fine-tuning Gap}, where models leverage \textit{Dataset-Specific Optimization} (DSO) paradigm for one domain at the cost of reduced effectiveness in others, leads to inflated performances on academic benchmarks. Second, the suboptimal performance in practical settings is primarily attributed to the
View details →Negative / Null Result ReportOpen accessComputer Science
Christopher Brady, Xu Wu · 2025 · arXiv
The Organization for Economic Cooperation and Development (OECD) Working Party on Nuclear Criticality Safety (WPNCS) proposed a benchmark exercise to assess the performance of current nuclear data adjustment techniques applied to nonlinear applications and experiments with low correlation to applications. This work introduces Bayesian Inverse Uncertainty Quantification (IUQ) as a method for nuclear data adjustments in this benchmark, and compares IUQ to the more traditional methods of Generalized Linear Least Squares (GLLS) and Monte Carlo Bayes (MOCABA). Posterior predictions from IUQ showed
View details →Negative / Null Result ReportOpen accessComputer Science
Meimingwei Li, Yuanhao Ding, Esteban Garces Arias et al. · 2026 · arXiv
Recent work has identified a counterintuitive phenomenon termed "Hyperfitting", where fine-tuning Large Language Models (LLMs) to near-zero training loss on small datasets surprisingly enhances open-ended generation quality and mitigates repetition in greedy decoding. While effective, the underlying mechanism remains poorly understood, with the extremely low-entropy output distributions suggesting a potential equivalence to simple temperature scaling. In this work, we demonstrate that this phenomenon is fundamentally distinct from distribution sharpening; entropy-matched control experiments re
View details →Negative / Null Result ReportOpen accessComputer Science
Ping Chen, Zezhou Chen, Xingpeng Zhang et al. · 2026 · arXiv
Current 2D-to-3D conversion methods achieve geometric accuracy but are artistically deficient, failing to replicate the immersive and emotionally resonant experience of professional 3D cinema. This is because geometric reconstruction paradigms mistake deliberate artistic intent, such as strategic zero-plane shifts for pop-out effects and local depth sculpting, for data noise or ambiguity. This paper argues for a new paradigm: Artistic Disparity Synthesis, shifting the goal from physically accurate disparity estimation to artistically coherent disparity synthesis. We propose Art3D, a preliminar
View details →Negative / Null Result ReportOpen accessComputer Science
Chenglong Ma, Ziqi Xu, Yongli Ren et al. · 2025 · arXiv
Traditional offline evaluation methods for recommender systems struggle to capture the complexity of modern platforms due to sparse behavioural signals, noisy data, and limited modelling of user personality traits. While simulation frameworks can generate synthetic data to address these gaps, existing methods fail to replicate behavioural diversity, limiting their effectiveness. To overcome these challenges, we propose the Personality-driven User Behaviour Simulator (PUB), an LLM-based simulation framework that integrates the Big Five personality traits to model personalised user behaviour. PU
View details →Negative / Null Result ReportOpen accessComputer Science
Hirotaka Tahara, Hikaru Sasaki, Hanbit Oh et al. · 2022 · arXiv
Robust imitation learning using disturbance injections overcomes issues of limited variation in demonstrations. However, these methods assume demonstrations are optimal, and that policy stabilization can be learned via simple augmentations. In real-world scenarios, demonstrations are often of diverse-quality, and disturbance injection instead learns sub-optimal policies that fail to replicate desired behavior. To address this issue, this paper proposes a novel imitation learning framework that combines both policy robustification and optimal demonstration learning. Specifically, this combinato
View details →Negative / Null Result ReportOpen accessComputer Science
Damian Ruck, R. Alexander Bentley, Alberto Acerbi et al. · 2017 · arXiv
Here we test Neutral models against the evolution of English word frequency and vocabulary at the population scale, as recorded in annual word frequencies from three centuries of English language books. Against these data, we test both static and dynamic predictions of two neutral models, including the relation between corpus size and vocabulary size, frequency distributions, and turnover within those frequency distributions. Although a commonly used Neutral model fails to replicate all these emergent properties at once, we find that modified two-stage Neutral model does replicate the static a
View details →Negative / Null Result ReportOpen accessComputer Science
Zichao Wang, Alexa Siu · 2026 · arXiv
Large language models (LLMs) have shown strong performance on standardized social science instruments, but their value for product discovery remains unclear. We investigate whether interview-informed generative agents can simulate user responses in concept testing scenarios. Using in-depth workflow interviews with knowledge workers, we created personalized agents and compared their evaluations of novel AI concepts against the same participants' responses. Our results show that agents are distribution-calibrated but identity-imprecise: they fail to replicate the specific individual they are gro
View details →Negative / Null Result ReportOpen accessComputer Science
Shintaro Sakai, Jisun An, Migyeong Kang et al. · 2025 · arXiv
Prior clinical psychology research shows that Western individuals with depression tend to report psychological symptoms, while Eastern individuals report somatic ones. We test whether Large Language Models (LLMs), which are increasingly used in mental health, reproduce these cultural patterns by prompting them with Western or Eastern personas. Results show that LLMs largely fail to replicate the patterns when prompted in English, though prompting in major Eastern languages (i.e., Chinese, Japanese, and Hindi) improves alignment in several configurations. Our analysis pinpoints two key reasons
View details →