Negative / Null Result ReportOpen accessComputer Science
Morris de Haan, Philipp Hager · 2024 · arXiv
Despite the popularity of the two-tower model for unbiased learning to rank (ULTR) tasks, recent work suggests that it suffers from a major limitation that could lead to its collapse in industry applications: the problem of logging policy confounding. Several potential solutions have even been proposed; however, the evaluation of these methods was mostly conducted using semi-synthetic simulation experiments. This paper bridges the gap between theory and practice by investigating the confounding problem on the largest real-world dataset, Baidu-ULTR. Our main contributions are threefold: 1) we s
View details →Negative / Null Result ReportOpen accessComputer Science
Nils Dycke, Iryna Gurevych · 2025 · arXiv
Large Language Models (LLMs) have great potential to accelerate and support scholarly peer review and are increasingly used as fully automatic review generators (ARGs). However, potential biases and systematic errors may pose significant risks to scientific integrity; understanding the specific capabilities and limitations of state-of-the-art ARGs is essential. We focus on a core reviewing skill that underpins high-quality peer review: detecting faulty research logic. This involves evaluating the internal consistency between a paper's results, interpretations, and claims. We present a fully au
View details →Negative / Null Result ReportOpen accessComputer Science
Bryan Y. Siow · 2025 · arXiv
This paper describes a practical approach of using supervised machine learning (ML) models to assist safety investigators to classify aviation occurrences into either incident or serious incident categories. Our implementation currently deployed as a ML web application is trained on a labelled dataset derived from publicly available aviation investigation reports. A selection of five supervised learning models (Support Vector Machine, Logistic Regression, Random Forest Classifier, XGBoost and K-Nearest Neighbors) were evaluated. This paper showed the best performing ML algorithm was the Random
View details →Negative / Null Result ReportOpen accessComputer Science
Mitsuru Kakizaki, Shigeki Matsumoto, Yoshio Sato et al. · 2005 · arXiv
We point out that Kaluza-Klein (KK) dark matter physics is drastically affected by second KK particles. In this work various interesting phenomena caused by the second KK modes are discussed. In particular, we reevaluate the annihilation cross section and thermal relic density of the KK dark matter quantitatively in universal extra dimensions, in which all the standard model particles propagate. In these models, the first KK mode of $B$ boson is a viable dark matter candidate by virtue of KK-parity. We demonstrate that the KK dark matter annihilation cross section can be enhanced, compared wit
View details →Negative / Null Result ReportOpen accessComputer Science
Shaaban Khalil · 1999 · arXiv
The cosmological relic density of the lightest supersymmetric particle (LSP) of type I string derived model is calculated. This model can accommodate large values of CP violating phases, and the electron and neutron electric dipole moments satisfy the experimental constraint. We show that the constraint from the electric dipole moment on the ratio between the gaugino masses implies that the mass of the LSP, which is bino like, is close to the lightest chargino. The co-annihilation between them is very important to reduce the LSP relic density to an interesting region. We show that the SUSY pha
View details →Negative / Null Result ReportOpen accessComputer Science
Klaudia Guzij, Michael Fröhlich, Florian Fincke et al. · 2022 · arXiv
The voluntary carbon market is an important building block in the fight against climate change. However, it is not trivial for consumers to verify whether carbon offset projects deliver what they promise. While technical solutions for measuring their impact are emerging, there is a lack of understanding of how to translate this data into interface designs that mediate the establishment of trust. With interaction between users and offset projects mainly happening online, it is critical to meet this design challenge. To this end, we designed and evaluated interfaces with varying trust cues for c
View details →Negative / Null Result ReportOpen accessComputer Science
Peggy Pei-Ying Lu, Makoto Konishi, Shin Sano et al. · 2022 · arXiv
Recent advances in technology have allowed an automation system to recognize its errors and repair trust more actively than ever. While previous research has called for further studies of different human factors and design features, their effect on human-automation trust repair scenarios remains unknown, especially concerning emotions. This paper seeks to fill such gaps by investigating the impact of anthropomorphism, users' individual differences, and emotional responses on human-automation trust repair. Our experiment manipulated various types of trust violations and apology messages with di
View details →Negative / Null Result ReportOpen accessComputer Science
Anaïs Halin, Marc Van Droogenbroeck, Christel Devue · 2025 · arXiv
In this simulator study, we adopt a human-centered approach to explore whether and how drivers' cognitive state and driving environment complexity influence reliance on driving automation features. Besides, we examine whether such reliance affects driving performance. Participants operated a vehicle equipped with adaptive cruise control (ACC) in a simulator across six predefined driving scenarios varying in traffic conditions while either performing a cognitively demanding task (i.e., responding to mental calculations) or not. Throughout the experiment, participants had to respect speed limits
View details →Negative / Null Result ReportOpen accessComputer Science
Andreas Vogelsang, Alexander Korn, Giovanna Broccia et al. · 2025 · arXiv
Large language models (LLMs) are increasingly used to generate software artifacts, such as source code, tests, and trace links. Requirements play a central role in shaping the input prompts that guide LLMs, as they are often used as part of the prompts to synthesize the artifacts. However, the impact of requirements formulation on LLM performance remains unclear. In this paper, we investigate the role of requirements smells-indicators of potential issues like ambiguity and inconsistency-when used in prompts for LLMs. We conducted experiments using two LLMs focusing on automated trace link gene
View details →Negative / Null Result ReportOpen accessComputer Science
Salman Khawar, Yingdan Lu, Yilang Peng et al. · 2026 · arXiv
The rapid proliferation of visual content raises fundamental questions about how different visual formats and features shape perceived credibility. Drawing on processing fluency theory, this research examines how visuals shape credibility judgments. We focus on three popular formats-photos, infographics, and data visualizations-comparing them to text-only posts, and test how two visual features, aesthetic appeal and production quality, influence credibility through processing fluency as a mediating mechanism. Through a preregistered experiment with 1200 US participants, we found that visual po
View details →Negative / Null Result ReportOpen accessComputer Science
Ihsan Kahveci, Timothy A. Thomas, Nathalie E. Williams et al. · 2026 · arXiv
Home eviction poses a significant threat to housing stability, a critical determinant of health. This study examines the relationship between eviction and health and substance use within the unhoused population of King County, Washington. Using a sample of 1,106 individuals experiencing homelessness, we employed a quasi-experimental design to compare the health outcomes of those who have experienced eviction with those who have not. Our findings reveal eviction is associated with an 8.3% point increase (SE = 0.039) in the likelihood of reporting poor general health and an 9.5% increase (SE = 0
View details →Negative / Null Result ReportOpen accessComputer Science
Maximilian Warsinke, Maurizio Vergari, Tanja Kojić et al. · 2025 · arXiv
This study explores how prior exposure to physical objects influences the quality and realism perception of Digital Twins (DT) with varying levels of fidelity in Virtual Reality (VR). In a mixed experimental design, 24 participants were divided into two equal groups: an exposure group, in which members were shown physical objects before inspecting and rating their replicas in VR, and a control group without prior knowledge. Three objects were presented, each under four fidelity conditions with varying texture resolution and geometric detail. Participants rated perceived quality and realism thr
View details →Negative / Null Result ReportOpen accessComputer Science
Rahul Arora, Jiannan Li, Gongyi Shi et al. · 2021 · arXiv
Size and distance perception in Virtual Reality (VR) have been widely studied, albeit in a controlled laboratory setting with a small number of participants. We describe a fully remote perceptual study with a gamified protocol to encourage participant engagement, which allowed us to quickly collect high-quality data from a large, diverse participant pool (N=60). Our study aims to understand medium-field size and egocentric distance perception in real-world usage of consumer VR devices. We utilized two perceptual matching tasks -- distance bisection and size matching -- at the same target dista
View details →Negative / Null Result ReportOpen accessComputer Science
Samuel Westby, Richard J. Radke, Christoph Riedl et al. · 2023 · arXiv
Voice assistants are increasingly prevalent, from personal devices to team environments. This study explores how voice type and contribution quality influence human-agent team performance and perceptions of anthropomorphism, animacy, intelligence, and trustworthiness. By manipulating both, we reveal mechanisms of perception and clarify ambiguity in previous work. Our results show that the human resemblance of a voice assistant's voice negatively interacts with the helpfulness of an agent's contribution to flip its effect on perceived anthropomorphism and perceived animacy. This means human tea
View details →Negative / Null Result ReportOpen accessComputer Science
F. W. Bopp, Yu. M. Shabelski · 2006 · arXiv
The process of baryon number transfer due to string junction propagation in rapidity space is considered. It leads to a significant effect in the net baryon production in pA collisions at mid-rapidities and an even more significant effect in the forward hemisphere for the cases of πA interactions. The results of numerical calculations in the framework of the Quark-Gluon String Model are in reasonable agreement with the data. Special consideration is given to Λproduced in {π^-}-A collisions extracted from data of WA89 Collaboration.
View details →Negative / Null Result ReportOpen accessComputer Science
Meryem Yilmaz Soylu, Jeonghyun Lee, Jui-Tse Hung et al. · 2025 · arXiv
As Artificial Intelligence (AI) tools become increasingly embedded in higher education, understanding how students interact with these systems is essential to supporting effective learning. This study examines how students' AI literacy and prior exposure to AI technologies shape their perceptions of Socratic Mind, an interactive AI-powered formative assessment tool. Drawing on Self-Determination Theory and user experience research, we analyze relationships among AI literacy, perceived usability, satisfaction, engagement, and perceived learning effectiveness. Data from 309 undergraduates in Com
View details →Negative / Null Result ReportOpen accessComputer Science
S. Mouslih, M. Jakha, S. El Asri et al. · 2021 · arXiv
Choosing a specific direction for the propagation of laser field waves often presents a challenge for researchers studying laser-assisted ultrafast quantum processes. They are faced with the question of why exactly this direction and not another. This paper resolves the discussion in this issue regarding decay processes. Therefore, we study theoretically the pion decay process in the presence of a circularly polarized laser field propagating along an arbitrary general direction. Using the first Born approximation and the Dirac-Volkov states for charged particles, we derive an analytic expressi
View details →Negative / Null Result ReportOpen accessComputer Science
Bhaskar Gurram · 2026 · arXiv
Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, but this assumption has rarely been validated against human annotation. We introduce AgentProp-Bench, a 2,000-task benchmark with 2,300 traces across four domains, nine production LLMs, and a 100-label human-validated subset. We quantify judge reliability, characterize error propagation, and evaluate a runtime mitigation. Substring-based judging agrees with human annotation at kappa=0.049 (chance-level); a three-LLM ensemble reaches kappa=0.432 (moderate) with a conservative bias. Under valid
View details →Negative / Null Result ReportOpen accessComputer Science
Pappu Kumar Yadav, J. Alex Thomasson, Robert G. Hardin et al. · 2022 · arXiv
Plastic shopping bags that get carried away from the side of roads and tangled on cotton plants can end up at cotton gins if not removed before the harvest. Such bags may not only cause problem in the ginning process but might also get embodied in cotton fibers reducing its quality and marketable value. Therefore, it is required to detect, locate, and remove the bags before cotton is harvested. Manually detecting and locating these bags in cotton fields is labor intensive, time-consuming and a costly process. To solve these challenges, we present application of four variants of YOLOv5 (YOLOv5s
View details →Negative / Null Result ReportOpen accessComputer Science
Jun-Shuai Wang, Chang Liu, Ju Chen et al. · 2025 · arXiv
In this work, we systematically investigate the capability of space-based gravitational wave detectors in constraining parameters of non-tensor polarization modes. Using Bayesian inference and Fisher Information Matrix methods, we analyze gravitational wave signals from the inspiral phase of supermassive binary black hole mergers. By starting with time-domain signals and applying Fourier transforms, we avoid the use of the stationary phase approximation. We found an asymmetry in the estimation of the vector-mode parameter $α_x$ at inclination angles $ι= 0$ and $ι= π$, which has not been explic
View details →Negative / Null Result ReportOpen accessComputer Science
Jingchao Fang, Victoria Xiaohan Wen, Mina Lee · 2026 · arXiv
The growing capability of artificial intelligence (AI) leads to its increasing adoption in writing, spurring discussions around whether writers should disclose their AI use in writing. What influences the perceived necessity of disclosure? We look into this question from three dimensions: perspective (reader or writer of the text), purpose (the goal of reading or writing), and procedural factors (how AI was used in the writing process in terms of replaceability, effortfulness, intentionality, and directness). In a vignette study (N = 727), we find that readers consider disclosure to be more ne
View details →Negative / Null Result ReportOpen accessComputer Science
Wenqing Zhang, Trang Nguyen, Elizabeth A. Stuart et al. · 2025 · arXiv
Systematic reviews are crucial for synthesizing scientific evidence but remain labor-intensive, especially when extracting detailed methodological information. Large language models (LLMs) offer potential for automating methodological assessments, promising to transform evidence synthesis. Here, using causal mediation analysis as a representative methodological domain, we benchmarked state-of-the-art LLMs against expert human reviewers across 180 full-text scientific articles. Model performance closely correlated with human judgments (accuracy correlation 0.71; F1 correlation 0.97), achieving
View details →Negative / Null Result ReportOpen accessComputer Science
Dustin Eisenhardt, Timothy Schaumlöffel, Alperen Kantarci et al. · 2026 · arXiv
Deep learning models for computer vision often suffer from poor generalization when deployed in real-world settings, especially when trained on synthetic data due to the well-known Sim2Real gap. Despite the growing popularity of style transfer as a data augmentation strategy for domain generalization, the literature contains unresolved contradictions regarding three key design axes: the diversity of the style pool, the role of texture complexity, and the choice of style source. We present a systematic empirical study that isolates and evaluates each of these factors for driving scene understan
View details →Negative / Null Result ReportOpen accessComputer Science
Antonino Flachi, Masato Minamitsuji · 2009 · arXiv
We discuss the localization of scalar, fermion, and gauge field zero modes on a $3-$brane that resides at the intersection of two $4-$branes in six-dimensional anti-de Sitter space. This set-up has been introduced in the context of brane world models and, higher-dimensional versions of it, in string theory. In both six- and ten-dimensional cases, it has been shown that four-dimensional gravity can be reproduced at the intersection, due to the existence of a massless, localized graviton zero-mode. However, realistic scenarios require also the Standard Model to be localized on the $3-$brane. In
View details →Negative / Null Result ReportOpen accessComputer Science
Dominique Machuletz, Rainer Böhme · 2019 · arXiv
The European Union's General Data Protection Regulation (GDPR) requires websites to ask for consent to the use of cookies for \emph{specific purposes}. This enlarges the relevant design space for consent dialogs. Websites could try to maximize click-through rates and positive consent decision, even at the risk of users agreeing to more purposes than intended. We evaluate a practice observed on popular websites by conducting an experiment with one control and two treatment groups ($N=150$ university students in two countries). We hypothesize that users' consent decision is influenced by (1) the
View details →Negative / Null Result ReportOpen accessComputer Science
Mohammed Alsobay, David M. Rothschild, Jake M. Hofman et al. · 2025 · arXiv
Group decision-making often suffers from uneven information sharing, hindering decision quality. While large language models (LLMs) have been widely studied as aids for individuals, their potential to support groups of users, potentially as facilitators, is relatively underexplored. We present a pre-registered randomized experiment with 1,475 participants assigned to 281 five-person groups completing a hidden profile task--selecting an optimal city for a hypothetical sporting event--under one of four facilitation conditions: no facilitation, a one-time message prompting information sharing, a
View details →Negative / Null Result ReportOpen accessComputer Science
Xinyi Liu, Weiguang Wang, Hangfeng He · 2025 · arXiv
With the growing adoption of Large Language Models (LLMs) for open-ended tasks, accurately assessing epistemic uncertainty, which reflects a model's lack of knowledge, has become crucial to ensuring reliable outcomes. However, quantifying epistemic uncertainty in such tasks is challenging due to the presence of aleatoric uncertainty, which arises from multiple valid answers. While bias can introduce noise into epistemic uncertainty estimation, it may also reduce noise from aleatoric uncertainty. To investigate this trade-off, we conduct experiments on Visual Question Answering (VQA) tasks and
View details →Negative / Null Result ReportOpen accessComputer Science
Lennart Meincke, Ethan Mollick, Lilach Mollick et al. · 2025 · arXiv
This is the third in a series of short reports that seek to help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. In this report, we investigate two commonly held prompting beliefs: a) offering to tip the AI model and b) threatening the AI model. Tipping was a commonly shared tactic for improving AI performance and threats have been endorsed by Google Founder Sergey Brin (All-In, May 2025, 8:20) who observed that 'models tend to do better if you threaten them,' a claim we subject to empirical testing here. We evaluate model p
View details →Negative / Null Result ReportOpen accessComputer Science
Danai Korre · 2023 · arXiv
Embodied conversational agents (ECAs) are paradigms of conversational user interfaces in the form of embodied characters. While ECAs offer various manipulable features, this paper focuses on a study conducted to explore two distinct levels of presentation realism. The two agent versions are photorealistic and animated. The study aims to provide insights and design suggestions for speech-enabled ECAs within serious game environments. A within-subjects, two-by-two factorial design was employed for this research with a cohort of 36 participants balanced for gender. The results showed that both th
View details →Negative / Null Result ReportOpen accessComputer Science
Márton Tápai, Zoltán Keresztes, László Árpád Gergely · 2016 · arXiv
We derive the conservative secular evolution of precessing compact binaries to second post-Newtonian order accuracy, with leading-order spin-orbit, spin-spin and mass quadrupole-monopole contributions included. The emerging closed system of first-order differential equations evolves the pairs of polar and azimuthal angles of the spin and orbital angular momentum vectors together with the periastron angle. In contrast with the instantaneous dynamics, the secular dynamics is autonomous. This secular dynamics reliably characterizes the system over timescales starting from a few times the radial p
View details →