e-ISSN: Pending
Replication FailureOpen accessComputer Science· cited by 62

The reliability of acceptability judgments across languages

Tal Linzen; Yohei Oseki · 2018 · Glossa a journal of general linguistics

WASTE classifies this as Replication Failure · AI classification, approximate

A previously reported effect did not replicate here — verify it holds before you build on it.

Abstract

The reliability of acceptability judgments made by individual linguists has often been called into question. Recent large-scale replication studies conducted in response to this criticism have shown that the majority of published English acceptability judgments are robust. We make two observations about these replication studies. First, we raise the concern that English acceptability judgments may be more reliable than judgments in other languages. Second, we argue that it is unnecessary to replicate judgments that illustrate uncontroversial descriptive facts; rather, candidates for replicatio

Abstract by Tal Linzen; Yohei Oseki, Glossa a journal of general linguistics (2018) — licensed CC BY 4.0.

About to run something similar?

Run an AI Precheck on your own design to catch failure modes like this one before you spend the time. Your first desk check is free.

WASTE indexes this work — it does not host or republish it. Failure-type classification is automated and approximate.

Metadata source: OpenAlex · DOI 10.5334/gjgl.528