Oracle and Human Baselines for Native Language Identification
Shervin Malmasi; Joel Tetreault; Mark Dras · 2015
WASTE classifies this as Negative / Null Result Report · AI classification, approximate
The study found no significant effect — useful as a negative control or null benchmark for your own design.
Abstract
We examine different ensemble methods, including an oracle, to estimate the upper-limit of classification accuracy for Native Language Identification (NLI). The oracle outperforms state-of-the-art systems by over 10% and results indicate that for many misclassified texts the correct class label receives a significant portion of the ensemble votes, often being the runner-up. We also present a pilot study of human performance for NLI, the first such experiment. While some participants achieve modest results on our simplified setup with 5 L1s, they did not outperform our NLI system, and this perf
Abstract by Shervin Malmasi; Joel Tetreault; Mark Dras (2015) — licensed CC BY 4.0.
About to run something similar?
Run an AI Precheck on your own design to catch failure modes like this one before you spend the time. Your first desk check is free.
Related failures
Leakage and the reproducibility crisis in machine-learning-based science
Negative / Null Result ReportDefining and detecting quantum speedup
Negative / Null Result ReportService robots in hotels: understanding the service quality perceptions of human-robot interaction
Negative / Null Result ReportBoosting methods for multi-class imbalanced data classification: an experimental review
Negative / Null Result ReportFINANCIAL DEVELOPMENT AND ECONOMIC GROWTH: A META‐ANALYSIS
Negative / Null Result ReportThe impact of site-specific digital histology signatures on deep learning model accuracy and bias
WASTE indexes this work — it does not host or republish it. Failure-type classification is automated and approximate.
Metadata source: OpenAlex · DOI 10.3115/v1/w15-0620
