YEDDA: A Lightweight Collaborative Text Span Annotation Tool. Code for ACL 2018 Best Demo Paper Nomination.
-
Updated
Feb 19, 2023 - Python
YEDDA: A Lightweight Collaborative Text Span Annotation Tool. Code for ACL 2018 Best Demo Paper Nomination.
Data release for the ImageInWords (IIW) paper.
Video Generation Benchmark
[NeurIPS'25] MLLM-CompBench evaluates the comparative reasoning of MLLMs with 40K image pairs and questions across 8 dimensions of relative comparison: visual attribute, existence, state, emotion, temporality, spatiality, quantity, and quality. CompBench covers diverse visual domains, including animals, fashion, sports, and scenes
Code for the AAAI 2023 Paper "Real or Fake Text?: Investigating Human Ability to Detect Boundaries Between Human-Written and Machine-Generated Text"
Color-matching makeup shades using k-means clustering and Claude multimodal LLM, evaluated with Delta E / CIELAB. Includes human annotation pipeline and Streamlit app.
A collection of MTurk templates designed to make complex tasks easier for human annotators.
Code and data for inherent biases in grammatical error correction and text simplification, including 2,500 human-written correction and simplification tasks.
Repository of the paper "Supporting Online Toxicity Detection with Knowledge Graphs" (ICWSM 22).
Repository for the journal article 'SHAMSUL: Systematic Holistic Analysis to investigate Medical Significance Utilizing Local interpretability methods in deep learning for chest radiography pathology prediction'
Design-aware temporal reliability auditing for human annotation studies
Human-annotated NLP analysis of 262 Phasmophobia Steam reviews with blinded labeling, duplicate-aware validation, TF-IDF models, and an interactive dashboard.
Code repository for the paper "Estimating Ground Truth in a Low-labelled Data Regime: A Study of Racism Detection in Spanish" (NEATClasS 2022)
A dataset-auditing pipeline that flags and reviews label-quality issues across three sentiment classification datasets: the human-annotated SST-2 benchmark, and two LLM-generated synthetic datasets
Human-annotated benchmark for the completeness and relevance of long answers (Language Resources and Evaluation, 2026)
To associate your repository with the human-annotation topic, visit your repo's landing page and select "manage topics."