Stars
Repository for "K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts"
"I didn’t Make the Micro Decisions": Measuring, Inducing, and Exposing Goal-Level AI Contributions in Collaboration (COLM 2026)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues (ACL 2026)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
The most modern LLM evaluation toolkit
Dataset and code for paper: "Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese".
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean