Skip to content
View rladmstn1714's full-sized avatar
😎
😎

Highlights

  • Pro

Block or report rladmstn1714

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Repository for "K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts"

Python 14 Updated Jun 2, 2026

LLM-wiki

TeX 1 Updated Jul 10, 2026

"I didn’t Make the Micro Decisions": Measuring, Inducing, and Exposing Goal-Level AI Contributions in Collaboration (COLM 2026)

Python 8 Updated Jun 16, 2026
Python 6 Updated Jun 5, 2025

Interactive Leaderboard with Benchub&HRET

Python 6 Updated May 13, 2026

Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues (ACL 2026)

Python 1 Updated Jun 28, 2026

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation

Python 11 Updated May 7, 2026

The most modern LLM evaluation toolkit

Python 70 12 Updated Apr 30, 2026

Dataset and code for paper: "Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese".

Python 16 6 Updated Nov 21, 2024

CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean

48 1 Updated Dec 23, 2024

Interview-based evaluation of LLMs

Python 31 1 Updated May 21, 2026