Skip to content
View Battam1111's full-sized avatar
馃幆
Focusing
馃幆
Focusing

Organizations

@polyunlp

Block or report Battam1111

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don鈥檛 include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user鈥檚 behavior. Learn more about reporting abuse.

Report abuse
battam1111/README.md

Yanjun Chen

PhD candidate at The Hong Kong Polytechnic University (joint training with EIT, Ningbo), working on reinforcement learning for large language models and embodied agents. Models are trainable; the environments that train them are not. I want to make the environment trainable, the way models are.

HomepageEmailGoogle ScholarORCID 路 Hong Kong

What I build

  • C3: exact per-decision credit for cooperative LLM agents by transcript replay, with a method-agnostic audit of credit quality. [paper]
  • AccuracyParadox-RLHF: reproducible RLHF training and evaluation pipelines and reference reward models. [paper, EMNLP 2024]

Pinned Loading

  1. Myco Myco Public

    The living armor an AI agent inhabits: eternal devouring, eternal evolution, eternal amplification.

    Rust 64 7

  2. EIT-EAST-Lab/C3 EIT-EAST-Lab/C3 Public

    Official implementation of the paper "Contextual Counterfactual Credit Assignment for Multi-Agent Reinforcement Learning in LLM Collaboration". (by Yanjun Chen)

    Python 36

  3. AccuracyParadox-RLHF AccuracyParadox-RLHF Public

    [EMNLP 2024 Main] Official implementation of the paper "The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models".

    Python 8