Skip to content
View jinnaiyuu's full-sized avatar

Highlights

  • Pro

Block or report jinnaiyuu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

slime is an LLM post-training framework for RL Scaling.

Python 7,857 1,132 Updated Aug 11, 2026

J-tau: A Japanese tau-bench for Benchmarking Tool-Agent-User Interaction in Real-World Domains

Python 6 1 Updated Jul 10, 2026
Python 15 3 Updated Aug 6, 2026

A fast and soft pattern search for trillion-scale corpora.

Python 239 11 Updated Feb 28, 2026
Python 5 Updated Feb 3, 2026

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems (Submitted on 26 Sep 2025)

Python 5 3 Updated Jan 15, 2026

[EMNLP2025] Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling

Python 16 1 Updated Nov 20, 2025

Muon is an optimizer for hidden layers in neural networks

Python 2,777 128 Updated May 24, 2026

Helpful kernel tutorials, examples and SKILLs for tile-based GPU programming

Python 796 82 Updated Aug 11, 2026

We introduce COMMUNITYNOTES dataset for predicting note helpfulness and its reasons, propose an automatic reason-optimization framework, and show its benefits for evidence sufficiency and fact-chec…

Python 4 1 Updated May 2, 2026
Python 77 7 Updated Apr 20, 2026
7 2 Updated May 30, 2026

A simple job queue using 'tmux'

Shell 4 Updated Jan 23, 2025
Python 3 Updated Nov 10, 2024
SAS 2 2 Updated Dec 15, 2024

Dataset for evaluating the knowledge of Yokai in language models.

Python 3 Updated Jun 25, 2026

Code of "Evaluation of Best-of-N Sampling Strategies for Language Model Alignment"

Python 6 Updated Feb 19, 2025

Code of "Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment" (2025).

Python 14 1 Updated Apr 4, 2025

Evaluate your LLM's response with Prometheus and GPT4 đŸ’¯

Python 1,105 68 Updated Apr 25, 2025

AirLLM 70B inference with single 4GB GPU

Jupyter Notebook 30,768 3,272 Updated Aug 11, 2026

Code of Annotation-Efficient Language Model Alignment via Diverse and Representative Response Texts (EMNLP Findings 2025)

Python 10 1 Updated May 29, 2024

A curated list of research papers and resources on Cultural LLM.

55 3 Updated Sep 26, 2024

Code of "Model-Based Minimum Bayes Risk Decoding for Text Generation" 2024

Jupyter Notebook 8 1 Updated Aug 15, 2024

Flexible evaluation tool for language models

Python 61 4 Updated Aug 7, 2026

A library for minimum Bayes risk (MBR) decoding

Python 53 7 Updated Nov 2, 2025
Python 10 Updated Aug 7, 2026

Robust recipes to align language models with human and AI preferences

Python 5,658 489 Updated May 26, 2026

[EMNLP 2024] Introducing Filtered Direct Preference Optimization (fDPO) that enhances language model alignment with human preferences by discarding lower-quality samples compared to those generated…

Jupyter Notebook 16 1 Updated Nov 27, 2024

The Prism Alignment Project

Jupyter Notebook 93 3 Updated Apr 25, 2024
Next