Skip to content
View pxyWaterMoon's full-sized avatar

Highlights

  • Pro

Block or report pxyWaterMoon

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results
Python 120 Updated Aug 13, 2026

🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.

Python 3,929 283 Updated Aug 13, 2026

主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题

HTML 14,903 1,461 Updated Jun 14, 2026

专业的 LaTeX 简历模板,专为大模型与 Agent 算法工程师设计 | Professional LaTeX resume template for LLM & Agent algorithm engineers

TeX 390 211 Updated Aug 12, 2026

大模型算法岗面试题(含答案):常见问题和概念解析 "大模型面试题"、"算法岗面试"、"面试常见问题"、"大模型算法面试"、"大模型应用基础"

Jupyter Notebook 1,990 134 Updated Jul 20, 2026

Sky-T1: Train your own O1 preview model within $450

Python 3,398 344 Updated Jul 12, 2025

An unofficial LaTeX thesis template for ShanghaiTech University.

TeX 125 27 Updated Jan 24, 2024

This repository contains code for the paper Direct Preference Optimization with an Offset (ODPO).

Python 21 2 Updated Feb 17, 2025

Experiment task scheduling made easy.

Python 33 5 Updated Aug 10, 2026
Jupyter Notebook 8 Updated Jan 14, 2025

Code for numerical experiments in the QPDE paper.

Jupyter Notebook 3 3 Updated Jun 4, 2023

Deep Reinforcement Learning of Partial Differential Equations

Jupyter Notebook 2 Updated Mar 19, 2025

Resources to crack the next coding interview

2,031 688 Updated Aug 3, 2024

ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors [EMNLP 2024 Findings]

Python 231 10 Updated Sep 29, 2024

【ACL 2024】 SALAD benchmark & MD-Judge

Python 176 15 Updated Mar 8, 2025

[NeurIPS 2024] SACPO (Stepwise Alignment for Constrained Policy Optimization)

Python 9 1 Updated Dec 23, 2024

论文写作与资料分享

3,286 649 Updated Aug 7, 2022

PyTorch version of Stable Baselines, reliable implementations of reinforcement learning algorithms.

Python 13,684 2,170 Updated Jul 25, 2026

Daily updated LLM papers. 每日更新 LLM 相关的论文,欢迎订阅 👏 喜欢的话动动你的小手 🌟 一个

Python 1,314 59 Updated Aug 13, 2026

【deepin源移植】Debian/Ubuntu上的QQ/微信快速安装方式

Python 5,287 376 Updated Jan 7, 2025

JMLR: OmniSafe is an infrastructural framework for accelerating SafeRL research.

Python 1,148 161 Updated Mar 17, 2025

An unofficial LaTeX beamer template for ShanghaiTech students.

TeX 8 1 Updated May 9, 2024

A library for constrained RLHF.

Python 13 Updated Feb 19, 2024

Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback

Python 1,613 133 Updated Nov 24, 2025

The Missing Semester of Your CS Education 📚

CSS 5,979 1,419 Updated Aug 9, 2026

The official PyTorch implementation of the paper "Human Motion Diffusion Model"

Python 4,086 457 Updated Oct 1, 2025

Pytorch实现:使用ResNet18网络训练Cifar10数据集,测试集准确率达到95.46%(从0开始,不使用预训练模型)

Python 324 29 Updated Mar 3, 2026

An open-source PyTorch code for crowd counting

Jupyter Notebook 733 198 Updated Mar 30, 2024

A very simple and easy to understand RISC-V core.

C 1,507 243 Updated Nov 9, 2023

https://hrl.boyuai.com/

Jupyter Notebook 4,932 827 Updated Nov 22, 2022
Next