Lists (3)
Sort Name ascending (A-Z)
Starred repositories
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题
专业的 LaTeX 简历模板,专为大模型与 Agent 算法工程师设计 | Professional LaTeX resume template for LLM & Agent algorithm engineers
大模型算法岗面试题(含答案):常见问题和概念解析 "大模型面试题"、"算法岗面试"、"面试常见问题"、"大模型算法面试"、"大模型应用基础"
Sky-T1: Train your own O1 preview model within $450
lirundong / shtthesis
Forked from mohuangrui/ucasthesisAn unofficial LaTeX thesis template for ShanghaiTech University.
This repository contains code for the paper Direct Preference Optimization with an Offset (ODPO).
Code for numerical experiments in the QPDE paper.
Deep Reinforcement Learning of Partial Differential Equations
ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors [EMNLP 2024 Findings]
[NeurIPS 2024] SACPO (Stepwise Alignment for Constrained Policy Optimization)
PyTorch version of Stable Baselines, reliable implementations of reinforcement learning algorithms.
Daily updated LLM papers. 每日更新 LLM 相关的论文,欢迎订阅 👏 喜欢的话动动你的小手 🌟 一个
JMLR: OmniSafe is an infrastructural framework for accelerating SafeRL research.
An unofficial LaTeX beamer template for ShanghaiTech students.
tedmoskovitz / ConstrainedRL4LMs
Forked from allenai/RL4LMsA library for constrained RLHF.
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
The Missing Semester of Your CS Education 📚
The official PyTorch implementation of the paper "Human Motion Diffusion Model"
Pytorch实现:使用ResNet18网络训练Cifar10数据集,测试集准确率达到95.46%(从0开始,不使用预训练模型)
An open-source PyTorch code for crowd counting
A very simple and easy to understand RISC-V core.