Skip to content
View Alex-xixiang's full-sized avatar
💭
I may be slow to respond.
💭
I may be slow to respond.

Block or report Alex-xixiang

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A Framework for LLM-based Multi-Agent Reinforced Training and Inference

Python 538 48 Updated Apr 14, 2026

✨✨Latest Advances on Multimodal Large Language Models

17,956 1,128 Updated Jul 2, 2026

RAGEN leverages reinforcement learning to train LLM reasoning agents in interactive, stochastic environments.

Python 2,755 229 Updated Apr 14, 2026

Samples for CUDA Developers which demonstrates features in CUDA Toolkit

C++ 9,417 2,385 Updated May 27, 2026

CUDA/Metal accelerated language model inference

C 646 33 Updated May 29, 2025

📚LeetCUDA: Modern CUDA Learn Notes with PyTorch for Beginners🐑, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.🎉

Cuda 11,616 1,215 Updated Jul 24, 2026

LLM training in simple, raw C/CUDA

Cuda 30,628 3,708 Updated Jun 26, 2025

基于通义千问 Qwen2.5-Omni 的实时语音对话系统,使用在线API服务,支持实时语音交互、动态语音活动检测和流式音频处理。A real-time voice conversation system based on Qwen2.5-Omni Online-API, supporting real-time voice interaction, dynamic voice activi…

Python 91 15 Updated May 11, 2025
Python 1,062 309 Updated Jan 29, 2023

✅(已完结)超级全面的 深度学习 笔记【土堆 Pytorch】【李沐 动手学深度学习】【吴恩达 深度学习】【大飞 大模型Agent】

Jupyter Notebook 22,878 2,582 Updated Jun 30, 2026

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,644 4,272 Updated Jul 24, 2026

✨终生持续更新✨ 计算机基础自学笔记/心得/实验/资源汇总;课程:数据结构、操作系统(MIT6.S081)、分布式系统(MIT6.824)等

Jupyter Notebook 582 73 Updated Feb 1, 2026

A Rust machine learning framework.

Rust 4,710 329 Updated May 30, 2026

Using ChatGPT to connect with FreeSWITCH, creating an intelligent phone robot.

HTML 130 37 Updated Apr 30, 2024

official implementation of ICLR'2025 paper: Rethinking Bradley-Terry Models in Preference-based Reward Modeling: Foundations, Theory, and Alternatives

Python 73 4 Updated Apr 2, 2025

简单实现VAD+声纹锁+SenseVoice完成类语音实时转录的小项目

Python 42 4 Updated Sep 23, 2024

PyTorch version of Stable Baselines, reliable implementations of reinforcement learning algorithms.

Python 13,603 2,162 Updated Jul 24, 2026

A library with extensible implementations of DPO, KTO, PPO, ORPO, and other human-aware loss functions (HALOs).

Python 908 51 Updated Sep 30, 2025

Implementations of selected inverse reinforcement learning algorithms.

Python 1,083 235 Updated Oct 21, 2022

This repository collects papers for "A Survey on Knowledge Distillation of Large Language Models". We break down KD into Knowledge Elicitation and Distillation Algorithms, and explore the Skill & V…

1,295 72 Updated Mar 9, 2025

每个人都能看懂的大模型知识分享,LLMs春/秋招大模型面试前必看,让你和面试官侃侃而谈

Jupyter Notebook 7,009 655 Updated May 31, 2026

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

Python 18,972 1,483 Updated Jul 23, 2026

使用vllm加速cosyvoice2的推理

Jupyter Notebook 498 66 Updated Apr 26, 2025

My course work solutions and quiz answers

Jupyter Notebook 61 29 Updated Dec 24, 2025

Train transformer language models with reinforcement learning.

Python 18,916 2,863 Updated Jul 24, 2026

RL Scaling and Test-Time Scaling (ICML'25)

116 1 Updated Jan 23, 2025

ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search (NeurIPS 2024)

Python 709 51 Updated Jan 20, 2025
Next