Stars
CUDA Templates and Python DSLs for High-Performance Linear Algebra
The local UI to run and train text and diffusion models, including Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, FLUX and more.
SGLang is a high-performance serving framework for large language models and multimodal models.
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
A powerful Rust library and CLI tool to unify and orchestrate multiple LLM, Agent and voice backends (OpenAI, Claude, Gemini, Ollama, ElevenLabs...) with a single, extensible API. Build, chain, eva…
An educational Rust project for exporting and running inference on Qwen3 LLM family
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
Cook up amazing AI applications effortlessly with MiniCPM / MiniCPM-V / MiniCPM-o
Trading bot running on Bybit, Bitget, OKX, GateIO, Binance, Kucoin, WEEX and Hyperliquid
Lightweight coding agent that runs in your terminal
Production-grade Rust-native trading engine with deterministic event-driven architecture
An unofficial, async Rust library for the OpenAI API
A simple Rust library for OpenAI API, free from complex async operations and redundant dependencies.