- Earth
-
05:48
(UTC +09:00) - https://damonxue.envs.net/
- https://github.com/0xdx2
Starred repositories
📚 Freely available programming books
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
《动手学深度学习》:面向中文读者、能运行、可讨论。中英文版被70多个国家的500多所大学用于教学。
Apache Superset is a Data Visualization and Data Exploration Platform
🌟 The Multi-Agent Framework: First AI Software Company, Towards Natural Language Programming
🧠 Train a 64M-parameter LLM from scratch in just 2h!
Qlib is an AI-oriented Quant investment platform that aims to use AI tech to empower Quant Research, from exploring ideas to implementing productions. Qlib supports diverse ML modeling paradigms, i…
Kronos: A Foundation Model for the Language of Financial Markets
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
⚡️ IPTV直播源自动更新工具:自动采集、校验、测速并生成可播放结果,支持 M3U/TXT/API 输出、自定义频道、IPv4/IPv6、Docker、GitHub Actions、CLI 与 GUI 多端部署
Educational framework exploring ergonomic, lightweight multi-agent orchestration. Managed by OpenAI Solution team.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Convert documentation websites, GitHub repositories, and PDFs into Claude AI skills with automatic conflict detection
A Conversational Speech Generation Model
The official GitHub page for the survey paper "A Survey of Large Language Models".
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
🌍 Discover our global repository of countries, states, and cities! 🏙️ Get comprehensive data in JSON, SQL, PSQL, SQLSERVER, MONGODB, SQLITE, XML, YAML, and CSV formats. Access ISO2, ISO3 codes, cou…
👀 Train a 65M-parameter VLM from scratch in just 2h!
Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
[CVPR 2025] EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation
Tencent Hunyuan3D-1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation
MusePose: a Pose-Driven Image-to-Video Framework for Virtual Human Generation
Evaluate your speech-to-text system with similarity measures such as word error rate (WER)
Our vision is to provide communication capabilities for intelligent agents, allowing them to connect with each other to form a collaborative network of intelligent agents.
A toolkit for reproducible evaluation, diagnostic, and error analysis of speaker diarization systems
Apache ShenYu Client SDK for python.