Lists (1)
Sort Name ascending (A-Z)
Stars
"AI-Trader: 100% Fully-Automated Agent-Native Trading"
OpenMMLab Detection Toolbox and Benchmark
PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT
[CVPR 2024] Official RT-DETR (RTDETR paddle pytorch), Real-Time DEtection TRansformer, DETRs Beat YOLOs on Real-time Object Detection. 🔥 🔥 🔥
《Hello 算法》:动画图解、一键运行的数据结构与算法教程。支持简中、繁中、English、日本語,提供 Python, Java, C++, C, C#, JS, Go, Swift, Rust, Ruby, Kotlin, TS, Dart 等代码实现
YOLOv10: Real-Time End-to-End Object Detection [NeurIPS 2024]
tulip-berkeley / open_clip
Forked from mlfoundations/open_clipAn open source implementation of CLIP (With TULIP Support)
Ongoing research training transformer models at scale
Research Code for Multimodal-Cognition Team in Ant Group
EVA Series: Visual Representation Fantasies from BAAI
Retrieval and Retrieval-augmented LLMs
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
Janus-Series: Unified Multimodal Understanding and Generation Models
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
DataComp: In search of the next generation of multimodal datasets
GLM-4 series: Open Multilingual Multimodal Chat LMs | 开源多语言多模态对话模型
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
An open source implementation of CLIP.
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.
OCET, torch, transformers, DeepLOB,limit-order-books
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
一个基于HuggingFace开发的大语言模型训练、测试工具。支持各模型的webui、终端预测,低参数量及全参数模型训练(预训练、SFT、RM、PPO、DPO)和融合、量化。