Stars
Auto-regressive 3D CT volume generation using Latent Video Flow Matching with a Spatial-Temporal DiT (STDiT), conditioned on CT report embeddings.
Diffusion model for synthetic 3D CT scan video generation — IDEA Lab, FAU Erlangen-Nürnberg
🏛️ 三省六部制 · OpenClaw Multi-Agent Orchestration System — 9 specialized AI agents with real-time dashboard, model config, and full audit trails
ChatDev 2.0: Dev All through LLM-powered Multi-Agent Collaboration
Diffusion Models in Medical Imaging (Published in Medical Image Analysis Journal)
Reference PyTorch implementation and models for DINOv3
Official codebase used to develop Vision Transformer, SigLIP, MLP-Mixer, LiT and more.
OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image genera…
Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object detection, Segmentation, Image similarity and more.
为GPT/GLM等LLM大语言模型提供实用化交互接口,特别优化论文阅读/润色/写作体验,模块化设计,支持自定义快捷按钮&函数插件,支持Python和C++等项目剖析&自译解功能,PDF/LaTex论文翻译&总结功能,支持并行问询多种LLM模型,支持chatglm3等本地模型。接入通义千问, deepseekcoder, 讯飞星火, 文心一言, llama2, rwkv, claude2, m…
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
A curated list of awesome computer vision resources