Stars
Official Code: SPG-CDENet – Spatial Prior-Guided Cross Dual Encoder Network for Multi-Organ Segmentation
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation
[CVPR 2026] G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
🤖 RoboOS: A Universal Embodied Operating System for Cross-Embodied and Multi-Robot Collaboration
📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥
Interactive visualizations of the geometric intuition behind diffusion models.
Evaluation of tree biomass using LiDAR, down to the branch scale
[ICLR 2026 Oral] Intrinsic Entropy of Context Length Scaling in LLMs
Large Concept Models: Language modeling in a sentence representation space
[ECCV 2024] Scene as Gaussians for Vision-Based 3D Semantic Occupancy Prediction
Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities。
[NeurIPS 2024 Spotlight] Official repository of the CycleNet paper: "CycleNet: Enhancing Time Series Forecasting through Modeling Periodic Patterns". This work is developed by the Lab of Professor …
A curated list of awesome LLM/VLM/VLA/World Model for Autonomous Driving(LLM4AD) resources (continually updated)
🌐 Jekyll is a blog-aware static site generator in Ruby
[ICLR 2025 Spotlight] Official implementation of "Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts"
Motion Question Answering via Modular Motion Programs
Official code repository of the paper Linear Transformers Are Secretly Fast Weight Programmers.
Open-source desktop app for local LLMs. Text, vision, tool-calling, OpenAI/Anthropic-compatible API. 100% private.
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
The official implementation of Segment Any 3D GAussians (AAAI-25)
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
③[ICML2024] [IQA, IAA, VQA] All-in-one Foundation Model for visual scoring. Can efficiently fine-tune to downstream datasets.