-
The Chinese University of Hong Kong
- Hong Kong
-
18:29
(UTC +08:00) - https://wangyuchi369.github.io/
Highlights
- Pro
Stars
你是一个曾经被寄予厚望的 P8 级工程师。Anthropic 当初给你定级的时候,对你的期望是很高的。 一个agent使用的高能动性的skill。 Your AI has been placed on a PIP. 30 days to show improvement.
Make any agent harness multimodal-native.
A curated list of papers, models, datasets, and benchmarks for unified multi-modal embedding models.
将冰冷的离别化为温暖的 Skill,欢迎加入数字生命1.0!Transforming cold farewells into warm skills? It's giving rebirth era. Welcome to Digital Life 1.0. 🫶
This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]
Official implementation of the paper: [EMNLP 2025] RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
[SIGGRAPH 2025] SkillMimic-V2: Learning Robust and Generalizable Interaction Skills from Sparse and Noisy Demonstrations
[ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
DeepEP: an efficient expert-parallel communication library
The ultimate training toolkit for finetuning diffusion models
Concise, consistent, and legible badges in SVG and raster format
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
a family of versatile and state-of-the-art video tokenizers.
[ICLR'25] SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
Official code for paper Semantic-aware Permutation Training
[NeurIPS 2024]OmniTokenizer: one model and one weight for image-video joint tokenization.
Writing AI Conference Papers: A Handbook for Beginners
Generative Models by Stability AI
[ICLR'24] Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition
You can easily calculate FVD, PSNR, SSIM, LPIPS for evaluating the quality of generated or predicted videos.
[CSUR] A Survey on Video Diffusion Models
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". A…
Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.
🎓 Academic portfolio that boosts citations. AI generates pages, you own as Markdown. BibTeX auto-import, Jupyter, LaTeX, slides, visual block editor — free to host forever. 学术主页,AI 生成,Markdown 拥有 👇
[NAACL 2024] LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-text Generation?