-
Tsinghua University
- Haidian, Beijing
Highlights
- Pro
Stars
[ACL 2025 main] FR-Spec: Frequency-Ranked Speculative Sampling
LLM papers I'm reading, mostly on inference and model compression
Fourier Controller Networks (FCNet) for Real-Time Decision-Making in Embodied Learning, ICML 2024
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
TorchMultimodal is a PyTorch library for training state-of-the-art multimodal multi-task models at scale.
Official implementation of our LREC-COLING 2024 paper "Generative Multimodal Entity Linking".
A playbook for systematically maximizing the performance of deep learning models.
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), ga…
Support extracting BUTD features for NLVR2 images.
Research code for ECCV 2020 paper "UNITER: UNiversal Image-TExt Representation Learning"
Official implementation for "Multimodal Chain-of-Thought Reasoning in Language Models" (stay tuned and more will be updated)
Simple image captioning model
Power CLI and Workflow manager for LLMs (core package)
CapDec: SOTA Zero Shot Image Captioning Using CLIP and GPT2, EMNLP 2022 (findings)
🍀 Pytorch implementation of various Attention Mechanisms, MLP, Re-parameter, Convolution, which is helpful to further understand papers.⭐⭐⭐
Some scripts and configuration files for personal use.
A collection of resources on multimodal knowledge graph, including datasets, papers and contests.