Highlights
- Pro
Stars
Modular Cognitive Architecture Emerges in Large Language Models
Official codebase for "Next-Latent Prediction Transformers Learn Compact World Models"
A curated list of research and projects on world models
[ArXiv 2025] A survey about controllable video generation: This repo is the official awesome of "Controllable video generation: A survey"
Paper: "Zero-shot World Models Are Developmentally Efficient Learners"
Your personal voice interface for any app. Speak naturally and your words appear wherever your cursor is, with fully customizable AI voice dictation. Open source alternative to Wispr Flow.
A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.
Official PyTorch Implementation of KL-Tracing: Taming generative video models for zero-shot optical flow extraction.
Official PyTorch Implementation of Opt-CWM: Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals.
[3DV25] Official code for "Towards Foundation Models for 3D Vision: How Close Are We?"
OpenMMLab Self-Supervised Learning Toolbox and Benchmark
The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (V…
A framework for evaluating models on their alignment to brain and behavioral measurements (100+ benchmarks)
Job Application Assistant: bagged transformer-based huggingface embedding models with tf-idf to recommend web-scraped real-time jobs based on uploaded resume.
Official codebase for I-JEPA, the Image-based Joint-Embedding Predictive Architecture. First outlined in the CVPR paper, "Self-supervised learning from images with a joint-embedding predictive arch…
PyTorch code and models for V-JEPA self-supervised learning from video.
An approach to building pure vision foundation models by prompting masked predictors with "counterfactual" visual inputs.
Papers from the intersection of deep learning and neuroscience
A curated list of Large Language Model (LLM) Interpretability resources.
This repository contains code to quantitatively evaluate instruction-tuned models such as Alpaca and Flan-T5 on held-out tasks.