Stars
An incredibly fast proxy checker & IP rotator with ease.
[ICML 2026] The official implementation of paper "Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key to Unification"
Unofficial PyTorch implementation of EOSTok: end-to-end autoregressive image generation with a 1D semantic tokenizer (arXiv:2605.00503). Runs on MacBook/MPS via an MNIST demo config; ImageNet Table…
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Official implementation of AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling.
SoulX-FlashHead: A unified 1.3B-parameter framework designed for high-fidelity, infinite-length, and real-time streaming portrait video generation.
We propose LeapTalk, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos.
Spectrum-based acceleration for ComfyUI’s native MiniMax H3 audio-video model. Forecasts post-transformer features with Chebyshev ridge regression to skip selected transformer evaluations, with ada…
Mixture-of-experts (MoE) training megakernel for NVL72s
SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE
https://wavespeed.ai/ Context parallel attention that accelerates DiT model inference with dynamic caching
[Tech Report] Context Scaling: Scaling Properties of Text Conditioning in Visual Generation
Reverse Engineering / Authorized Penetration Testing / Security Research Skill Router Pack AI-powered routing + On-demand toolchain bootstrapping + Self-evolving knowledge base Supports Claude Code…
PyTorch Code for Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
[ICLR 2025] Official PyTorch Implementation of Gated Delta Networks: Improving Mamba2 with Delta Rule
High-performance single-GPU inference for selected model checkpoints and GPUs.
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Everything you need to know about LLM inference
OpenAI's Codex Security CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities. npm: https://www.npmjs.com/package/@openai/codex-security
You like pytorch? You like micrograd? You love tinygrad! ❤️
a language for fast, portable data-parallel computation