Lists (2)
Sort Name ascending (A-Z)
Stars
Lean implementation of various multi-agent LLM methods, including Iteration of Thought (IoT)
An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.
[ICLR'26] "Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment"
An Open Large Reasoning Model for Real-World Solutions
ImOV3D: Learning Open Vocabulary Point Clouds 3D Object Detection from Only 2D Images (NeurIPS2024)
Differentiable Ray Tracing Toolbox for Radio Propagation Simulations
Official GPU implementation of the paper "PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance"
PyTorch code for "ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning"
This repository aims to collect Transformer-based sound event detection (SED) algorithms.
Official repository for Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning [ICLR 2025]
A collection of projects designed to help developers quickly get started with building deployable applications using the Claude API
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
[EMNLP 2025 Findings] Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
g1: Using Llama-3.1 70b on Groq to create o1-like reasoning chains
A NodeJS RAG framework to easily work with LLMs and embeddings
[NeurIPS 2024] Official Repository of The Mamba in the Llama: Distilling and Accelerating Hybrid Models
The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure
Find API quality and security issues via your OpenAPI spec
OpenResearcher, an advanced Scientific Research Assistant
[ICLR 2025] LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
ClickAttention: Click Region Similarity Guided Interactive Segmentation
real time face swap and one-click video deepfake with only a single image
[ECCV2024] ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
TypeScript AI platform with AI chat, Autonomous agents, Software developer agents, chatbots and more
Convert REPL interactions into example-based tests.