Skip to content
View Durgesh92's full-sized avatar
🎯
Focusing
🎯
Focusing
  • Infolabs Global
  • Dubai

Block or report Durgesh92

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results
Python 91 14 Updated Jul 18, 2026

Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.

Rust 13,526 925 Updated Aug 8, 2026

SOTA-Class TTS at Compact Scale

Python 421 41 Updated Aug 9, 2026

An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python depend…

C++ 1,231 157 Updated Aug 8, 2026

[ACL 2026] Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models

Python 20 1 Updated May 23, 2026

Real-time, low-latency AEC at 48 kHz — designed for efficiency and low CPU usage.

Python 10 4 Updated Jul 28, 2026

AVTR-1: Avatars that listen back

Python 410 73 Updated May 26, 2026

Lightweight on-device keyword spotting engine for iOS using CoreML and real-time audio streaming.

Swift 16 2 Updated Jun 14, 2025

Elucidated Text-To-Audio (ETTA) is a SOTA text-to-audio model with a holistic understanding of the design space and trained with synthetic captions.

Python 135 12 Updated Mar 3, 2026

A LiveKit TTS plugin for self hosted open source models. Expose any model through a WebSocket compatible API (low latency) and swap engines without changing agent code.

Python 7 Updated Jun 20, 2026
Python 1 Updated Jun 11, 2026

"ViMax: Agentic Video Generation (Director, Screenwriter, Producer, and Video Generator All-in-One)"

Python 11,699 1,748 Updated Jul 29, 2026

The retrieval layer for production AI systems. Lightning-fast (<10ms) search without vector databases. Built for browser, edge, on-device, and cloud.

Python 663 91 Updated Jul 29, 2026

A standalone desktop/smartTV overlay that translates system audio into 3D Sign Language animation in real-time.

TypeScript 4 Updated Aug 7, 2026

Open-source American english TTS model. 6 voices and a high performance inference library for Apple Silicon.

Python 17 2 Updated May 20, 2026

[SIGGRAPH 2026] Pixal3D: Pixel-Aligned 3D Generation from Images

Python 2,089 205 Updated Jun 23, 2026

tiny-world-builder

JavaScript 1,508 202 Updated Aug 9, 2026

Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).

Python 13,411 1,302 Updated Jun 26, 2026

[ECCV 2026 Oral] Implementation of "Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length"

Python 2,329 271 Updated Jul 26, 2026

Self hosted, real-time digital human agent platform. Build voice-first AI agents with WebRTC, persona memory, tools, RAG, and optional digital-human video.

Python 1,551 215 Updated Aug 5, 2026

FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST.…

C++ 495 60 Updated Aug 8, 2026
Python 109 19 Updated Aug 7, 2026

Browser-based text-to-speech powered by OmniVoice. Runs entirely locally via WebGPU and WebAssembly.

JavaScript 16 3 Updated Jul 2, 2026

Open source video conferencing app powered by LiveKit. Built with Django and React.

Python 2,213 270 Updated Aug 8, 2026

Self-hosted DTLN noise suppression plugin for LiveKit Agents — no cloud API, no per-minute fees

Python 44 10 Updated Apr 16, 2026

Building actual open source including dataset Multilingual TTS more than 150 languages with Voice Cloning.

Jupyter Notebook 56 4 Updated Jul 14, 2026

A framework for efficient model inference with omni-modality models

Python 5,972 1,430 Updated Aug 9, 2026

🎙️ VoxSherpa TTS Offline Neural Text-to-Speech Engine for Android ⚡ Sherpa-ONNX powered 🔊 Natural voice synthesis 📱 Fully offline processing 🚀 No cloud • No limits

Java 192 29 Updated Aug 6, 2026

Vietnamese TTS with instant voice cloning • On-device • Real-time CPU inference • 24kHz audio quality • Chuyển văn bản thành giọng nói tiếng Việt • Text to speech tiếng Việt • TTS tiếng Việt

Python 2,306 688 Updated Aug 3, 2026
Next