-
01:57
(UTC +03:00) - https://huggingface.co/ysdede
- in/ysdede
Lists (10)
Sort Name ascending (A-Z)
- All languages
- Arduino
- Assembly
- AutoHotkey
- Batchfile
- Bikeshed
- C
- C#
- C++
- CSS
- Cuda
- Cython
- Dart
- Dockerfile
- Eagle
- GLSL
- Go
- Groovy
- HTML
- Java
- JavaScript
- Jinja
- Julia
- Jupyter Notebook
- Kotlin
- Lua
- MDX
- Makefile
- Meson
- Objective-C
- PHP
- Pascal
- Perl
- PowerShell
- Python
- Rich Text Format
- Ruby
- Rust
- Shell
- Solidity
- Svelte
- Swift
- TSQL
- Tcl
- TeX
- TypeScript
- V
- Visual Basic .NET
- Vue
- Zig
- kvlang
- reStructuredText
Starred repositories
A FastAPI wrapper for NVIDIA's new parakeet 0.6b v3 TTS 600m model designed for high-quality multilingual speech recognition, beating Whisper Large v3 and Whisper Large v3 turbo with blazing fast s…
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
A GUI for masking/rotoscoping video using AI models
Speech recognition with word-level timestamps, optimized for batch inference.
A comprehensive collection of Agent Skills for context engineering, multi-agent architectures, and production agent systems. Use when building, optimizing, or debugging agent systems that require e…
A two-level morphological analyzer for Turkish.
Lean neural real-time acoustic echo cancellation with soft delay estimation - GGML and PyTorch inference
A platform plugin that enables AI agents (Claude, ChatGPT, etc.) to drive, observe, execute, and verify the Unreal Engine Editor and Runtime through semantic commands via HTTP, WebSocket, and MCP t…
Industrial audio online policy distillation (OPD) training stack for ASR and TTS, distilling compact audio models from stronger teacher models.
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
Describe Anything, Anywhere, at Any Moment (DAAAM), a novel approach to real-time, large-scale, spatio-temporal memory
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
Lore is a next-generation, open source version control system
Code Release for "OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data"
Extract a target speaker’s clean, non-overlapped speech from multi-speaker audio and export word-safe LJSpeech-style TTS datasets.
Generative AI extensions for onnxruntime
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
iPad-controlled ST 2110 / NMOS broadcast-over-IP multiviewer & monitoring rig
Turkish general and medical ASR, benchmarking, and readability post-processing
CCEdit: Creative and Controllable Video Editing via Diffusion Models
high-performance inference and serving library for interactive autoregressive video and world models
Voice-to-text dictation app with local (Nvidia Parakeet/Whisper) and cloud models (BYOK). Privacy-first and available cross-platform.
Zonos2 is a leading open-weight text-to-speech MoE.
Whisfusion: Parallel ASR Decoding via a Diffusion Transformer
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling
Local-first healthcare AI: clinical NER & HIPAA PII de-identification that runs 100% on-device. 2,200+ medical models, 21 languages, Apple MLX + Python, no cloud, no patient data leaving your netwo…