Hi, I'm Sakharam Thorat 👨💻
AI Voice Systems Engineer | Real-time Speech AI | Conversational Agents
🚀 I specialize in building production-grade real-time voice AI systems — from telephony pipelines to speech-to-speech agents.
My work spans across:
- 🎙️ Speech AI (STT, TTS, Speech-to-Speech)
- ⚡ Low-latency real-time systems
- 📡 SIP / VoIP / streaming architectures
- 🧠 Voice agents & conversational AI
- 🧩 On-device AI & optimization
- 📍 Based in Pune, India
- 💼 10+ years in real-time communication systems (Vonage, Avaya)
- 🔊 Built end-to-end voice AI pipelines (STT → LLM → TTS)
- ⚡ Focused on low-latency conversational systems (<200ms–2s)
- 🧠 Exploring AI voice agents, speech models & edge AI
📧 Reach me at: srt.2011@outlook.com
🔗 LinkedIn
- Real-time speech-to-speech AI (Azure OpenAI)
- Multi-provider STT (Deepgram, Google, Azure, Whisper)
- TTS pipelines (Edge, streaming synthesis)
- Voice agents with multi-turn conversations
- SIP / RTP / VoIP pipelines
- WebRTC + WebSocket streaming
- Codec optimization (G.711, G.729, Opus)
- Packetization (ptime), jitter & packet loss handling
- Noise suppression (RNN-based training on A100)
- CTC-based ASR experimentation
- NVIDIA NeMo, Vosk (Kaldi-based)
- Audio quality evaluation (PESQ, MOS)
- Whisper.cpp, Parakeet (on-device STT)
- Model quantization for low-memory inference
- CPU-optimized real-time pipelines
- aura-sde-interview-agent
Real-time voice AI agent for Google SDE interviews
→ Full loop: Behavioural, Coding, System Design, Debugging
→ Built with Gemini Live + Vertex AI
→ Live grading + voice-only interaction
-
nemotron-stt (In Progress)
High-concurrency WebSocket ASR server using Nemotron-0.6B -
whisper-stt
Production-grade streaming STT using faster-whisper turbo -
qwen3-stt
Voice AI agent design + STT experimentation
-
VeloxTx (In Progress)
Multilingual low-latency translation engine
→ “Train anywhere, run anywhere” philosophy -
velox-realtime (In Progress)
CPU-optimized real-time ASR (<30ms latency)
-
x-benchmark-tests
Benchmarking STT models (accuracy, latency, throughput) -
stt-systems-lab (In Progress)
Experimental lab for:- Nemotron
- Whisper Turbo
- Qwen3
- ONNX / browser inference
-
freeswitch-speech-ai
Real-time:- Transcription
- Translation
- Sentiment
- NER
- Call summarization
-
pjsua2-python
Prebuilt Python bindings for PJSIP (SIP stack)
- slm_framework_full (In Progress)
Unified framework for:- STT, TTS, Vocoder
- Small Language Models
- Real-time inference
- STT: Whisper, Deepgram, Azure, Google
- TTS: Edge, streaming pipelines
- Speech Models: NeMo, Vosk, Kaldi-based
- Metrics: WER, PESQ, MOS
- FastAPI, WebSockets, WebRTC
- LiveKit, Cloud Run, Vertex AI
- Kubernetes, Docker
- Python (primary)
- TypeScript / JavaScript
- Java
- Real-time AI voice agents with tool-calling
- Ultra-low latency speech-to-speech systems
- On-device conversational AI (edge-first)
- Production-grade multilingual voice pipelines
⭐ If you're working on Voice AI / Speech Systems / Real-time AI — let’s collaborate!