Stars
🤗 ml-intern: an open-source ML engineer that reads papers, trains models, and ships ML models
A general purpose scientific writer
VIP cheatsheet for Stanford's CME 295 Transformers and Large Language Models
E-ink calendar integrating google calendar og OWM onto a 7.5 inch Waveshare screen based on an ESP32 LOLIN32 board
ESP32 e-ink calendar display integrated with Home Assistant
Synthetic Dialog Generation and Analysis with LLMs
Causal streaming adaptation of OpenAI Whisper for real-time transcription on small audio chunks.
SlamKit is an open source tool kit for efficient training of SpeechLMs. It was used for "Slamming: Training a Speech Language Model on One GPU in a Day"
Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.
Have a natural, spoken conversation with AI!
Official code for "MAmmoTH2: Scaling Instructions from the Web" [NeurIPS 2024]
Foundational Model for Speech Recognition Tasks
A course on aligning smol models.
Add n-gram and large language model (LLM) support to Whisper models.
Experiments to turn LLMs into high quality simultaneous translation engines
TFDS data loaders for sign language datasets.
a text-to-gloss-to-pose-to-video pipeline for spoken to signed language translation
CHIME-7/8 diarization champion system: neural speaker diarization using memory-aware multi-speaker embedding with sequence-to-sequence architecture
Code for our INTERSPEECH paper Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
Awesome speech/audio LLMs, representation learning, and codec models
Whisper realtime streaming for long speech-to-text transcription and translation
A python true casing utility that restores case information for texts