-
Special Technological Center Ltd.
- Saint Petersburg
- https://scholar.google.ru/citations?user=T-kDfn4AAAAJ&hl=ru
Stars
A curated collection of skills for AI coding agents. Skills are packaged instructions and scripts that extend agent capabilities across development, documentation, planning, and professional workfl…
Algorithm powering the For You feed on X
Tongyi Deep Research, the Leading Open-source Deep Research Agent
The context API to search, scrape, and interact with the web at scale. 🔥
MiMo-Audio: Audio Language Models are Few-Shot Learners
Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.
Flink Agents is an Agentic AI framework based on Apache Flink
CRYFISH: On deep audio analysis with Large Language Models
A simple library for Fréchet Audio Distance (FAD) calculation
An extremely fast Python package and project manager, written in Rust.
🤗 smolagents: a barebones library for agents that think in code.
Fully open-source command-line AI assistant inspired by OpenAI Codex, supporting local language models.
Lightweight coding agent that runs in your terminal
A benchmark for evaluating audio encoders on various audio tasks.
Template for creating audio encoders compatible with X-ARES
Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
Code for SuDoRm-Rf networks for efficient audio source separation. SuDoRm-Rf stands for SUccessive DOwnsampling and Resampling of Multi-Resolution Features which enables a more efficient way of sep…
GitHub Repository for the Survey Paper on Audio-Language Datasets for Scenes and Events
PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models
SALMONN family: A suite of advanced multi-modal LLMs
Fast Food Memes monolith https://t.me/ffmemesbot
Lightweight highly configurable Python launcher based on microkernel architecture
Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus.
The official repository of Dynamic-SUPERB.
The PyTorch-based audio source separation toolkit for researchers