Stars
Python tool for converting files and office documents to Markdown.
14MB foundation model for tiny devices; phones, wearables, smart home, and robots.
Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
夫子•明察司法大模型是由山东大学、浪潮云、中国政法大学联合研发,以 ChatGLM 为大模型底座,基于海量中文无监督司法语料与有监督司法微调数据训练的中文司法大模型。该模型支持法条检索、案例分析、三段论推理判决以及司法对话等功能,旨在为用户提供全方位、高精准的法律咨询与解答服务。
[中文法律大模型] DISC-LawLLM: an intelligent legal system powered by large language models (LLMs) to provide a wide range of legal services.
Instruct-tune LLaMA on consumer hardware
Code for loralib, an implementation of "LoRA: Low-Rank Adaptation of Large Language Models"
Code & Data for our Paper "Alleviating Hallucinations of Large Language Models through Induced Hallucinations"
qiguanjie / ICD
Forked from HillZhang1999/ICDCode & Data for our Paper "Alleviating Hallucinations of Large Language Models through Induced Hallucinations"
Unsupervised text tokenizer for Neural Network-based text generation.
alibaba / Megatron-LLaMA
Forked from NVIDIA/Megatron-LMBest practice for training LLaMA models in Megatron-LM
A No-Recurrence Sequence-to-Sequence Model for Speech Recognition
Text Normalization & Inverse Text Normalization
基于PaddlePaddle实现端到端中文语音识别,从入门到实战,超简单的入门案例,超实用的企业项目。支持当前最流行的DeepSpeech2、Conformer、Squeezeformer模型
CDCPP: Cross-Domain Chinese Punctuation Prediction
Robust Speech Recognition via Large-Scale Weak Supervision
Data repository for pretrained NLP models and NLP corpora.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
Universal Romanizer that can convert any unicode script to roman (latin) script
Unsupervised phone and word segmentation using dynamic programming on self-supervised VQ features.