Stars
An open-source RAG-based tool for chatting with your documents.
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
[EMNLP'25] s3 - ⚡ Efficient & Effective Search Agent Training via RL for RAG (RLVR for Search with Minimal Data)
Toolkit to segment text into sentences or other semantic units in a robust, efficient and adaptable way.
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
This repository contains the codes of "A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild", published at ACM Multimedia 2020. For HD commercial model, please try out Sync Labs
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
Face super resolution based on ESRGAN
This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
Official implementation of ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining (AAAI 2024)
Official implementation of UPOCR: Towards unified pixel-level OCR interface (ICML 2024)
Inpaint anything using Segment Anything and inpainting models.
The official project of paper "Visual Text Processing: A Comprehensive Review and Unified Evaluation""
A curated list of image inpainting and video inpainting papers and resources
Code for the paper "UVDoc: Neural Grid-based Document Unwarping"
[IEEE TPAMI] Hi-SAM: Marrying Segment Anything Model for Hierarchical Text Segmentation
Scene text removal via cascaded text stroke detection and erasing
LPIPS metric. pip install lpips
[Open-Source Project] Combining MMOCR with Segment Anything & Stable Diffusion. Automatically detect, recognize and segment text instances, with serval downstream tasks, e.g., Text Removal and Text…
This dataset contains re-annotations of 4 popular Latin/English scene text recognition datasets.
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.