-
Tsinghua University
- Beijing
- http://conghui.github.com/
Stars
Codex skill for converting slide images, PDFs, and image-based PPTX files into editable PowerPoint decks.
GPT-Image-2 PPT Generator Skill for Creating Image-Based PowerPoint Presentations in Codex and Other Skill-Compatible Agents
SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
Teams-first Multi-agent orchestration for Claude Code
conghui / replaycode
Forked from sanbuphy/learn-coding-agentReplayCode — first open-source rebuild of Claude Code that actually runs. Built from decompiled source with Node.js/esbuild
[ECCV 2026] A diffusion-based framework for document OCR that replaces autoregressive decoding with block-level parallel diffusion decoding.
📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
MinerU-HTML: An SLM-powered HTML main content extractor that outputs clean HTML bodies. Perfect for Deep Research Agents, RAG applications, and training data generation.
Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
Data browser based on s3. 一个基于 S3 的数据(json / jsonl / parquet / html / md等)可视化工具。👇 Try online.
A simple screen parsing tool towards pure vision based GUI agent
Implementation of Reinforcement Learning Algorithms. Python, OpenAI Gym, Tensorflow. Exercises and Solutions to accompany Sutton's Book and David Silver's course.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
[ICLR 2025 Spotlight] The official implementation of the paper “LOKI:A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models”
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
This is the repo for the paper Multi-Agent Collaborative Data Selection for Efficient LLM Pretraining.
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
A Comprehensive Toolkit for High-Quality PDF Content Extraction
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
This list of writing prompts covers a range of topics and tasks, including brainstorming research ideas, improving language and style, conducting literature reviews, and developing research plans.
[ACL 2024 Main Conference] Chinese commonsense benchmark for LLMs
小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B 站视频 | 评论爬虫、微博帖子 | 评论爬虫、百度贴吧帖子 | 百度贴吧评论回复爬虫 | 知乎问答文章|评论爬虫
Awesome-LLM: a curated list of Large Language Model
The official GitHub page for the survey paper "A Survey of Large Language Models".
AAAI 2024: Visual Instruction Generation and Correction