Skip to content
View darkrush's full-sized avatar

Highlights

  • Pro

Block or report darkrush

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
5 stars written in Python
Clear filter

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

Python 80,550 6,724 Updated Sep 24, 2026

A Comprehensive Toolkit for High-Quality PDF Content Extraction

Python 10,023 752 Updated Jan 3, 2025

MinerU-HTML: An SLM-powered HTML main content extractor that outputs clean HTML bodies. Perfect for Deep Research Agents, RAG applications, and training data generation.

Python 292 28 Updated Mar 27, 2026

[ACL 2025 Best Theme Paper] This is the official implementation for the paper: "Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models"

Python 197 16 Updated Aug 29, 2025

This is the repo for the paper Multi-Agent Collaborative Data Selection for Efficient LLM Pretraining.

Python 49 3 Updated Aug 22, 2025