⭐ 如果这个项目对你有帮助,请给我们一个Star!⭐
🔥 AI驱动 | 🎯 开箱即用 | 💡 多LLM支持 | 🌍 国际化
✨ 特性 • 🚀 快速开始 • 📖 使用指南 • 🤝 贡献 • 📄 许可证
💝 喜欢这个项目?给个Star⭐支持一下吧!
Become a Sponsor · Join Discussions
🎯 为什么选择 MED-Dataset Platform?
✅ AI原生设计 - 专为医疗AI训练数据生成而打造 ✅ 零门槛使用 - 无需编程背景,医疗专业人员也能轻松上手 ✅ 企业级质量 - 支持大规模数据集生成,满足商业项目需求 ✅ 开源免费 - 完全开源,可商业使用
🔥 热门功能一览:
- 📄 智能文档解析(PDF/Word/Markdown)
- 🤖 AI自动问答生成
- 🎯 多种数据集格式导出
- 🌐 多LLM模型支持
- 💻 跨平台桌面应用
MED-Dataset Platform 是专门为医疗领域AI数据集生成而设计的专业平台。通过直观的界面,让医疗研究人员和AI开发者能够快速创建、处理和管理高质量的训练数据集,助力医疗AI研究和临床决策支持系统的发展。
🚀 一键将医疗文档转换为结构化AI训练数据集
| 功能模块 | 描述 | 亮点 |
|---|---|---|
| 📄 智能文档处理 | 支持PDF、Word、Markdown等格式 | 🎯 专业医疗内容识别 |
| ✂️ 智能文本分割 | 针对医疗内容优化的分割算法 | 🧠 保持语义完整性 |
| 🤖 AI问题生成 | 从医疗文本自动生成高质量问题 | ⚡ 批量生成,效率倍增 |
| 📊 数据集构建 | 生成结构化问答数据集 | 🎯 ML训练就绪 |
| 🔗 多LLM支持 | OpenAI、Ollama、智谱AI等 | 🌐 灵活选择模型 |
| 💾 数据管理 | 安全的数据存储和项目管理 | 🔒 企业级安全 |
| 📤 多格式导出 | Alpaca、ShareGPT、JSON等 | 🎨 适配各种训练框架 |
| 🖥️ 跨平台支持 | Web、桌面、Docker | 📱 随时随地使用 |
- 🎯 零学习成本 - 直观的可视化界面
- ⚡ 高效工作流 - 从文档到数据集一站式完成
- 🌍 国际化支持 - 中英文双语界面
- 🔧 高度可定制 - 灵活的配置选项
ed3.mp4
💡 提示: 完整功能演示请查看 详细演示视频
🎯 推荐方式:一键下载桌面应用
| Windows | MacOS | Linux | |
|
Setup.exe |
Intel |
M |
AppImage |
🛠️ 开发者本地运行指南
- 克隆项目:
git clone https://github.com/2023Anita/med-dataset-platform.git
cd med-dataset-platform- 安装依赖:
npm install
# 或者使用 pnpm (推荐)
pnpm install- 启动应用:
npm run build && npm run start
# 访问 http://localhost:1717💡 提示: 首次运行会自动初始化数据库,请稍等片刻
🚀 最简单的部署方式
- 克隆并构建:
git clone https://github.com/2023Anita/med-dataset-platform.git
cd med-dataset-platform
docker build -t med-dataset-platform .- 运行容器:
docker run -d \
-p 1717:1717 \
-v ./local-db:/app/local-db \
--name med-dataset-platform \
med-dataset-platform- 访问应用:
http://localhost:1717
📝 注意: 数据将保存在
./local-db目录中
- 点击首页"创建项目"按钮
- 输入项目名称和描述
- 配置您偏好的LLM API设置
- 在"文本分割"页面上传文件(支持PDF、Markdown、txt、DOCX)
- 查看并调整自动分割的文本片段
- 查看并调整全局领域树
- 基于文本块批量构造问题
- 查看并编辑生成的问题
- 使用标签树组织问题
- 基于问题批量构造数据集
- 使用配置的LLM生成答案
- 查看、编辑和优化生成的答案
- 在数据集页面点击"导出"按钮
- 选择首选格式(Alpaca或ShareGPT)
- 选择文件格式(JSON或JSONL)
- 根据需要添加自定义系统提示
- 导出您的数据集
med-dataset/
├── app/ # Next.js application directory
│ ├── api/ # API routes
│ │ ├── llm/ # LLM API integration
│ │ │ ├── ollama/ # Ollama API integration
│ │ │ └── openai/ # OpenAI API integration
│ │ ├── projects/ # Project management API
│ │ │ ├── [projectId]/ # Project-specific operations
│ │ │ │ ├── chunks/ # Text chunk operations
│ │ │ │ ├── datasets/ # Dataset generation and management
│ │ │ │ ├── generate-questions/ # Batch question generation
│ │ │ │ ├── questions/ # Question management
│ │ │ │ └── split/ # Text splitting operations
│ │ │ └── user/ # User-specific project operations
│ ├── projects/ # Frontend project pages
│ │ └── [projectId]/ # Project-specific pages
│ │ ├── datasets/ # Dataset management UI
│ │ ├── questions/ # Question management UI
│ │ ├── settings/ # Project settings UI
│ │ └── text-split/ # Text processing UI
│ └── page.js # Homepage
├── components/ # React components
│ ├── datasets/ # Dataset-related components
│ ├── home/ # Homepage components
│ ├── projects/ # Project management components
│ ├── questions/ # Question management components
│ └── text-split/ # Text processing components
├── lib/ # Core libraries and tools
│ ├── db/ # Database operations
│ ├── i18n/ # Internationalization
│ ├── llm/ # LLM integration
│ │ ├── common/ # Common LLM tools
│ │ ├── core/ # Core LLM clients
│ │ └── prompts/ # Prompt templates
│ │ ├── answer.js # Answer generation prompts (Chinese)
│ │ ├── answerEn.js # Answer generation prompts (English)
│ │ ├── question.js # Question generation prompts (Chinese)
│ │ ├── questionEn.js # Question generation prompts (English)
│ │ └── ... other prompts
│ └── text-splitter/ # Text splitting tools
├── locales/ # Internationalization resources
│ ├── en/ # English translations
│ └── zh-CN/ # Chinese translations
├── public/ # Static resources
│ └── imgs/ # Image resources
└── local-db/ # Local file database
└── projects/ # Project data storage
- View the demo video of this project: MED-Dataset Demo Video
MED-Dataset × LLaMA Factory: Enabling LLMs to Efficiently Learn Domain Knowledge
We welcome contributions from the community! If you'd like to contribute to MED-Dataset, please follow these steps:
- Fork the repository
- Create a new branch (
git checkout -b feature/amazing-feature) - Make your changes
- Commit your changes (
git commit -m 'Add some amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request (submit to the DEV branch)
Please ensure that tests are appropriately updated and adhere to the existing coding style.
Contact information available in the project repository.
This project is licensed under the AGPL 3.0 License - see the LICENSE file for details.