Skip to content

About

MED-Dataset Platform:专业医疗AI数据集生成与管理平台(Next.js + Electron + Prisma)。AI原生设计,零门槛上手,多LLM与多格式导出,适合科研与企业落地。

Topics

Resources

Contributing

Stars

4 stars

Watchers

0 watching

Forks

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

MED-Dataset Platform

GitHub Repo stars GitHub forks GitHub Downloads GitHub Release AGPL 3.0 License GitHub last commit Build Status

🏥 MED-Dataset Platform

🚀 专业医疗数据集生成与管理平台 | Professional Medical AI Dataset Generator

⭐ 如果这个项目对你有帮助,请给我们一个Star!⭐

🔥 AI驱动 | 🎯 开箱即用 | 💡 多LLM支持 | 🌍 国际化

🇨🇳 简体中文 | 🇺🇸 English

✨ 特性 • 🚀 快速开始 • 📖 使用指南 • 🤝 贡献 • 📄 许可证

💝 喜欢这个项目?给个Star⭐支持一下吧!
Become a Sponsor · Join Discussions

🌟 项目亮点

🎯 为什么选择 MED-Dataset Platform?

✅ AI原生设计 - 专为医疗AI训练数据生成而打造 ✅ 零门槛使用 - 无需编程背景,医疗专业人员也能轻松上手 ✅ 企业级质量 - 支持大规模数据集生成,满足商业项目需求 ✅ 开源免费 - 完全开源,可商业使用

🔥 热门功能一览:

  • 📄 智能文档解析(PDF/Word/Markdown)
  • 🤖 AI自动问答生成
  • 🎯 多种数据集格式导出
  • 🌐 多LLM模型支持
  • 💻 跨平台桌面应用

💡 核心优势

MED-Dataset Platform 是专门为医疗领域AI数据集生成而设计的专业平台。通过直观的界面,让医疗研究人员和AI开发者能够快速创建、处理和管理高质量的训练数据集,助力医疗AI研究和临床决策支持系统的发展。

🚀 一键将医疗文档转换为结构化AI训练数据集

✨ 核心功能

🔥 主要特性

功能模块 描述 亮点
📄 智能文档处理 支持PDF、Word、Markdown等格式 🎯 专业医疗内容识别
✂️ 智能文本分割 针对医疗内容优化的分割算法 🧠 保持语义完整性
🤖 AI问题生成 从医疗文本自动生成高质量问题 ⚡ 批量生成,效率倍增
📊 数据集构建 生成结构化问答数据集 🎯 ML训练就绪
🔗 多LLM支持 OpenAI、Ollama、智谱AI等 🌐 灵活选择模型
💾 数据管理 安全的数据存储和项目管理 🔒 企业级安全
📤 多格式导出 Alpaca、ShareGPT、JSON等 🎨 适配各种训练框架
🖥️ 跨平台支持 Web、桌面、Docker 📱 随时随地使用

🎨 用户体验

  • 🎯 零学习成本 - 直观的可视化界面
  • ⚡ 高效工作流 - 从文档到数据集一站式完成
  • 🌍 国际化支持 - 中英文双语界面
  • 🔧 高度可定制 - 灵活的配置选项

🎬 演示视频

ed3.mp4

💡 提示: 完整功能演示请查看 详细演示视频


🚀 快速开始

📦 下载客户端

🎯 推荐方式:一键下载桌面应用

Windows MacOS Linux

Setup.exe

Intel

M

AppImage

💻 源码安装

🛠️ 开发者本地运行指南

  1. 克隆项目:
git clone https://github.com/2023Anita/med-dataset-platform.git
cd med-dataset-platform
  1. 安装依赖:
npm install
# 或者使用 pnpm (推荐)
pnpm install
  1. 启动应用:
npm run build && npm run start
# 访问 http://localhost:1717

💡 提示: 首次运行会自动初始化数据库,请稍等片刻

🐳 Docker 部署

🚀 最简单的部署方式

  1. 克隆并构建:
git clone https://github.com/2023Anita/med-dataset-platform.git
cd med-dataset-platform
docker build -t med-dataset-platform .
  1. 运行容器:
docker run -d \
  -p 1717:1717 \
  -v ./local-db:/app/local-db \
  --name med-dataset-platform \
  med-dataset-platform
  1. 访问应用: http://localhost:1717

📝 注意: 数据将保存在 ./local-db 目录中


📖 使用指南

🆕 创建项目

  1. 点击首页"创建项目"按钮
  2. 输入项目名称和描述
  3. 配置您偏好的LLM API设置

📄 处理文档

  1. 在"文本分割"页面上传文件(支持PDF、Markdown、txt、DOCX)
  2. 查看并调整自动分割的文本片段
  3. 查看并调整全局领域树

🤖 生成问题

  1. 基于文本块批量构造问题
  2. 查看并编辑生成的问题
  3. 使用标签树组织问题

📊 创建数据集

  1. 基于问题批量构造数据集
  2. 使用配置的LLM生成答案
  3. 查看、编辑和优化生成的答案

📤 导出数据集

  1. 在数据集页面点击"导出"按钮
  2. 选择首选格式(Alpaca或ShareGPT)
  3. 选择文件格式(JSON或JSONL)
  4. 根据需要添加自定义系统提示
  5. 导出您的数据集

Project Structure

med-dataset/
├── app/                                # Next.js application directory
│   ├── api/                            # API routes
│   │   ├── llm/                        # LLM API integration
│   │   │   ├── ollama/                 # Ollama API integration
│   │   │   └── openai/                 # OpenAI API integration
│   │   ├── projects/                   # Project management API
│   │   │   ├── [projectId]/            # Project-specific operations
│   │   │   │   ├── chunks/             # Text chunk operations
│   │   │   │   ├── datasets/           # Dataset generation and management
│   │   │   │   ├── generate-questions/ # Batch question generation
│   │   │   │   ├── questions/          # Question management
│   │   │   │   └── split/              # Text splitting operations
│   │   │   └── user/                   # User-specific project operations
│   ├── projects/                       # Frontend project pages
│   │   └── [projectId]/                # Project-specific pages
│   │       ├── datasets/               # Dataset management UI
│   │       ├── questions/              # Question management UI
│   │       ├── settings/               # Project settings UI
│   │       └── text-split/             # Text processing UI
│   └── page.js                         # Homepage
├── components/                         # React components
│   ├── datasets/                       # Dataset-related components
│   ├── home/                           # Homepage components
│   ├── projects/                       # Project management components
│   ├── questions/                      # Question management components
│   └── text-split/                     # Text processing components
├── lib/                                # Core libraries and tools
│   ├── db/                             # Database operations
│   ├── i18n/                           # Internationalization
│   ├── llm/                            # LLM integration
│   │   ├── common/                     # Common LLM tools
│   │   ├── core/                       # Core LLM clients
│   │   └── prompts/                    # Prompt templates
│   │       ├── answer.js               # Answer generation prompts (Chinese)
│   │       ├── answerEn.js             # Answer generation prompts (English)
│   │       ├── question.js             # Question generation prompts (Chinese)
│   │       ├── questionEn.js           # Question generation prompts (English)
│   │       └── ... other prompts
│   └── text-splitter/                  # Text splitting tools
├── locales/                            # Internationalization resources
│   ├── en/                             # English translations
│   └── zh-CN/                          # Chinese translations
├── public/                             # Static resources
│   └── imgs/                           # Image resources
└── local-db/                           # Local file database
    └── projects/                       # Project data storage

Documentation

Community Practice

MED-Dataset × LLaMA Factory: Enabling LLMs to Efficiently Learn Domain Knowledge

Contributing

We welcome contributions from the community! If you'd like to contribute to MED-Dataset, please follow these steps:

  1. Fork the repository
  2. Create a new branch (git checkout -b feature/amazing-feature)
  3. Make your changes
  4. Commit your changes (git commit -m 'Add some amazing feature')
  5. Push to the branch (git push origin feature/amazing-feature)
  6. Open a Pull Request (submit to the DEV branch)

Please ensure that tests are appropriately updated and adhere to the existing coding style.

Join Discussion Group & Contact the Author

Contact information available in the project repository.

License

This project is licensed under the AGPL 3.0 License - see the LICENSE file for details.

Star History

Star History Chart

Built with ❤️ by ConardLi • Follow me: WeChat Official Account|Bilibili|Juejin|Zhihu|Youtube

About

MED-Dataset Platform:专业医疗AI数据集生成与管理平台(Next.js + Electron + Prisma)。AI原生设计,零门槛上手,多LLM与多格式导出,适合科研与企业落地。

Topics

Resources

Contributing

Stars

4 stars

Watchers

0 watching

Forks

Releases

Sponsor this project

Packages

Contributors

Languages