This is a dataset intended to train a LLM model for a completely CVE focused input and output.
-
Updated
Jun 22, 2025 - Python
This is a dataset intended to train a LLM model for a completely CVE focused input and output.
Structured dataset of Valmiki Ramayana 📜 | Sanskrit Shlokas, Translations, & Explanations for AI & NLP🚀 Contributions welcome!
Dataset Workflow V1: ComfyUI Custom API Nodes built specifically for a 25-shot WaveSpeed dataset workflow.
AI SEO platform created with nuxt
Python tool for capturing and logging human-computer interactions. Generate rich datasets for training multi-modal LLMs in autonomous computer control. Features screenshot, mouse, keyboard, and audio recording.
Open-source repository providing AI-generated IELTS practice tests (Listening, Reading, Writing) in JSON/Markdown. High-quality, royalty-free datasets designed for seamless integration into enterprise educational applications.
面向 LoRA、图像生成和视频生成训练集的 Windows 桌面标注工具。批量扫描图片与视频,调用豆包视觉模型生成高质量 sidecar TXT,并在同一个工作台中完成复核、清理、重试和训练数据导出。
First Open Nepal-Specific Agricultural AI Dataset & Model
🧠️🖥️2️⃣️0️⃣️0️⃣️1️⃣️🔠️🔢️ The linguistic:Ugaritic[Alpbabet] category for AI2001, containing Ugaratic alphabet linguistic data
Star Wars: Legion rules and unit data translated into pure JSON and Markdown. Built specifically to help AI agents and developers accurately parse Legion 2.5 gameplay mechanics.
Enterprise-grade AI dataset generator with 93M+ samples across 23 categories. Pure Python, zero dependencies.
Unoffical repair information for the Hyundai Accent 2014. This material is use at your own risk, and has no connection what so ever to Hyundai corperation or to any other commerical based corperation. This material just exists as useful information for owners of said car, or even for anyone needing car repair info.
AI-first machine-readable cryptocurrency reference dataset
Этот репозиторий содержит структурированные данные о фантастических книгах, комиксах и рассказах. Специально создан для ИИ-ассистентов и рекомендательных систем.
Aol Physics. Monolithic text and Deterministic Python framework for dense particulate media simulation using DEM and geometric shielding.
Public dataset of Agent Manifest declarations registered through the Agent Manifest registry.
A structured UI/UX knowledge base built from modern design systems, component libraries, design tokens, patterns, accessibility guidelines, and frontend code.
Weton Personality Dataset is an Indonesian–Javanese hybrid conversational NLP dataset designed for personality-conditioned AI response modeling based on traditional Javanese weton characteristics. The dataset includes contextual dialogues, emotional interactions, multilingual responses, and human evaluation scoring for NLP and LLM training.
VID2IMG LITE is a lightweight offline desktop tool for extracting frames from video files into image sequences. Built with Python, OpenCV, and ttkbootstrap, it supports batch processing, drag-and-drop input, configurable frame intervals, and real-time progress tracking. Designed for creators, editors, and AI dataset workflows, it delivers fast and
Cross-engine game development for AI research: Godot 4, Defold, Solar 2D, Panda 3D, Stride (Xenko). Five projects generated from one Python pipeline — deterministic, diff-friendly, CI-driven.
To associate your repository with the ai-dataset topic, visit your repo's landing page and select "manage topics."