ContextGem: Effortless LLM extraction from documents
-
Updated
Aug 13, 2026 - Python
ContextGem: Effortless LLM extraction from documents
Use LLMs to robustly extract web data
Simplifies the retrieval, extraction, and training of structured data from various unstructured sources.
🕵️♂️ Privacy-focused AI job scraper, local storage, and interactive dashboard. Auto-scrapes AI/ML roles from top companies using ScrapeGraph-AI + LLM and LangGraph Agents, filters for relevance, and provides a Streamlit UI for tracking applications. Built for developers seeking AI careers.
HTML to Markdown with CSS selector and XPath annotations
Lightfeed SDK to search and filter web data
Extract structured data from documents with LLMs. Define schemas in Pydantic and get typed entities with exact source spans. Handles PDFs, HTML, long text, and OCR.
CORSA is a Python tool for scraping, cleaning, and analyzing AUA course data from SONIS and GenEd sources.
HAIR is a semantic Hardware Abstraction IR describing silicon devices with normalized peripherals, registers, fields, timing, and constraints. Every element carries provenance, enabling reliable generation of SVDs, PACs, HALs, simulators, and documentation.
AI-powered competitive intelligence tool that scrapes company websites, uses LLMs for dynamic page selection and extraction, and generates executive prospectus reports.
Benchmarked invoice and document extraction pipeline: PII redaction before any model call, layout-aware extraction into a strict schema, field-level confidence scoring, validation rules, a human review queue, and corrections fed back as exemplars. Runs offline, no API key.
Clean Web-to-Markdown & Lead Intelligence API for LLMs and AI Agents (0ms Cold Start, 0-Token Cost)
AI-Native Spec-Driven Development framework for Claude Code com integracao nativa ao GitHub
Pipeline automatizzata per la ricerca di acceleratori in Europa e l'analisi dei portfolio startup.
Generates FAQs for any website using Firecrawl
Competitive exam question extraction pipeline using llm free-teirs
Valoracion inmobiliaria por comparables en Madrid: circuito reproducible de texto libre a Excel y panel. Validado contra 22.643 anuncios reales que el motor no habia visto (error mediano 12,0 %), comparado siempre contra lineas base mas simples.
AI-agent-driven venue governance database. Extracts editorial boards and program committees from journal websites using local LLMs, with entity resolution against OpenAlex.
To associate your repository with the llm-extraction topic, visit your repo's landing page and select "manage topics."