The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).
-
Updated
Sep 24, 2026 - Java
The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).
基于 Spring Boot 4.1、Java 25、Spring AI 2.0、React、PostgreSQL/pgvector、Redis 和 RustFS 构建的开源 AI 面试平台,支持简历智能分析、模拟面试、语音面试和知识库 RAG。
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
Elasticsearch File System Crawler (FS Crawler)
A crawl workstation: View, Control, and Crawl. Solr CrawlDB, Tika, Vue 3.
A cross-platform command line tool for parallelised content extraction and analysis.
Use the Java Tika text extraction library on the .NET platform
Code for Machine Learning with TensorFlow: 2nd Edition Published by Manning Publications
Apache Tika bindings for PHP: extract text and metadata from documents, images and other formats
Tika-Similarity uses the Tika-Python package (Python port of Apache Tika) to compute file similarity based on Metadata features.
RADiX overlay on Mnemosyne 1.11.0: Apache Solr 10, Apache Tika, Paddle/RapidOCR (TrOCR/Donut optional). Ingest in place, MIME/EXIF + OCR into Solr. Vue OPSUI at /opsui/. ImageSpace (search, CLIP similar, fg/bg, Lens) on :8090.
R Interface to Apache Tika
Distributed, fault tolerant batch processing for Natural Language Applications and Search, using remote partitioning
Geographic Place, Date/time, and Pattern entity extraction toolkit along with text extraction from unstructured data and GIS outputters.
Apache NiFi Custom Processor Extracting Text From Files with Apache Tika
To associate your repository with the tika topic, visit your repo's landing page and select "manage topics."