The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).
-
Updated
Nov 26, 2024 - Java
The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).
Elasticsearch File System Crawler (FS Crawler)
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
Spark-Crawler: Apache Nutch-like crawler that runs on Apache Spark.
A cross-platform command line tool for parallelised content extraction and analysis.
Use the Java Tika text extraction library on the .NET platform
Code for Machine Learning with TensorFlow: 2nd Edition Published by Manning Publications
Viewers for statistics and dashboarding of Domain Search Engine data
Apache Tika bindings for PHP: extract text and metadata from documents, images and other formats
Tika-Similarity uses the Tika-Python package (Python port of Apache Tika) to compute file similarity based on Metadata features.
ImageCat is an Apache OODT RADIX application that uses Apache Solr, Apache Tika and Apache OODT to ingest 10s of millions of files (images,but could be extended to other files) in place, and to extract metadata and OCR information from those files/images using Tika and Tesseract OCR.
Interactive Image similarity and Visual Search and Retrieval application
R Interface to Apache Tika
Extract and Visualize location from any file
Geographic Place, Date/time, and Pattern entity extraction toolkit along with text extraction from unstructured data and GIS outputters.
Distributed, fault tolerant batch processing for Natural Language Applications and Search, using remote partitioning
Apache NiFi Custom Processor Extracting Text From Files with Apache Tika
Add a description, image, and links to the tika topic page so that developers can more easily learn about it.
To associate your repository with the tika topic, visit your repo's landing page and select "manage topics."