Shredword

A fast and efficient tokenizer library for natural language processing tasks, built with Python and optimized C backend.

Features

High Performance: Fast tokenization powered by optimized C libraries
Multiple Encodings: Support for various tokenization models and vocabularies
Flexible API: Easy-to-use Python interface with comprehensive functionality
Special Tokens: Built-in support for special tokens and custom vocabularies
Fallback Mechanisms: Robust error handling with fallback tokenization
BPE Support: Byte Pair Encoding implementation for subword tokenization
Word Tokenization: Fast word-level tokenization with contraction handling
TF-IDF Embeddings: Built-in TF-IDF vectorization with dense and sparse representations

Installation

pip install shredword

Quick Start

BPE Tokenization

from shred import load_encoding

tokenizer = load_encoding("pre_16k")

tokens = tokenizer.encode("Hello, world!")
print(tokens)

text = tokenizer.decode(tokens)
print(text)

print(f"Vocabulary size: {tokenizer.vocab_size}")
print(f"Special tokens: {tokenizer.special_tokens}")

Word Tokenization & TF-IDF Embeddings

from shred import WordTokenizer, TfidfEmbedding

tokenizer = WordTokenizer()
tokens = tokenizer.tokenize("Hello, world! This is a test.")
print(tokens)

embedding = TfidfEmbedding()
embedding.add_documents([
  "The quick brown fox jumps over the lazy dog",
  "Python programming is fun and exciting"
])

ids = embedding.encode_ids("The lazy fox")
dense_vec = embedding.encode_tfidf_dense("The lazy fox")
indices, values = embedding.encode_tfidf_sparse("The lazy fox")

embedding.save("vocab.txt")
loaded = TfidfEmbedding.load("vocab.txt")

Documentation

For detailed usage instructions, API reference, and examples, please see our User Documentation.

Supported Encodings

Shredword supports various pre-trained tokenization models. The library automatically downloads vocabulary files from the official repository when needed.

Contributing

We welcome contributions! Please feel free to submit issues, feature requests, or pull requests.

Development Setup

Clone the repository
Install development dependencies: pip install -r requirements.txt (there are none!)
Run tests: python -m pytest

Guidelines

Follow PEP 8 style guidelines
Add tests for new features
Update documentation as needed
Ensure all tests pass before submitting PRs

License

This project is licensed under the Apache 2.0 License - see the LICENSE file for details.

Support

Issues: Report bugs or request features on GitHub Issues
Discussions: Join community discussions on GitHub Discussions

Name		Name	Last commit message	Last commit date
Latest commit History 65 Commits
docs		docs
shred		shred
tests		tests
.gitignore		.gitignore
CMakeLists.txt		CMakeLists.txt
LICENSE		LICENSE
Manifest.in		Manifest.in
README.md		README.md
pyproject.toml		pyproject.toml

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Repository files navigation

Shredword

Features

Installation

Quick Start

BPE Tokenization

Word Tokenization & TF-IDF Embeddings

Documentation

Supported Encodings

Contributing

Development Setup

Guidelines

License

Support

About

Uh oh!

Releases

Packages

Uh oh!

Languages

License

delveopers/Shredword

Folders and files

Latest commit

History

Repository files navigation

Shredword

Features

Installation

Quick Start

BPE Tokenization

Word Tokenization & TF-IDF Embeddings

Documentation

Supported Encodings

Contributing

Development Setup

Guidelines

License

Support

About

Topics

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Languages

Packages