Extract multi-word domain terms from text using the C-value algorithm and a complete Python pipeline.
-
Updated
Sep 1, 2026 - Jupyter Notebook
Extract multi-word domain terms from text using the C-value algorithm and a complete Python pipeline.
A definitive blueprint for mastering Natural Language Processing. Bridging computational linguistics with deep learning, this rigorous curriculum systematically covers text pre-processing, statistical modeling, distributed word semantics, sequence-to-sequence architectures, and attention mechanisms for robust text analytics.
🌟 Build and manage your tasks effortlessly with blabla, a simple yet powerful tool for enhanced productivity and organization.
🌟 Build an LSTM-based sentiment analysis pipeline with robust tokenization, class handling, and efficient training for accurate results.
Voynich Manuscript Decoded: Elu-Sinhala Phonetic Transcription & Vocabulary Toolkit 2026
Reproducible, provenance-first computational research lab for the Voynich Manuscript.
Benchmarks vector representations of the Hebrew Psalms against scholarly annotations of poetic parallelism and genre, using retrieval-based evaluation with bootstrap confidence intervals and permutation inference.
A deeply structured edition of Littré's Dictionnaire de la langue française. Available as TEI Lex-0 XML and SQLite.
List of personal Kaggle contributions.
Cavar's homepage
Category-theoretic treatment of linguistic case integrated with Active Inference and the CEREBRUM architecture: DisCoPy string diagrams spanning typology, categorial grammar, topos theory, and quantum extensions, generating 30 publication figures and a 24-section manuscript. 1,207 tests, 95.96% coverage, zero mocks.
Release package for SlopShape: Identifying AI-Generated Commercial Web Content (verification artifacts, instrument, prompts, code)
Data and software for building the ACL Anthology.
Orthographic Normalization package that cleans up Yorùbá text encoding. It also draws the line most preprocessing code misses: underdots change the word, tone marks do not.
Morphological analyzer and rule engine for Turkish. Derives word structure from root and affix rules instead of pattern-matching the surface, so every phonological event comes with evidence.
A course introducing computational linguistics to advanced undergraduates and early graduate students in linguistics.
A declarative, phonology-oriented transliteration engine for Cantonese and other languages.
Run BGE-small embeddings with low latency and minimal RAM using a dependency-free C engine for Python, Node, Go, Rust, and C.
Generate endless game worlds through conversational storytelling with multi-agent support for scenes, stats, and replay features.
To associate your repository with the computational-linguistics topic, visit your repo's landing page and select "manage topics."