Stars
Builds wordpiece(subword) vocabulary compatible for Google Research's BERT
Code for CEDR: Contextualized Embeddings for Document Ranking, accepted at SIGIR 2019.
Datasets for EMNLP-IJCNLP 2019 paper "NCLS:Neural Cross-Lingual Summarization"
Snips Python library to extract meaning from text
中文人名语料库。人名生成器。中文姓名,姓氏,名字,称呼,日本人名,翻译人名,英文人名。可用于中文分词、人名实体识别。
Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
Distributed training framework for TensorFlow, Keras, PyTorch, and Apache MXNet.
Convolutional neural network and word embeddings for Chinese word segmentation
An Open Source Machine Learning Framework for Everyone