Tags: phdgil/meta
Tags
release MeTA v1.3.7
Bug fixes
- Fix inconsistent behavior of `utf::segmenter` (and thus `icu_tokenizer`) for
different locales. Thanks @CanoeFZH and @tng-konrad for helping debug
this!
Enhancements
- Allow for specifying the language and country for locale generation in
setting up `utf::segmenter` (and thus `icu_tokenizer`)
- Allow for suppression of `<s>` and `</s>` tags within `icu_tokenizer`,
mostly useful for information retrieval experiments with unigram words.
Thanks @HusseinHazimeh for the suggestion!
- Add a `default-unigram-chain` filter chain preset which is suitable for
information retrieval experiments using unigram words. Thanks
@HusseinHazimeh for the suggestion!
release MeTA v.1.3.4
New features
- Support building with biicode
- Add Vagrantfile for virtual machine configuration
- Add Dockerfile for Docker support
Enhancements
- Improve `ir_eval` unit tests
Bug fixes
- Fix `ir_eval::ndcg` incorrect log base and addition instead of subtraction in
IDCG calculation
- Fix `ir_eval::avg_p` incorrect early termination
release MeTA v1.3.3
Bug fixes:
- Fix issues with system-defined integer widths in binary model files
(mainly impacted the greedy tagger and parser); please re-download any
parser model files you may have had before
- Fix bug where parser model directory is not created if a non-standard
prefix is used (anything other than "parser")
Enhancements:
- Silence inconsistent missing overrides warning on clang >= 3.6
release MeTA v1.3
New features:
- additions to the graph library:
* myopic search
* BFS
* preferential attachment graph generation model (supports node
attractiveness from different distributions)
* betweenness centrality
* eigenvector centrality
- added a new natural language parsing library:
* parse tree library (visitor-based)
* shift-reduce constituency parser for generating phrase structure
trees
* reimplementation of evalb metrics for evaluating parsers
* new filter for Penn Treebank-style normalization
- added a greedy averaged Perceptron-based tagger
- demo application for various basic text processing (profile)
- basic iostreams that support gzip compression (if compiled with ZLib
support)
- added iteration method for `stats::multinomial` seen events
- added expected value and entropy functions to `stats` namespace
- added `linear_model`: a generic multiclass classifier storage class
- added `gz_corpus`: a compressed version of `line_corpus`
- added macros for generating type safe identifiers with user defined
literal suffixes
- added a persistent stack data structure to `meta::util`
Enhancements:
- added operator== for `util::optional<T>`
- better CMake support for building the libsvm modules
- better CMake support for downloading unit-test data
- improved setup guide in README (for OS X, Ubuntu, Arch, and EWS/ENGRIT)
- tree analyzers refactored to use the new parser library (removes
dependency on outside toolkits for generating tree files)
- analyzers that are not part of the "core" have been moved into their
respective folders (so `ngram_pos_analyzer` is in `src/sequence`,
`tree_analyzer` is in `src/parser`)
- `make_index` now checks if the files exist before loading an index, and
if they are missing creates a new one (as opposed to just throwing an
exception on a nonexistent file)
- cpptoml upgraded to support TOML v0.4.0
- enable extra warnings (-Wextra) for clang++ and g++
Bug fixes:
- fix `sequence_analyzer::analyze() const` when applied to untagged
sequences (was throwing when it shouldn't)
- ensure that the inverted index object is destroyed first before
uninverting occurs in the creation of a `forward_idnex`
- fix bug where `icu_tokenizer` would output spaces as tokens
- fix bugs where index objects were not destroyed before trying to delete
their files in the unit tests
- fix bug in `sparse_vector::find()` where it would return a non-end
iterator when asked to find an element that does not exist
release MeTA v1.2 New features: - demo application for CRF-based POS tagging - nearest_centroid classifier - basic statistics library for representing relevant probability distributions - sparse_vector utility class Enhancements: - ngram_pos_analyzer now uses the CRF internally (issue \meta-toolkit#46) - knn classifier now supports weighted knn - classifier cross validation can now optionally create even class splits - filesystem::copy_file() no longer hangs without progress reporting with large files - CMake build system now includes INTERFACE targets (better inclusion as subproject in external projects) - MeTA can now (optionally) be built with C++14 support Bug fixes: - language_model_ranker scoring function corrected (issue \meta-toolkit#50) - naive_bayes classifier scoring corrected - several incorrect instances of numeric_limits<double>::min() replaced with numeric_limits<double>::lowest() - compilation fixed with versions of ICU < 4.4
PreviousNext