Repository navigation
Tags: idiap/fast-transformers
Tags
Release v0.4.0 This release contains mostly improvements on kernels and a few features and fixes. Namely: - We have a new super fast causal-linear kernel written by NVIDIA's Julien Demouth - We have faster clustered broadcast and clustered aggregate kernels written by Apoorv Vyas What should have been in this release but isn't because I didn't have time to work on it :-) : - Fancier masking that allows for different masks per sample while maintaining backwards compatibility - Checkpointing for training huge models on single GPU machines - 16-bit kernels for linear, local and clustering
Release v0.2.0
- Add transformer decoders
* Add TransformerDecoderLayer and TransformerDecoder for batch processing of
sequences and training
* Add RecurrentTransformerDecoderLayer and RecurrentTransformerDecoder for
one element at a time decoding and inference
* Recurrent decoders also require a recurrent cross attention version that
avoids recomputing the projections for the keys and values
- Refactor the builders
* Attention builders are now available that simplify building just the
attention modules
* Attention registry allows dynamic registration of attention implementations
so that fancy attention implementations can be separately implemented as
plugins
* TransformerDecoderBuilder and RecurrentDecoderBuilder simplify building
decoders