Skip to content

Tags: idiap/fast-transformers

Tags

v0.4.0

Toggle v0.4.0's commit message
Release v0.4.0

This release contains mostly improvements on kernels and a few features and
fixes. Namely:

- We have a new super fast causal-linear kernel written by NVIDIA's Julien
  Demouth
- We have faster clustered broadcast and clustered aggregate kernels written by
  Apoorv Vyas

What should have been in this release but isn't because I didn't have time to
work on it :-) :

- Fancier masking that allows for different masks per sample while maintaining
  backwards compatibility
- Checkpointing for training huge models on single GPU machines
- 16-bit kernels for linear, local and clustering

v0.3.0

Toggle v0.3.0's commit message
Bump the version to 0.3

v0.2.2

Toggle v0.2.2's commit message
Bugfix RecurrentTransformerDecoder mask import

0.2.1

Toggle 0.2.1's commit message
Bugfix random number generation for cluster_cpu

v0.2.0

Toggle v0.2.0's commit message
Release v0.2.0

- Add transformer decoders
  * Add TransformerDecoderLayer and TransformerDecoder for batch processing of
    sequences and training
  * Add RecurrentTransformerDecoderLayer and RecurrentTransformerDecoder for
    one element at a time decoding and inference
  * Recurrent decoders also require a recurrent cross attention version that
    avoids recomputing the projections for the keys and values
- Refactor the builders
  * Attention builders are now available that simplify building just the
    attention modules
  * Attention registry allows dynamic registration of attention implementations
    so that fancy attention implementations can be separately implemented as
    plugins
  * TransformerDecoderBuilder and RecurrentDecoderBuilder simplify building
    decoders

v0.1.3

Toggle v0.1.3's commit message
Fix the PyPI distribution

v0.1.2

Toggle v0.1.2's commit message
Fix typo in README