PyTorch implementation of a collections of scalable Video Transformer Benchmarks.
-
Updated
May 4, 2022 - Python
PyTorch implementation of a collections of scalable Video Transformer Benchmarks.
Developed the ViViT model for medical video classification, enhancing 3D organ image analysis using transformer-based architectures.
Python script to fine tune Open source Video Vision Transformer (ViVit) using HuggingFace Trainer Library
The dataset used for the "A non-contact SpO2 estimation using video magnification and infrared data" publication
Video vision transformers for hierarchical anomaly detection in video scenes.
A comparative study of ViViT, CNN-GRU sequence models for video action recognition using the UCF101 dataset
Unofficial Tensorflow implementation of the ViViT model architecture
This repository hosts the deep learning framework for a next-generation Digital Point-of-Care Testing (POCT) system.
Implementation and benchmarking experiment with the unfactorised Video Vision Transformer (ViViT).
LSC50 — multimodal Colombian Sign Language recognition over 50 signs across IMU, body/face/hand landmarks, and ViViT video pipelines.
Slip detection with Franka Emika and GelSight Sensors
A deep learning framework for predicting deforestation with spatio-temporal deep learning models (ResUNet, ConvLSTM, ViViT) and satellite data. Developed as part of my Master's Thesis.
Unofficial Home Assistant integration for Vivit Energy / Repsol Luz y Gas with usage, costs, bills and virtual battery.
Action recognition using VideoMAE, TimeSFormer and ViViT | 1,250 video clips | 93.62% Top-1 | 98.40% Top-5 | spatio-temporal localisation
Comparing video (ViViT) vs image (ConvNeXT) models for intracranial hemorrhage detection from CT scan sequences
To associate your repository with the vivit topic, visit your repo's landing page and select "manage topics."