-
Beijing University of Posts and Telecommunications
- Beijing
-
05:13
(UTC +08:00)
Stars
Official PyTorch implementation of the paper: UniGaze: Towards Universal Gaze Estimation via Large-scale Pre-Training.
This is the official implementation of our work entitled "Multistream Gaze Estimation with Anatomical Eye Region Isolation by Synthetic to Real Transfer Learning" accepted in IEEE Transactions on A…
An official implementation of "OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs"
⏰ Agenticly track worldwide conference deadlines (Website, Python Cli, Wechat Applet)
AI Agent Assistant & development framework that integrates lots of IM platforms, LLMs, plugins and AI feature, and can be your openclaw alternative. ✨
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
[CVPR2026] Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
[ICCV'25]DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
[TIP 2026] ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model
Code for the Gaze360: Physically Unconstrained Gaze Estimation in the Wild Dataset
aQiaoz / gaze-estimation
Forked from yakhyo/gaze-estimation👀 | MobileGaze: Real-Time Gaze Estimation models using ResNet 18/34/50, MobileNet v2 and MobileOne s0-s4 | In PyTorch >> ONNX Runtime Inference
Real-time gaze estimation with ResNet, MobileNet and MobileOne - PyTorch training, ONNX Runtime inference, pretrained weights.
Original reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering"
Original reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering"
Repository of notes, code and notebooks in Python for the book Pattern Recognition and Machine Learning by Christopher Bishop
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
[CVPR 2025] OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.
Customize your arXiv recommendation every day.
OpenMMLab Pose Estimation Toolbox and Benchmark.
An end-to-end library for editing and rendering motion of 3D characters with deep learning [SIGGRAPH 2020]
Efficient 3D human pose estimation in video using 2D keypoint trajectories
Phase-Functioned Neural Networks for Character Control
This repository contains the codes of "A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild", published at ACM Multimedia 2020. For HD commercial model, please try out Sync Labs