Stars
BioMedArena: a state-of-the-art biomedical harness for evaluating AI agents at scale - 100+ benchmarks, 70+ tools
📖 This is a repository for organizing papers, codes, and other resources related to personalized video generation and editing.
[npj Digital Medicine] A multimodal multidomain multilingual medical foundation model for zero shot clinical diagnosis
[EMNLP2024] Benchmark for "Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive Benchmark"
[Nature Reviews Bioengineering🔥] Application of Large Language Models in Medicine. A curated list of practical guide resources of Medical LLMs (Medical LLMs Tree, Tables, and Papers)
[NeurIPS 2022] Code for "Retrieve, Reason, and Refine: Generating Accurate and Faithful Discharge/Patient Instructions"
Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
The PyTorch code of the AAAI2021 paper "Non-Autoregressive Coarse-to-Fine Video Captioning".
Code and Resources for the Transformer Encoder Reasoning and Alignment Network (TERAN), accepted for publication in ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)
A curated list of radiology report generation (medical report generation) and related areas. :-)
本项目用于Multimodal领域新手的学习路线,包括该领域的经典论文,项目及课程。旨在希望学习者在一定的时间内达到对这个领域有较为深刻的认知,能够自己进行的独立研究。
A Beamer Theme of PKU for academic report, thesis and talk.
This repository focus on Image Captioning & Video Captioning & Seq-to-Seq Learning & NLP
Vision-Language Pre-training for Image Captioning and Question Answering
Unofficial pytorch implementation for Self-critical Sequence Training for Image Captioning. and others.
Code for "Understanding and Improving Layer Normalization"
Code for "simNet: Stepwise Image-Topic Merging Network for Generating Detailed and Comprehensive Image Captions" (EMNLP 2018)
We are building an open database of COVID-19 cases with chest X-ray or CT images.
Code for ICLR 2020 paper "VL-BERT: Pre-training of Generic Visual-Linguistic Representations".
Research code for ECCV 2020 paper "UNITER: UNiversal Image-TExt Representation Learning"
Script that crawls meta data from ICLR OpenReview webpage. Tutorials on installing and using Selenium and ChromeDriver on Ubuntu.
Meshed-Memory Transformer for Image Captioning. CVPR 2020
Code for paper "Attention on Attention for Image Captioning". ICCV 2019
A curated list of image captioning and related area resources. :-)
A simple module consistently outperforms self-attention and Transformer model on main NMT datasets with SoTA performance.
Shared repository for open-sourced projects from the Google AI Language team.
Code for the paper "VisualBERT: A Simple and Performant Baseline for Vision and Language"