A hybrid testbed for evaluating top open-source LLMs (like gpt-oss-20b and Llama 3.3) on local, cloud GPUs, and AWS Inferentia2/Trainium instances, focusing on vLLM optimization, capacity management, kernel bypass, hardware-software co-design, as well as supporting infrastructure such as NCCL, RDMA, NVMeoF.
aws gpu rdma nvme kernel-bypass nccl gpudirect nvlink nvmeof llm vllm trainium vllm-serve inferentia2 software-hardware-co-design aws-ofi-nccl
-
Updated
Apr 21, 2026 - Python