Hi there, this is Punit Vara. Please feel free to contact me on LinkedIn
Software/ML systems engineer with experience across backend engineering, distributed data systems, cloud/MLOps, and LLM infrastructure — increasingly specializing in GPU-accelerated inference, distributed AI systems, and high-performance ML infrastructure. I like understanding systems from first principles rather than treating frameworks as black boxes.
🌇 Experience
- Staff ML Engineer · Visa · Aug.2025-Present
- Machine Learning Engineer - III · moneyview · Oct.2023-Jul.2025
- Machine Learning Engineer · Levi Strauss & Co. · Aug.2021-Oct.2023
- Machine Learning Engineer · Intel Corporation · Jan.2019-Aug.2021
- Zephyr Kernel Engineer · Intel Corporation · Dec.2016-Dec.2018
🔬 Focus Areas
LLM Inference & Serving (KV cache, quantization, llama.cpp) · GPU Systems & Collective Communication (NCCL, All-Reduce) · Distributed Data Systems (Kafka, Spark, Trino) · ML Infrastructure & MLOps (Kubernetes, Terraform, MLflow)
🌱 Currently Exploring
GPU/NCCL communication primitives and distributed inference performance bottlenecks
📫 Contact
- LinkedIn: in/punitvara
- YouTube: @punitvara.
- Twitter/X: @punitvara