Sanjoy Das
Los Altos, California, United States
10K followers
500+ connections
View mutual connections with Sanjoy
Sanjoy can introduce you to 10+ people at NVIDIA
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Sanjoy
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
I'm a director of software engineering at NVIDIA, where I run several key cross-BU…
Articles by Sanjoy
-
Libraries Are Dead. Long Live Libraries.
Libraries Are Dead. Long Live Libraries.
In this new world of AI-based coding, the classical value proposition of libraries is beginning to look shaky…
98
7 Comments -
No compiler is sufficiently smartJan 8, 2026
No compiler is sufficiently smart
For any self-hosting compiler (i.e.
225
17 Comments
Activity
10K followers
-
Sanjoy Das reposted thisSanjoy Das reposted thisI'm building a new team at NVIDIA and looking for great engineers who an help build a new approach to execution on NVIDIA GPUs. This will be a great opportunity for someone to join on the ground floor on a project that has the potential to have a big impact within NVIDIA and across the ecosystem. We are looking for both junior and senior engineers who have deep understanding of state-of-the-art LLM models and who can build high-performance GPU implementations for training, RL, and disaggregated inference. https://lnkd.in/ggzeptkR
-
Sanjoy Das reposted thisSanjoy Das reposted thisCuTeDSL 4.7 is out. The DSL is now general purpose, as of 4.7, it's a language you can write any kind of program in. The short version: → CUTLASS Primitives — close-to-metal abstractions → A Task Scheduling framework — provides static analysis of execution schedules for warp-specialized kernels. Large release, lots of people behind it. https://lnkd.in/ethdmqxjGitHub - NVIDIA/cutlass: CUDA Templates and Python DSLs for High-Performance Linear AlgebraGitHub - NVIDIA/cutlass: CUDA Templates and Python DSLs for High-Performance Linear Algebra
-
Sanjoy Das shared thisHighly visible, high impact role in an excellent team!Sanjoy Das shared thisWorking at NVIDIA is difficult to describe—it is something you experience. The ambition is enormous, the technical bar is high, and the work crosses organizational boundaries. You may join one team, but you quickly realize that the larger NVIDIA mission is what brings people together. I am hiring for Principal Product Managers to work with me: Reinforcement Learning Product Manager https://lnkd.in/gzhUP2cn PyTorch Compiler Product Manager https://lnkd.in/gh4V323Y These roles sit at the intersection of open-source AI frameworks, accelerated computing, researchers, engineers, and state of the art models and workloads A few examples of work our team has contributed to recently: World-record MoE pretraining performance with open-source frameworks including Megatron Core, PyTorch-native TorchTitan, and JAX: https://lnkd.in/gXxQtm72 Post-training technology enablement - Miles and on-policy distillation: https://lnkd.in/guDE3caf We are looking for people who enjoy technically deep problems, can turn emerging technology into a clear product strategy, and want to help shape how the AI ecosystem trains and post-trains the next generation of models. Come work with me.
-
Sanjoy Das reposted thisSanjoy Das reposted this10x tokens per megawatt. First measured silicon numbers on NVIDIA Vera Rubin NVL72, from CoreWeave. Not projections — real results from live hardware. 10x more tokens per second per megawatt than GB200 NVL72 at matched interactivity, running DeepSeek-R1 with NVIDIA TensorRT-LLM and NVIDIA Dynamo.First-Ever Measured Vera Rubin NVL72 Silicon Performance Stats | CoreWeave BlogFirst-Ever Measured Vera Rubin NVL72 Silicon Performance Stats | CoreWeave Blog
-
Sanjoy Das posted thisSomeone recently asked me how to think about software engineering -- and their career -- as AI changes the industry. My take is that there are two phases to this transition. In the near term, the winners will be the people who learn how to use AI tools to become more productive without sacrificing quality, maintainability, or good engineering practices. In my corner of the industry, this first phase is well understood, and a lot of people are already acting on it. But I think much bigger changes are coming. Computing has gone through several major shifts: from mainframes, to desktop computing, to cloud and mobile computing, with the browser and smartphone acting as thin clients. Each shift created new winners -- usually the people who saw the change coming and prepared for it. I expect something similar to happen to the nature of software itself. The exact shape of this next "Software 4.0" era is still pretty hazy since we're very early. But one signal I'm watching closely is the emergence of products and business models that simply could not have existed without AI. For instance, think of Uber and Google Docs in the cloud-and-mobile era, or Excel and Word in the desktop era. These were not just "faster horses". They were systems native to the new computing paradigm that could not have existed earlier. I think the most important AI-native products will follow the same pattern. They won't just be existing applications with an AI button. They'll do things that were previously impossible, creating a discontinuity in what software is and where durable value ultimately accrues.
-
Sanjoy Das reposted thisSanjoy Das reposted thisI’m #hiring a Senior Systems Software Engineer to join NVIDIA’s Compute Stack Acceleration effort. We’re looking for an experienced systems engineer who enjoys working across compilers and toolchains, build and CI systems, diagnostics, profiling, and performance optimization, and who can turn promising platform capabilities into measurable improvements across real software components. This is a broad, cross-stack role with close collaboration across compiler, platform, performance, library, and application teams. If that sounds like you or someone in your network, please apply or reshare: Senior Systems Software Engineer, Compute Stack Acceleration #NVIDIA #Hiring #SystemsSoftware #Compilers #PerformanceEngineering
-
Sanjoy Das reposted thisSanjoy Das reposted thisOur push for inference performance and efficiency continues - with Claude running on NVIDIA GB300, we take another big step towards higher speed and cost efficiency resulting in significantly lower TCO. "With Claude in Foundry running on NVIDIA GB300 NVL72 systems with NVIDIA Quantum-X800 InfiniBand networking, enterprises can now build and run more powerful agentic systems, including autonomous and specialized sub-agents that can work across business domains to perform advanced tasks. NVIDIA is working with Anthropic to extend developer capabilities by integrating NVIDIA tools into the Anthropic stack. That integration enables enterprises to give Claude agents domain-specific abilities. Through NVIDIA verified agent skills, enabled by access to NVIDIA accelerated computing, enterprises can embed AI agents deeply into their business and use them as the operating system for the organization. Enterprises can run Claude agents on Azure by using the NVIDIA Secure Agent Workspace Reference Design. It provides a blueprint for running autonomous agents in a governed environment where identity, network access, credentials and runtime policy are controlled at the infrastructure level." https://lnkd.in/gncVWw2dClaude Meets Blackwell Ultra: Anthropic’s Models Now Run on NVIDIA GB300 in AzureClaude Meets Blackwell Ultra: Anthropic’s Models Now Run on NVIDIA GB300 in Azure
-
Sanjoy Das posted thisWe're hiring compiler engineers with deep systems expertise to work on LPUs. You may be a strong fit if you've worked on compilers for GPUs, LPUs, or performance-sensitive ML systems. Compared to GPUs, LPUs make different architectural bets to enable ultra-low-latency LLM inference: deterministic execution, on-chip SRAM as primary storage, and software-scheduled chip-to-chip communication. That pushes a lot of the tough architectural problems into the compiler stack. If this kind of work interests you, please reach out on LinkedIn.
-
Sanjoy Das reposted thisSanjoy Das reposted thisOver the past couple of years, we have been working on improving the reliability and safety of our GPU programs. A major milestone of that effort is cuTile Rust, a programming model in Rust which brings Rust's fearless concurrency to the GPU, and our paper describing it: "Fearless Concurrency on the GPU." Rust gives you fearless concurrency on the CPU, but GPU kernel programming still requires unsafe code. cuTile Rust carries Rust's ownership model across the launch boundary: You partition a mutable output into disjoint tiles, and each tile program gets an exclusive mutable (&mut) view of its piece plus shared read-only (&) access to the inputs. You get compile-time data-race freedom on the safe surface, and it's extensible in the usual Rust way: You can wrap unsafe code yourself into safe abstractions. On the host side, you compose a static graph of GPU computations and execute it synchronously, asynchronously, or as a replayable CUDA graph. The safety is effectively free. On a B200, the safe GEMM is competitive with cuBLAS: about 2 PFlop/s, roughly 92% of the GPU's dense f16 peak, and within 0.3% of a hand-written low-level version. The safe, high-level kernel runs as fast as the hand-tuned one. Element-wise kernels hit ~7 TB/s, around 91% of peak memory bandwidth. It holds up end to end, too. Grout, a Qwen3 inference engine we built on cuTile Rust with Hugging Face, reaches 171 tokens/s for Qwen3-4B on a consumer RTX 5090 and 82 tokens/s for Qwen3-32B on a B200 at batch-1 decode. By our HBM roofline analysis, that's competitive with state of the art on memory-bound inference. This is an early-stage research release. It's co-evolved with community feedback since we made the repo public, so feedback and contributions are welcome. We'll also be giving the talk about this work at RustConf 2026. Please come check it out if you're attending. Thanks to my co-authors Jared Roesch, Isaac Gelado, Michael Garland, and Eric Buehler (Hugging Face). And thank you to all of the teams at NVIDIA who have helped make this possible (PSA, ARG, CUDA, Tile IR). - Code: https://lnkd.in/gwKS36A4 - Paper: https://lnkd.in/g5kiBtzz - Talk: https://lnkd.in/gq_zEDzS #rustlang #GPU #CUDA #MachineLearning #compilersGitHub - NVlabs/cutile-rs: cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions.GitHub - NVlabs/cutile-rs: cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions.
-
Sanjoy Das liked thisSanjoy Das liked thisMassive open-weight model drop! Congratulations to Qwen on releasing the open weights for Qwen3.8-2.4T-A95B, a 2.4T parameter model with 95B active, designed for demanding reasoning and agentic workloads. Out of the box, the model achieves a throughput of 4K+ tokens/second per GPU and 350+ tokens/second per user on NVIDIA GB300 NVL72 in FP8 precision. Read the tech blog: https://nvda.ws/463kFdP
-
Sanjoy Das liked thisSanjoy Das liked thisI’m hiring two hands-on Solutions Architecture Managers to join my team at NVIDIA and lead our work with some of the world’s most advanced AI labs. These are leaders who can build strong technical teams while staying close to the work - designing, deploying, and debugging large-scale GPU systems and AI networking infrastructure in production. With deep experience in GPU systems, Linux, Ethernet and/or InfiniBand networking, data center infrastructure, cluster bring-up, and performance troubleshooting—along with a strong track record of developing people and partnering with customers. If this sounds like you—or someone in your network—I’d love to hear from you. Please apply or share. Open roles: Manager, Solutions Architecture – AI Labs https://lnkd.in/gzGZWzQn Manager, Solutions Architecture – Emerging AI Labs https://lnkd.in/gksWTAXB #NVIDIA #Hiring #AIInfrastructure #Systems #Networking #LeadershipManager, Solutions Architecture - Emerging AI labsManager, Solutions Architecture - Emerging AI labs
-
Sanjoy Das liked thisSanjoy Das liked thisI'm building a new team at NVIDIA and looking for great engineers who an help build a new approach to execution on NVIDIA GPUs. This will be a great opportunity for someone to join on the ground floor on a project that has the potential to have a big impact within NVIDIA and across the ecosystem. We are looking for both junior and senior engineers who have deep understanding of state-of-the-art LLM models and who can build high-performance GPU implementations for training, RL, and disaggregated inference. https://lnkd.in/ggzeptkR
-
Sanjoy Das liked thisSanjoy Das liked thisCuTeDSL 4.7 is out. The DSL is now general purpose, as of 4.7, it's a language you can write any kind of program in. The short version: → CUTLASS Primitives — close-to-metal abstractions → A Task Scheduling framework — provides static analysis of execution schedules for warp-specialized kernels. Large release, lots of people behind it. https://lnkd.in/ethdmqxjGitHub - NVIDIA/cutlass: CUDA Templates and Python DSLs for High-Performance Linear AlgebraGitHub - NVIDIA/cutlass: CUDA Templates and Python DSLs for High-Performance Linear Algebra
-
Sanjoy Das liked thisSanjoy Das liked thisAt NVIDIA, we are looking for software engineers to contribute to the design and development of libraries and tools to simplify and accelerate computing for unstructured sparsity in DL and HPC. If your background, interest, and energy align, consider applying today. https://lnkd.in/eqkwpnkZ
-
Sanjoy Das liked thisSanjoy Das liked thisGiving harsh feedback is easier than giving corrective feedback to a high performer. Which probably sounds a bit confusing. Telling someone "you broke our team's policy and your job is at risk" is emotionally hard in the moment, but it's intellectually simple. The rule is clear, and there's no ambiguity. You take a deep breath, point at the line they crossed, and you've done your job. Coaching a good performer is a different problem entirely. There's nothing obviously wrong, or you would have told them. However, coaching good performers brings much more value than coaching poor performers. Boosting a top performer by 10% is often worth more than improving your bottom performer by 50%. But so many managers spend the majority of their time on the (comparatively less valuable) poor performer. Here's an example I use when explaining how we should be dealing with high performers. A solid Level 6 engineering manager at Amazon finishes a presentation to a VP. It went well. They ask their manager how it went. The easy answer (which I bet 90% of solid performers will hear from their managers): "Nice job, keep doing what you're doing." I hate this phrase. It gives the person nothing to act on. Here's the type of feedback I think we should say instead: "The ticket data breakdown was exactly what's expected at your level. Good job. A Senior Manager would also break down root causes, with dates for fixes. Perhaps also linking to the tracking tickets for those fixes. You didn't need that for this audience, but the VP would have found it useful. As you grow toward that next level, start thinking about what we'll expect from someone at that level." Notice what's happening. I'm not saying "you failed to do X." I'm describing behavior at the next level up, not their current one. Good performers are usually doing their current job just fine. Unlike poor performance, we're not talking about how they've messed up, or their incompetence. We're instead giving them an example of the next level in concrete terms, because specific definitions are not written anywhere. The specific difference is these small nuances, and this is an excellent opportunity to share them. To learn more about the behaviors that separate senior engineers from mid-level ones, read on.The Behaviors That Turn Engineers Into Senior EngineersThe Behaviors That Turn Engineers Into Senior Engineers
Experience
Education
-
Indian Institute of Technology Kharagpur
-
-
Languages
-
English
-
-
Bengali
-
View Sanjoy’s full profile
-
See who you know in common
-
Get introduced
-
Contact Sanjoy directly
Other similar profiles
Explore more posts
-
VJ Anand
Automation Anywhere • 2K followers
DSPy: The compiler paradigm for language models fundamentally transforms language model programming from manual prompt engineering into systematic software development, treating prompts as learnable parameters that compile into optimized implementations. Here is my recent review of that framework - a game changing framework that removes the manual and tedious work of prompt design to a more structured program development. Please review my recent blog on this framework: https://lnkd.in/g8AZTJiv
6
-
Xinyi Chen
CommonWealth Magazine… • 1K followers
NVIDIA GTC just shipped Dynamo 1.0 at GTC. Everyone's watching Blackwell benchmarks. I'm looking at the integration list — and one name stands out: LMCache. Two weeks ago, my OSS-Investment-Scorecard evaluated LMCache Lab at 7.78/10 — Yellow Recommendation. The scores told a specific story: Technical Moat: 8.5 — KV-cache disaggregation across inference instances is a genuinely hard systems problem with measurable cost impact. This is the highest dimension score in the evaluation. Ecosystem: 8.0 — active vLLM integration, growing star count, steady PR cadence. Exit Path: 7.5 — natural acquisition targets include cloud providers and LLM serving platforms. Where it falls short: Commercialisation & PMF at 4.5. No public ARR, no disclosed enterprise customers. The vLLM ecosystem pull is a positive proxy, but it's not revenue. Then GTC happened. NVIDIA officially integrated LMCache into Dynamo's orchestration stack — alongside LangChain Labs, SGLang, and vLLM. That's a material signal for the PMF dimension that wasn't priced into my original evaluation. This is exactly what the scorecard is built to do: flag projects with strong technical foundations and clear ecosystem positioning before the market catches up. The gap between C-dimension strength (8.5) and D-dimension weakness (4.5) was the signal — a project solving a real infrastructure problem that hadn't yet found its commercial surface. NVIDIA AI just provided that surface. 🤩
18
1 Comment
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content