Renaud Gaubert
San Francisco Bay Area
1K followers
500+ connections
View mutual connections with Renaud
Renaud can introduce you to 10+ people at OpenAI
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Renaud
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
I am a driven Software Engineer experienced in Cloud, DevOps, Linux Containers…
Activity
1K followers
-
Renaud Gaubert shared thisCome work with Kenny and me on one of the most ambitious data center buildouts in the world! We’re looking for an exceptional TPM to help make it happen 💪Renaud Gaubert shared thisAs we pursue our mission to “build the world’s largest compute footprint so frontier AI can reach everyone and every workflow,” we need substantial support to scale and mature our engineering, business, and operational execution. We're hiring for the following in the compute space! - TPM, Compute Infra: https://lnkd.in/g4a2FbjH - TPM, Core Network & WAN Infra: https://lnkd.in/geP7ns7y - Data Center, Compute Infra: https://lnkd.in/gdgseEVH
-
Renaud Gaubert reposted thisRenaud Gaubert reposted thisi’m hiring! in this role, you’d define and drive the infrastructure powering our training and product workloads—from datacenter strategy and hardware choices to scaling next-gen systems. high ownership, high leverage. apply here: https://lnkd.in/gnhE6zEqTechnical Program Manager, Infrastructure Foundations | OpenAITechnical Program Manager, Infrastructure Foundations | OpenAI
-
Renaud Gaubert shared thisOpenAI is growing! We're building a Supercomputing (infra for research) team! If you're interested in huge systems, datacenters and large training jobs, do apply! https://lnkd.in/g56Kbw_5
-
Renaud Gaubert shared thisJacob is an amazing person to work with, and this is a super cool field to work in!Renaud Gaubert shared thisMy team is hiring! Help build our cloud native platform for edge AI. Remote candidates welcome to apply. #kubernetes #openshift
-
Renaud Gaubert shared thisRenaud Gaubert shared thisDear Network, All my wishes for this new year! Am looking for an extraordinary, high-energy individual to join our Santa Clara team and manage our AI startup program for North America. If you know of a good fit, please ask them to apply using the below link and to drop drop me a quick note as a heads up. https://lnkd.in/eTtF7_f #hiring #inception #artificialintelligence
-
Renaud Gaubert shared thisWe are looking for someone to help us scale Deep Learning, Containers, Kubernetes and GPUs in tomorrow's data center! If working at the intersection of Linux, Distributed Systems and Deep Learning sounds interesting to you, we'd be excited to talk to you! #gpus #datacenter #deeplearning #linux #kubernetes #containers
-
Renaud Gaubert shared thisWe are looking for someone to help us scale Deep Learning, Containers, Kube and GPUs in tomorrow's data centers! If working at the intersection of Linux, Distributed Systems and Deep Learning sounds interesting to you we'd be excited to talk to you! #gpus #datacenters #deeplearning #linux #kubernetes #containers
-
Renaud Gaubert shared this#ANSSI10 Hier après-midi les 41 membres de l' Agora41@ANSSI ont présenté l’avancée de leurs travaux au public. Lors de cette intervention, nous avons eu l’honneur de partager nos axes de travail avec Monsieur le député Cedric Villani. Alix Durand Pauline Flury Baptiste Fortin Caroline Faillet Jean-Baptiste Clais Xavier Lepage Romain Galesne-Fontaine Eric Danon Marie-Christine Dupuis-Danon Philippe Lavault Dr Helene Lavoix Clotilde Bômont Renaud Gaubert Emmanuel Germain
-
Renaud Gaubert liked thisI’m excited to launch Racktwin: https://racktwin.com Modern AI infrastructure is pushing physical systems to a level of complexity that deserves better tooling. We’ve built strong software layers for logical infrastructure. But the physical layer — topology, cabling, validation, deployment accuracy, and operational change — still too often lives across spreadsheets, diagrams, PDFs, and tribal knowledge. Racktwin is my attempt to change that. The idea is to make the physical datacenter more like software: something that can be modeled precisely, analyzed systematically, validated before errors happen, and used by other systems over time. The images below show both sides of that idea: a declarative spec and the physical infrastructure model it produces. It’s still early, but I’m excited to finally put it into the world. If you work on AI clusters, physical network design, datacenter deployment, or operations, I’d love your feedback. #AIInfrastructure #datacenter #networking #infrastructureascode #HPC #startup #fiberoptics
-
Renaud Gaubert liked thisRenaud Gaubert liked thisi've contributed towards ~200 of the world's largest computers in the past few years (i feel very lucky to say this) they produce and run systems i would've thought impossible as recently as 2022 i'm most looking forward to what the largest computer ever leads to in 2026 :)
-
Renaud Gaubert liked thisRenaud Gaubert liked thistl;dr: I’ve left Anthropic and am taking some time off, hopefully indefinitely. After a wild year at Anthropic, I left the company this past week. Of all of the organizations working on AI, they have my support, but I’ve been working on this stuff for many, many years now and am looking forward to a break. My team and I did great stuff at Anthropic. We hired amazing people and improved availability while scaling >20x in usage and 10x in revenue over the past year. The world discovered Claude and the models are critical, especially for software engineering work. The reliability (and correctness/quality) of these models is increasingly important. Anthropic will continue to improve the availability of the serving stack, and the quality of the underlying models. I’ve spent the last 17 years working on big ML systems and the last four working at the white hot center of the AI shenanigans. I built ML SRE at Google and helped oversee the transition of those teams’ infrastructure. I spent a wild ~year at Open AI starting a research reliability team from scratch and working on some of the early phases of Orion training (which probably became Chat GPT 4.5 after I left). And then I spent the next year at Anthropic building a reliability team from scratch. We went from zero people to over 20 people in two teams, serving and training. AI is changing the world, in some great and some problematic ways. There is no bottle to put the genie back into. We have to grapple with all of these changes together, while addressing broken politics and policy making. I trust Anthropic to make the most serious and effective efforts to smooth this transition. They are genuinely good people who care about humanity as a whole. They want to mitigate the potential challenges on employment, and education, and mental health, and information security that may come from thoughtless use of AI. I would bet on Anthropic getting these things right. My career has been super weird. In 1997, I got a sysadmin gig at a small nonprofit that happened to be the largest “Internet” service provider in New Mexico (AS2901). I spent the next 12 years in the middle of the most important technological transformation of the 20th century. Next, I got a job working on something called “SmartASS” at Google. This turned out to be one of the world's biggest ML systems (I knew nothing about ML prior). And, again, I stumbled across what is, so far, the most important technological transformation of the 21st century. When early career folks ask me for advice, I’m genuinely stumped. “Be lucky” is probably top of my list. That’s really not that helpful. I’m sorry. I’m going to try not working for a while and see how it goes. I’ve been working since I was about 15 years old. I’m going to reevaluate how I spend time, connect with friends and family I haven’t seen much for a while, and come up with some plans within a few months. Or not. I’m excited to see how it goes.
-
Renaud Gaubert liked thisRenaud Gaubert liked thisAfter 6.5 unforgettable years at Microsoft, I’m closing this chapter with a full heart. 💙 When I joined, I never imagined I’d spend these years helping build the supercomputers that powered the breakthroughs behind OpenAI and made ChatGPT possible. What started as an ambitious idea grew into the world’s largest AI GPU fleet—supporting Microsoft, OpenAI, Anthropic, and some of the most exciting innovation happening anywhere in technology. It’s been the privilege of a lifetime to help build the infrastructure that’s redefining what’s possible. But as proud as I am of the work, what I’ll miss most are the people. The teams who pushed boundaries, solved the impossible, and did it all with humility, grit, and a sense of mission. Thank you for the trust, the collaboration, the laughter, and the late-night problem-solving that somehow always ended in progress. This jacket captures the experience so well—it has been through data centers, war rooms, launches, milestones, and a whole lot of chaos and triumph. Chaos that felt like so much fun! Now on to the next chapter 💙 💙
-
Renaud Gaubert liked thisRenaud Gaubert liked thisGreat to be back at #AIWorld to share foundational innovations across Networking and AI infrastructure that are key to delivering the next generation of AI to our customers. Critical to our success is working closely with our partners - NVIDIA, AMD, Arista Networks and more. Our long-standing, customer first partnerships enable our customers, such as Uber , to provide market-leading experiences. Key themes from my keynote: - OCI's technology and architecture are helping Uber prioritize passenger safety through robust infrastructure and advanced solutions. Thanks Albert Greenberg - Oracle Acceleron, OCI’s suite of network software and architecture, is designed to help customers run any workload faster and more cost-effectively. - OCI Zettascale10 clusters, powered by NVIDIA AI infrastructure and Oracle Acceleron RoCE network architecture, deliver unprecedented multi-gigawatt AI capacity. Thanks Ian Buck. - OCI will be launch partner for the first publicly available AI supercluster powered by 50,000 AMD Instinct™ MI450 Series GPUs. Thanks Forrest Norrod - Together with Arista, we’ve pioneered networking innovations that power everything from Oracle Exadata to large-scale GPU clusters and deliver exceptional customer experiences. Thanks Kenneth Duda Excited to continue building with our partners and customers to deliver next-generation cloud infrastructure for the AI era!
-
Renaud Gaubert liked thisRenaud Gaubert liked thisToday is my last day at Microsoft. After almost 13 years, I’m leaving grateful for the people who taught me, challenged me, and cheered me on; for teams that cared about doing things right; and for the chance to learn every single day. I'm incredibly excited about what’s next! More to follow soon... For now: thank you to everyone I’ve worked with along the way. From mentors, partners, customers, and friends. You made this journey meaningful. Obligatory photos. #MicrosoftAlumni
-
Renaud Gaubert reacted on thisRenaud Gaubert reacted on thisWe too have fled the US. We made the decision in March, have been in Sweden since June, and are fully intending to stay here for the long term. We sincerely hope that we’ll look back in a year and laugh at ourselves for having punched the eject button, but despite all the stresses of relocation, this feels good, safe, and right. I know there’s a lot of folks who would do the same but can’t; our hearts go out to all of you, and I hope you find a safe path forward ❤️ OpenAI couldn’t move the job with me, unfortunately, but as a silver lining I’ve had nearly three months to focus on easing my family through this move, and I’ve been savouring it: it has been easily 20 years since I last took a break this long. But I have also been thinking about what comes next. It is astonishing how quickly the USA has devolved, and alarming that the architects of that tragedy have set their sights on Europe. There are fewer bastions of liberal democracy left, and it feels like the barbarians are at the gates. No technology is a magic bullet, but tech can contribute to defense, and perhaps AI in particular can contribute to the kinds of cultural defenses that we need to strengthen today. Rather than AI safety in the abstract, we need more application of AI to actively defending the values and norms we cherish. More broadly, we need a tech industry that bolsters liberal society, rather than exploiting it, if we hope to preserve it. I’m looking forward to finding a part to play.
-
Renaud Gaubert liked thisRenaud Gaubert liked thisMicrosoft just leaked their official compensation bands for engineers. We often forget that you can be a stable high-performing engineer with great work-life balance, be a BigTech lifer and comfortably retire with a net worth of ~$15M! Source: https://archive.is/biMgu
Experience
Education
-
Ecole de Guerre Economique
-
-
-
-
-
-
-
Languages
-
French
Native or bilingual proficiency
-
English
Full professional proficiency
View Renaud’s full profile
-
See who you know in common
-
Get introduced
-
Contact Renaud directly
Other similar profiles
Explore more posts
-
Satishkumar Dhule
Salesforce • 3K followers
mistake: a $2 million memory miscalculation wrecked a GPU demo. Memory is the constraint that reveals real system limits - not the glossy dashboards. When data scales, every byte counts. A mis-sized buffer or cache can cascade into page faults and a failed run. 🔍 Size matters: budget memory against worst cases and verify with load tests. ⚡ Rehearse under real pressure: memory pressure kills demos. 🎯 API-first design (GraphQL federation, tRPC) helps fetch only what’s needed. 🛡️ Observability is mandatory: track allocations, GC pauses, and leaks. 💡 Start conservative, then tune with real traffic using traces. Memory tells the truth fast; plan for it or pay the price. ───────────────────────── 🔗 Read the full article: https://lnkd.in/dZfmPgqG 🎯 Practice interview questions: https://lnkd.in/gmy5drNw #operatingsystems #million #memory #mistake
3
-
Rudra Pratap Singh
Alta School of Technology • 179 followers
While solving LeetCode problems in C++, I started experimenting with memory allocation and saw how much performance can be influenced by low-level optimizations. That made me wonder: If C++ is so fast, why isn't it the go-to language for web applications? And then I discovered frameworks like Drogon. C++ absolutely can be used to build high-performance web applications and APIs. Drogon provides routing, HTTP handling, database support, asynchronous programming and more. So the question isn't whether C++ can do web development. It can. The real question is: why isn't it the mainstream choice? A major reason is the trade-off between performance and developer productivity. C++ gives incredible control over memory, CPU and system resources, but that control also comes with additional complexity. For many web applications, the bottleneck isn't the programming language. It's often: → Database queries → Network latency → External APIs → I/O operations So shaving microseconds from your C++ code may not matter if your database query takes 100ms. This made me realize: The fastest language isn't always the best tool. The right choice depends on the problem you're solving. And that's one of the things I love about learning C++ — it makes you think about what is actually happening underneath the abstractions. 🚀 #CPlusPlus #Cpp #Drogon #WebDevelopment #BackendDevelopment #SoftwareEngineering #Performance #DSA #LeetCode #LearningInPublic
3
-
Sameer Bhardwaj
Layrs • 59K followers
A candidate interviewing for a Senior Engineer @ Meta was asked to design a Distributed Stream Processing System like Kafka. Another candidate at Google's L5 loop got hit with the same question. If this question shows up in your system design interview, you are not being tested on config flags. You are being tested on whether you understand when to use it and what tradeoffs you are making. Btw, If you are preparing for DSA or system design, try our mock interview tool on Layrs for free:http://layrs.me/interviews Here is how I would structure a clear answer. (Note: I am not going into nuances, this is a very high-level explanation) 1. What problem are we solving At a high level, Kafka is my go-to when I need: - A central backbone for events or logs coming from many services - High throughput writes with ordering per key - Consumers that can read at their own pace and be added or removed independently - The ability to replay history using offsets Typical use cases I call out: - Asynchronous workflows - Central logging and metrics - Ad click or analytics streams - Pub sub for notifications or chat That already tells the interviewer I know -why- Kafka is in the picture. 2. Requirements Functional - Accept events from many producers in different services and regions - Preserve order for events that share the same key - Allow many independent consumers +some doing real time processing +some doing batch jobs or analytics - Support replay from an offset for backfills and reprocessing Non functional - Handle very high write rates - Scale horizontally by adding machines - Survive broker failures without losing committed data - Let us tune retention for cost vs replay needs Rest of the breakdown: https://lnkd.in/gmg7u-ku
384
12 Comments -
Alan Kochukalam George
Myovine • 617 followers
⚙️ One library, real DSP speed. There’s a gap in the developer ecosystem. Teams doing DSP/ML often migrate to Python, because that’s where most tooling lives. But Python wasn’t built for high-throughput I/O — under load, it spawns new interpreters/processes, creating duplicated memory, slow cold starts, and more containers. Compute stays fast, but orchestration cost explodes. Meanwhile, Node.js/TS scale I/O well, but were never meant for serious compute. Pure-JS DSP hits performance walls. Even worse, serializing data between Redis ↔ DSP/ML adds latency and forces batching — slowing real-time pipelines. So devs are stuck: Python → strong math, weak concurrency Node → strong networking, weak compute dspx closes that gap — TypeScript DX + native C++/SIMD (AVX2/SSE3/SSE2/NEON) under the hood. 🧩 How dspx works FIR / FFT / Conv1D run in optimized C++/SIMD Memory-safe circular buffers → O(1) throughput TS handles Kafka / Redis / WebSockets Redis persistence < 0.5 ms Batched logging avoids I/O stalls At tiny batch sizes, N-API overhead can make naive JS look competitive. At realistic scale, batched native pipelines win. ⚡ Benchmarks Dell OptiPlex 3000 Micro (i5-12600T · AVX2 · Node 22) FFT: 2.6× faster than fft.js, ~9× faster than tfjs-node FIR (51-tap): ~4× faster than fili / naive JS Conv1D (128-kernel): ~3.5× faster than tfjs-node Moving Avg (O(1)): 559× more throughput/sec than naive JS Redis save/load: sub-ms Logging: <3% overhead 📊 Full tables + charts in carousel 💬 Observations SIMD FMA + contiguous memory → big FIR/Conv wins O(1) moving-average design → massive throughput gain Sub-ms Redis + low I/O overhead → real-time persistence in pure Node 🧠 Open Source dspx is Apache 2.0, free for commercial/academic use. Looking for contributors interested in: ARM NEON / Apple M / Graviton tuning Audio / sensor / biomedical DSP Visualization + benchmark tooling 📦 npm i dspx 🔗 https://lnkd.in/e-tWxAgu 🧠 Real-time DSP for Node, TypeScript & Redis 💼 Note I’m exploring opportunities in full-stack, real-time systems, performance engineering, and DSP. If your team works in this space, I’d love to connect. My next post will cover why sub-ms Redis latency isn’t just technical — it’s economic. Lower serialization + compute overhead → lower infra + energy cost. 🔖 Tags & Mentions #NodeJS #TypeScript #DSP #PerformanceEngineering #EdgeComputing #OpenSource #SIMD #Redis #Cplusplus #RealTime NodeJS Developer TensorFlow Google Microsoft JavaScript Developer Amazon Web Services (AWS) Vercel
5
1 Comment -
Dilan Hendadura
Hoplo • 6K followers
Anthropic released Opus 4.5. Top of SWE bench at 80.9 percent. Nice chart, but the real story is somewhere else. Everyone compares the LLM race to the cloud race. Same TAM, same few winners. Wrong. In the cloud you do not switch with a single line of code. With LLMs you do. One upgrade from Gemini, GPT, or Claude and teams move millions in spend overnight. That switching pressure pushes prices down. Great for builders. Tough for model vendors. Every month the cost of inference drops while product value goes up. So the war is not model vs model. The war is between three approaches. Single vendor teams that hope their model never falls behind. Multi model teams that treat models like swappable parts. Vertical products that own workflows, data, and users. We already saw a one person AI platform (Maor Shlomo‘s Base44) get acquired for $80M and pass $100M in revenue later. This happens when you ride the model race instead of competing inside it. Opus 4.5 winning the chart is the headline. The real story is who builds something durable while models keep getting cheaper and easier to replace. If you build in this space, ask one thing. If your favorite model changed tomorrow, would your product still stand strong or fall with it
2
-
Youssef Alaa
Bargou • 1K followers
Your LLM API isn't slow because of matrix math. It’s slow because you're treating memory access like compute. Single-token generation is rarely bottlenecked by raw FLOPS. It’s bottlenecked by memory bandwidth specifically, moving billions of parameter weights from VRAM to SRAM for every single token generated. If you want to cut atency, focus on these three memory architecture shifts: 1. Disaggregate Prefill & Decode Nodes - Prefill (prompt processing) is compute-heavy. - Decode (token generation) is memory-bandwidth bound. - Serving both on the same GPU causes prefill bursts to starve decode throughput. Splitting them onto dedicated node pools eliminates this resource contention. 2. Prefix Caching via Radix Trees If you run multi-turn chats or RAG, re-computing system prompts is wasted latency. Frame-level caching (like SGLang's RadixAttention) treats the KV cache as a dynamic tree, skipping prefill entirely for shared context and slashing Time-To-First-Token (TTFT). 3. Speculative Decoding Stop using a 70B parameter model to generate predictable tokens. Let a fast 1B draft model generate K token proposals, then use the 70B target model to run a single parallel verification pass over all of them in one go. High-performance ML engineering isn't about throwing more GPUs at the problem. It's about keeping data in cache and avoiding redundant memory passes. #MachineLearning #AIEngineering #SystemDesign #MLOps #vLLM #LLMInference #SoftwareEngineering
-
Nishantha Ruwan
IWROBOTX Software Inc. • 2K followers
The paper identifies a critical performance limitation in multi-turn, agentic large language model (LLM) inference workflows: the key-value (KV) cache I/O from external storage becomes the dominant bottleneck as context lengths grow. In traditional disaggregated inference systems, prefill engines (which generate KV cache data) saturate their storage network interfaces while decode engines sit underutilized, leading to poor throughput and inefficient resource use. To address this, the authors introduce DualPath, a novel inference architecture that adds a second KV-cache loading route directly from storage to the decode engines. Once loaded, the KV cache can be efficiently transferred to prefill engines via high-speed RDMA over the compute network, reducing pressure on the storage network and avoiding congestion with latency-sensitive computation traffic. A global scheduler dynamically balances workload across the prefill and decode engines to optimize utilization and performance. When evaluated on three different models under production-like agentic inference workloads, DualPath significantly improves throughput: up to 1.87× for offline inference scenarios and roughly 1.96× for online serving without violating service-level objectives (SLOs). These results suggest that reorganizing data flows at the system level can meaningfully break storage bandwidth barriers in large-context LLM deployments. https://lnkd.in/gFc7N99w
-
Sanchit Narula
Nielsen • 46K followers
The best advice I got on my promotions as a software engineer was from a Principal Architect at Amazon after he told me his story of getting promoted to Principal, and this has stayed with me my whole career. He said, very simply: > Your architecture can be a love letter to your technical brilliance, > but if it does not move customer and company goals, it will not move your level. That hurt a little. Because I recognised myself in it. A few years ago I wanted to design something with every shiny thing in the book: API Gateway, Lambda, DynamoDB, event bus, the works. On paper it was beautiful. In reality it was expensive, slow to ship, and did not match the skills of the team that would maintain it. So we shipped a hybrid architecture instead. Fewer moving parts, more reuse of existing services, cheaper to run, easier to hand over. Over time, here is how my thinking about promotions changed. 1. Start from the business goal, not the tech stack Before talking about Kafka vs SQS or MySQL vs DynamoDB, answer this first: What money, risk, or customer pain will this remove? Promotions track that, not the number of services in your diagram. 2. Ship finished products, not clever demos Leaders remember the feature that went live, stayed stable, and earned or saved real money. Drive the entire lifecycle yourself: design, build, launch, oncall, cleanup, docs. 3. Make the team faster, not just yourself Tools, scripts, templates, better CI, better dashboards. If people quietly say “things move faster when you are on this” your manager already has half a promotion case. 4. Do your share of the unsexy work Oncall, migrations, bugs, incident reviews. Senior folks who avoid this slow the entire org. Senior folks who lean into it become trusted very quickly. 5. Grow people, not just systems Mentor juniors, run design reviews, write clear docs, set standards. When your manager sees that you are already acting like the next level, formal promotion becomes a catch up, not a favour. 6. Be proactive about your promotion, not obsessed with it Have periodic one to ones. Talk about where you want to go and ask what evidence your manager would need to support that. Remember your manager has ten other people to manage. It is your responsibility to make your impact visible. Do not be desperate. Understand the timelines and work within them. As you get close to the boundary, collect the missing data points: impact metrics, incident write ups, design docs, peer feedback. Make the story easy to tell. Most importantly, work for the business, not for promotion-driven development. If your roadmap only exists to tick boxes on a promotion rubric, people can feel it. If your roadmap clearly creates value, saves cost, or reduces risk, promotion becomes the natural side effect. So instead of asking “What new tech can I use this year?” Ask, “Where can I create obvious business impact and leave the system and the people around me better than I found them?”
182
15 Comments
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content