Qiangjian Xi
San Mateo, California, United States
8K followers
500+ connections
View mutual connections with Qiangjian
Qiangjian can introduce you to 10+ people at Google
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Qiangjian
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Activity
8K followers
-
Qiangjian Xi liked thisQiangjian Xi liked thisSuper happy to share our intention to join forces with NVIDIA in a $12,930,300,000 acquisition 💛💚 10 years after starting Hugging Face, open-source AI is at an inflection point. Thanks to the community, we’ve shown that it can be a complement, and even an alternative, to closed-source APIs. But for it to happen at larger scale, it needs more compute, more support, more collaboration and more visibility. That’s why we went to talk to Jensen, who offered to do exactly that with us. In addition to doubling down on NVIDIA’s massive contributions to open-source AI (I called them the “King of American open-source AI” earlier this year), they’ve committed to strongly supporting Hugging Face and our mission while keeping the platform open, independent and compute agnostic. The founders and the team are all staying to keep pushing this mission forward. Together, we think we can make open source the default way to build AI, with the goal of empowering 100 million AI builders to own their intelligence rather than rent it. Excited about the next 10 years! 💛💚
-
Qiangjian Xi liked thisQiangjian Xi liked thisModern video search systems increasingly rely on multi-vector retrieval to capture visual, audio, and contextual signals. In this technical whitepaper, we break down how these architectures work and how teams deploy them in production.
-
Qiangjian Xi liked thisQiangjian Xi liked thisI’m thrilled to share that I joined 𝗠𝗲𝘁𝗮 𝗦𝘂𝗽𝗲𝗿𝗶𝗻𝘁𝗲𝗹𝗹𝗶𝗴𝗲𝗻𝗰𝗲 𝗟𝗮𝗯𝘀 as an AI Research Scientist two weeks ago. I’m excited to contribute to AI modeling, AI search, and AI personalization — problems that combine cutting-edge research with real impact for billions of users. It’s the scale, billions of users and data you can’t find anywhere else, that makes this such a compelling opportunity. Even in this short time, I've been incredibly impressed by the pace, ambition, and execution in Meta Superintelligence Labs (MSL). What stands out is the commitment: the caliber of people being hired, the scale of infrastructure and compute investment, and the clarity around how AI fits into the long-term strategy. Looking forward to building together! #Meta #MetaSuperintelligenceLabs #AI #LLM
-
Qiangjian Xi liked thisQiangjian Xi liked thisWhat will agentic commerce become in 5 years? And where we all are today? Today we are not limited by AI models. It’s limited by whether our data and systems are ready for real agent reasoning. What about in 5 years? Data and systems can be agent-ready. The open question is how responsibly and reliably we get there. NRF 2026 showed real momentum—agents are moving from recommendations toward decision-making and execution. From an engineering and data perspective, the challenge isn’t that models aren’t powerful enough. It’s that most systems weren’t designed for agents to reason over them. Today, across many industries, agents still depend on: • incomplete or inconsistent product and service data • flat schemas with limited context • unclear constraints around compatibility, policy, or compliance This isn’t a criticism—it’s simply how most commerce systems evolved. Over the next 3–5 years, progress will come from steady, foundational work: • improving data quality and structure • making relationships and constraints explicit • enabling explanations and auditability • designing systems agents can trust—and humans can trust back As models continue to improve and become more accessible, the real differentiator won’t be intelligence alone. It will be readiness: Are our systems prepared to let agents make decisions safely and transparently? At Paladio AI , we’re focused on that foundational layer—helping make product data more structured, compatible, and explainable—so agents can operate reliably in real production environments. Curious how others are approaching agent readiness in their systems—would love to hear what’s working (or not) in the comments. Or let’s grab a coffee in person with our Vamsi Putrevu at #NRF2026 #NRF2026 #AgenticCommerce #AIInfrastructure #DataEngineering #CatalogOps https://lnkd.in/eCiE5xxa
-
Qiangjian Xi liked thisQiangjian Xi liked thisPosting our JD -- Senior Staff Data + AI Engineer We are a fast-paced AI startup delivering end-to-end automation and LLM-driven solutions with scalability, compliance, and observability at the core. Our vision is to transform how companies are run by building the future of AI-powered enterprises. Role Overview We are seeking a Senior Staff Data + AI Engineer to drive the design, implementation, and operation of scalable data and AI infrastructure. This role blends data engineering, ML operations, and platform reliability—enabling our teams to train, deploy, and scale LLMs while ensuring robust pipelines, governance, and observability. You will lead with urgency, ownership, and technical depth, setting the bar for engineering excellence. Responsibilities Data & Infrastructure Architect and maintain distributed data pipelines and warehouses with Spark, Iceberg, Trino, DBT, Airbyte, Snowflake, Databricks. Build fast, reliable ETL/ELT workflows and streaming pipelines (Kafka, Temporal) to power ML. Implement governance, compliance, and security frameworks across data platforms. AI/ML Platform Engineering Design and operate Kubernetes (EKS/AKS), ECS, App Run clusters for high-performance LLM workloads. Optimize inference with vLLM, TensorRT-LLM, Triton, LangChain, LangGraph, LangFuse. Drive automation in CI/CD for training, validation, deployment, and monitoring. Ensure 99.9%+ uptime, cost efficiency, and observability (Prometheus, Grafana, Azure Monitor). Leadership & Collaboration Act fast: take bold ownership of end-to-end delivery, ensuring speed and quality in execution. Owner mindset: proactively solve problems, reduce operational friction, and champion platform adoption. Mentor engineers and data scientists in MLOps, distributed systems, and data best practices. Lead design/code reviews emphasizing security, scalability, and compliance. Partner with product and research teams to deliver ML-driven applications at scale. Qualifications Required 8+ years in software/data engineering with focus on infrastructure, big data, or MLOps. Strong in Python; experience with Go, Rust, or Java is a plus. Hands-on with Azure (AKS, Azure ML, DevOps); multi-cloud (AWS/GCP) experience preferred. Proven track record with LLMs (Llama, Mistral, Gemma, etc.) and inference optimization. Expertise in Kubernetes, ECS/EKS, App Run, Terraform, CI/CD, Docker, distributed systems. Strong knowledge of SQL/NoSQL, Spark, Iceberg, Trino, DBT, Airbyte, Snowflake, Databricks. Demonstrated leadership, mentoring, and cross-functional collaboration skills. Preferred Master’s or PhD in CS, ML, or related field. Experience designing data pipelines with Medallion Architecture. Experience with vector DBs (Milvus, Pinecone, Qdrant) and orchestration tools (Temporal, Airflow, Kubeflow, MLflow). Deep experience in GPU optimization, RLHF, fine-tuning, and hybrid/multi-cloud deployments. Familiarity with AI safety, governance, and regulated industry compliance
-
Qiangjian Xi liked thisQiangjian Xi liked thisFrom brainstorming to technical docs, Lucid helps engineering teams map out systems and ship code faster. Start diagramming today.
-
Qiangjian Xi liked thisQiangjian Xi liked this[Note: I have since left P-1 AI] Phew, I can finally share what I've been up to since last summer! We just raised a $23 million seed round!! 😅 I co-founded P-1 AI w/ Paul Eremenko (ex CTO of Airbus, UTC) and Adam Nagel (ex engineering director at Airbus) with a mission to build an engineering AGI for the built world. Our vision is simple: we want to build an engineering AGI for the real world to help us design airplanes, Dyson Spheres, cars, HVAC systems, etc. Our system is called Archie (my YouTube video coming soon :)). I just opened up our office in San Francisco, just next to the old OpenAI's office. We'll be hosting an office party very soon - stay tuned! A big thank you to our VC Radical Ventures for leading our round (they backed up Fei Fei Li's startup and some of the best AI startups out there) as well as our other investors! Also a huge thank you to our angels Jeff Dean (Google DeepMind), Peter Welinder (VP of Product at OpenAI), and my other friends Bob van Luijt, Luis F Voloch, Waikit L., etc. who believed in us! And last but definitely not least the amazing team behind P-1 AI! We'll be rapidly expanding the size of the team. Our GPU cluster is humming in the background, join while it has free cycles. :P I'm hiring for cracked ML research engineers: https://p-1.ai/ --- I have a lot of stories to share that I couldn't while we wre in stealth, stay tuned for that as well. :)) And yeah, a paper coming out soon as well :) --- Few articles: * Fortune article by the amazing Sharon Goldman! https://lnkd.in/dkAtvzgN thanks Sharon for covering us in an exclusive! * Business Wire: https://lnkd.in/d2M28Ttt
-
Qiangjian Xi liked thisQiangjian Xi liked thisI'm thrilled to share that I've joined Perplexity AI! I am joining the AI Training team, working on exciting LLM projects! As I wrap up my first week, I'm already amazed by the incredible talent and innovation surrounding me. Perplexity AI is at the forefront of AI-driven technology, and I'm excited to work with Aravind Srinivas, Denis Yarats, Johnny Ho, Alex Lang and many more! The potential impact of our work on how people access and interact with information is truly inspiring! Perplexity is actively hiring across various roles (https://lnkd.in/gGekUgku). If you're passionate about AI, please reach out! #Perplexity
-
Qiangjian Xi liked thisQiangjian Xi liked thisDelighted to have graduated from my MS degree in Computer Science from Georgia Institute of Technology with straight As in December. Really enjoyed writing a compiler from scratch, getting my head around quantum algorithms, and AI and human-computer interactions. And thrilled to be joining Wayve today to work on their data platform to help power self driving cars!
Experience
Education
Languages
-
Chinese
Native or bilingual proficiency
-
English
-
View Qiangjian’s full profile
-
See who you know in common
-
Get introduced
-
Contact Qiangjian directly
Other similar profiles
Explore more posts
-
Alan Kochukalam George
Myovine • 617 followers
⚙️ One library, real DSP speed. There’s a gap in the developer ecosystem. Teams doing DSP/ML often migrate to Python, because that’s where most tooling lives. But Python wasn’t built for high-throughput I/O — under load, it spawns new interpreters/processes, creating duplicated memory, slow cold starts, and more containers. Compute stays fast, but orchestration cost explodes. Meanwhile, Node.js/TS scale I/O well, but were never meant for serious compute. Pure-JS DSP hits performance walls. Even worse, serializing data between Redis ↔ DSP/ML adds latency and forces batching — slowing real-time pipelines. So devs are stuck: Python → strong math, weak concurrency Node → strong networking, weak compute dspx closes that gap — TypeScript DX + native C++/SIMD (AVX2/SSE3/SSE2/NEON) under the hood. 🧩 How dspx works FIR / FFT / Conv1D run in optimized C++/SIMD Memory-safe circular buffers → O(1) throughput TS handles Kafka / Redis / WebSockets Redis persistence < 0.5 ms Batched logging avoids I/O stalls At tiny batch sizes, N-API overhead can make naive JS look competitive. At realistic scale, batched native pipelines win. ⚡ Benchmarks Dell OptiPlex 3000 Micro (i5-12600T · AVX2 · Node 22) FFT: 2.6× faster than fft.js, ~9× faster than tfjs-node FIR (51-tap): ~4× faster than fili / naive JS Conv1D (128-kernel): ~3.5× faster than tfjs-node Moving Avg (O(1)): 559× more throughput/sec than naive JS Redis save/load: sub-ms Logging: <3% overhead 📊 Full tables + charts in carousel 💬 Observations SIMD FMA + contiguous memory → big FIR/Conv wins O(1) moving-average design → massive throughput gain Sub-ms Redis + low I/O overhead → real-time persistence in pure Node 🧠 Open Source dspx is Apache 2.0, free for commercial/academic use. Looking for contributors interested in: ARM NEON / Apple M / Graviton tuning Audio / sensor / biomedical DSP Visualization + benchmark tooling 📦 npm i dspx 🔗 https://lnkd.in/e-tWxAgu 🧠 Real-time DSP for Node, TypeScript & Redis 💼 Note I’m exploring opportunities in full-stack, real-time systems, performance engineering, and DSP. If your team works in this space, I’d love to connect. My next post will cover why sub-ms Redis latency isn’t just technical — it’s economic. Lower serialization + compute overhead → lower infra + energy cost. 🔖 Tags & Mentions #NodeJS #TypeScript #DSP #PerformanceEngineering #EdgeComputing #OpenSource #SIMD #Redis #Cplusplus #RealTime NodeJS Developer TensorFlow Google Microsoft JavaScript Developer Amazon Web Services (AWS) Vercel
5
1 Comment -
Byung-Kwan Lee
NVIDIA • 11K followers
What does comfortable, confident, safe autonomy feel like? Ride through San Francisco with NVIDIA founder and CEO Jensen Huang and NVIDIA VP of Automotive Xinzhou Wu as they discuss the technology powering the next generation of autonomous vehicles. ▶️ Watch now: https://bit.ly/4rz72M1
70
-
DoKyun-DK Kim
Samsung Electronics America • 431 followers
[ The Rubin Paradox: A 31% Drop in FP16 Efficiency?] NVIDIA's Rubin architecture undoubtedly pushes the raw performance envelope: * 2.75x increase for FP16 FLOPS * 5x increase for FP8 FLOPS However, this comes at a steep cost: * Power consumption has surged to 2,300W. * This is a massive increase, especially given that TDP remained constant at 700W from the H100 to the B100. The result? A critical shift in performance per Watt. The FLOPS/W metrics have significantly declined. In fact, FP16 FLOPS/W is down by a substantial 31% compared to the Blackwell (B100). It's clear why the GTC narrative steered toward Tokens/sec/W and Tokens/sec/User rather than absolute performance. This focus on inference efficiency over pure computational density is a telltale sign. This leads to a compelling question: Is planar process scaling finally at its limits? Perhaps this explains the push for new breakthroughs like: * Advanced Packaging - Die Stacking (Feynman) * Disaggregated Architectures (+Groq) ________________ H100 B100 Rubin GPU TDP[W] 700 700 2300 FP16[TFLOPS] 989 1750 4000 FP8[TFLOPS] 1979 3500 17500 #HBM4 #DieStacking #Feynman #DisaggregatedArchitecture #Dataflow #Groq #PowerEfficiency #TDP
6
-
Ivan Nardini
Google • 30K followers
Serving LLMs on TPUs with SGLang ! I've been exploring the JAX ecosystem for inference recently and came across SGL-JAX, a high-performance, JAX-based inference engine for LLMs, specifically optimized for Google TPUs. Some highlights: 🌳 The "Radix Tree" KV Cache: SGL-JAX implements a Radix Tree (similar to PagedAttention) to manage memory. This enables efficient prefix sharing for multi-turn chat or agentic workloads where the system prompt remains constant. ⚡ JAX-Native FlashAttention: It integrates a high-performance FlashAttention kernel directly into the JAX compute graph. This is critical for faster, more memory-efficient attention usage, particularly when dealing with long sequence lengths on TPUs. 🧩 Native Tensor Parallelism: It handles the sharding logic natively using JAX distributed primitives. It supports distributing large models (like the Qwen 3 MoE series) across multiple TPU devices without needing complex, manual mesh definitions. ⏳ Continuous Batching: It dynamically schedules incoming requests to maximize TPU core saturation. You aren't waiting for a fixed batch size to fill up while tail latency spikes. 🔧 Drop-in Compatibility: It exposes an OpenAI-compatible API standard. You can literally point your existing LangChain or LlamaIndex setup at the new endpoint and switch hardware backends seamlessly. For more info about the architecture and how it works, check out the repo below! #JAX #MachineLearning #TPU #LLMOps #GoogleCloud #AIEngineering
46
1 Comment -
Anyscale
63K followers
In this week’s AI infra chat, we sit down with Seiji Eicher, Distributed LLM Inference Engineer at Anyscale to unpack how large MoE models are actually served in practice, what it looks like to work deep in the LLM inference stack & more 🕐 Timestamps: 𝟬𝟬:𝟬𝟬 Why MoE comes up in modern AI infra 𝟬𝟬:𝟮𝟯 What Mixture of Experts actually changes inside a model 𝟬𝟭:𝟮𝟰 Why expert parallelism exists (and what becomes inefficient without it) 𝟬𝟮:𝟱𝟳 When MoE helps in practice 𝟬𝟰:𝟮𝟯 How Ray Serve fits into real inference stacks 𝟬𝟳:𝟬𝟳 What a Distributed LLM Inference Engineer works on day to day 𝟭𝟭:𝟭𝟵 Where the space is headed and how to learn more If you want to go deeper, Seiji gave a Ray Summit talk that dives into production-grade serving of MoE models, including KV-cache optimization, parallelism strategies, and multi-node orchestration with Ray Serve and vLLM. Check out the talk linked in the comments 👇
60
3 Comments -
Ramin Mehran
Google DeepMind • 4K followers
In this episode, we discuss ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models by Mingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu, Xin Dong, Yejin Choi, Jan Kautz, Yi Dong. This paper introduces ProRL, a new reinforcement learning training method that uncovers novel reasoning strategies beyond those found in base language models. Empirical results show that models trained with ProRL consistently outperform base models on challenging reasoning tasks, including cases where base models fail even with extensive attempts. The study demonstrates that prolonged RL can meaningfully expand reasoning capabilities by exploring new solution spaces over time, advancing understanding of how RL enhances language model reasoning.
2
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content