We just took #1 on all 3 MiniMax M3 inference leaderboards. 🏆 Artificial Analysis benchmarks every major provider serving the model across 3 boards: fastest, lowest latency, and lowest price. CoreWeave Serverless Inference now holds #1 on every one of them for this model. Output speed comes in at 357 tokens per second, 1.8x the next provider. Time to first answer token is 6.6 seconds, where the runner-up takes 12. Our $0.22 per million blended price matches the lowest on the board, so the speed carries no premium. MiniMax M3 ships with tool calling, vision input, JSON mode, and LoRA support, and it runs on the same infrastructure nine of the 10 leading foundation model providers rely on. See the data for yourself: https://lnkd.in/gKMqSTiq Try MiniMax M3 here: https://www.utm.io/urNL8
Weights & Biases
Software Development
San Francisco, California 93,698 followers
The AI developer platform.
About us
Weights & Biases: the AI developer platform. Build better models faster, fine-tune LLMs, develop GenAI applications with confidence, all in one system of record developers are excited to use. W&B Models is the MLOps solution used by foundation model builders and enterprises who are training, fine-tuning, and deploying models into production. W&B Weave is the LLMOps solution for software developers who want a lightweight but powerful toolset to help them track and evaluate LLM applications. Weights & Biases is trusted by over a 1,000 companies to productionize AI at scale including teams at OpenAI, Meta, NVIDIA, Cohere, Toyota, Square, Salesforce, and Microsoft. Sign up for a 30-day free trial today at http://wandb.me/trial.
- Website
-
https://wandb.ai/site
External link for Weights & Biases
- Industry
- Software Development
- Company size
- 201-500 employees
- Headquarters
- San Francisco, California
- Type
- Privately Held
- Founded
- 2017
- Specialties
- deep learning, developer tools, machine learning, MLOps, GenAI, LLMOps, large language models, llms, Generative AI, Developer Tools, Experiment Tracking, AI Governance, Model Monitoring, Inference, Open Source AI, Model Comparison, Evals & Scorers, Data Quality, Generative AI, AI Observability, Agentic Workflows, RAG (Retrieval-Augmented Generation), Prompt Engineering, Hyperparameter Tuning, Benchmarking, Large Language Models (LLMs), Reproducibility, Dataset Versioning, and Tracing
Products
Weights & Biases
Machine Learning Software
Weights & Biases helps AI developers build better models faster. Quickly track experiments, version and iterate on datasets, evaluate model performance, reproduce models, and manage your ML workflows end-to-end.
Locations
-
Primary
Get directions
400 Alabama St
San Francisco, California 94110, US
Employees at Weights & Biases
Updates
-
The numbers couldn't tell the one success from 39 failures. We ran DreamZero, NVIDIA GEAR Lab's 14B World Action Model (WAM), zero-shot across 65 rollouts in simulation on scenes it never trained on. A WAM works differently from a standard Vision-Language-Action model. Instead of going straight from camera frame to action, it first generates a few seconds of predicted video, a "dream," then acts on what it expects to see. The useful part for evals is that the dream is a real artifact you can score. You line up what the model predicted against what the simulator rendered, frame by frame. Two of three tasks held up. Cube into a bowl and banana into a bin both hit 100% success zero-shot. When we swept the cube task with position noise, the model stayed at 100% out to 10 cm off center. The can-into-a-mug task managed 1 success in 40 attempts, and rephrasing the goal never helped. The arm got there but missed on final placement. On that hard task, the scalar scorers all landed in the same range for the single success as for the failures. The numbers couldn't separate them. Watching the dream against the actual rollout was what showed us where it broke. Physics sim, 14B inference, and video encoding all ran in parallel on one CoreWeave cluster of RTX Pro 6000 Blackwell GPUs, tracked end to end with W&B Weave. Anu Vatsa's full write-up covers where the model breaks, plus an open question worth chasing. If the model's own prediction noise spikes before it fails, that could be a usable runtime warning signal. The repo ships every manifest. Link in the comments.
-
-
Weights & Biases reposted this
Why not build it yourself? At #FullyConnected26, dive into hands-on labs that put you inside real CoreWeave + Weights & Biases infrastructure. Across three days, you'll: ✅ Build an autoresearch agent that proposes, runs, and improves its own experiments ✅ Compare Serverless, Dedicated, and self-hosted CKS Inference tiers on real latency and cost data ✅ Profile a live training run on SUNK and find your MFU gap before it costs you ✅ Fine-tune a protein model and ship it to the W&B Model Registry with full lineage ✅ Teach a robot a new skill in simulation, then compete for the top of the leaderboard ✅ Run a live AI governance review, red-team an agent, then approve, change, or block it You leave with a working build, not a slide deck. 🗓️ September 29–October 1 📌 Moscone South, San Francisco 🔗 Register now: https://www.utm.io/upDdF
-
One click turned a confusing curve into a straight line. Vincent kept adding cities to his traveling salesman runs and watching the runtime climb. What he couldn't tell from the raw chart was the actual relationship. A YouTube commenter pointed him to the log scale button in Weights & Biases. Flip it, and the near-straight line on the log axis says it plainly: the approach scales exponentially. His words, an awesome feature hidden in plain sight.
-
A robot policy is only as reliable as the data it learned from. We're teaming up with Encord for a live session on fine-tuning World Action Models (WAMs), following the full loop from raw multimodal sensor data to evaluated robot behavior. Skander Fourati, ML Solutions Engineer at Encord, and Anushrav V., Sr. AI Solutions Engineer here at Weights & Biases, will get into: - Curating, annotating, and versioning multimodal datasets across vision, language, action, and sensor streams - Tracing failure modes back to the data underneath - Building a data flywheel so every experiment makes the model better - Running Encord and W&B together as one reproducible workflow 📆 Tuesday, August 4th 🕘 9am PT Live demos and open Q&A, and registrants get the recording either way. Save your seat: https://lnkd.in/g4fUvGCw
-
-
Weights & Biases reposted this
GPT-5.6 Sol took the top three spots on WolfBench: Codex: 86.74% Terminus-2: 85.17% Hermes: 84.49% So far, so benchmarky. Then I checked the tokens. With the same Sol model at maximum reasoning, Codex used 83.8 million tokens per run. Hermes used 170.9 million. Twice the tokens for a 2.25 percentage-point lower score. Hermes recorded $173.41 per run. Codex comes out at an estimated ~$89.79-$95.06. (That Codex number is reconstructed from token usage, not billed spend, but even its upper bound is 45% lower.) The GPT-5.5 comparison makes it even weirder: With Codex, GPT-5.6 gained 8.09 points, used 20% fewer tokens, and appears roughly half as expensive. With Terminus-2, it gained 8.31 points while cost fell 26.5%. With Hermes, it gained 10.34 points - but used 68% more tokens and cost 47% more. So is GPT-5.6 cheaper than GPT-5.5? Depends entirely on the agent. One last oddity: Across 15 Sol runs, the "configure git webserver" task was never solved. Terra solved it once. Luna solved it every time. The smaller variants are not simply weaker copies of Sol. My takeaway: Codex + Sol is the strongest full-agent stack we tested. Terminus-2 + Sol is the cleanest reference baseline. And for Hermes, I would seriously consider Terra: only 3.6 points behind Sol, but 44% cheaper. Benchmark the stack, not just the model! And run it more than once. Full results on WolfBench.ai:
-
-
Weights & Biases reposted this
Y'all should checkout the Bee Arena I just made. Weights & Biases Bee inside of a CoreWeave Arena. Testing Kimi K3 vs GPT 5.6 Sol. My one complaint with Kimi K3: WHO USES INVERTED CONTROLS ON GAMES! IF YOU DO THEN YOU HAVE THE WRONG OPINION. IDC SINCE I WAS A CHILD I HATED IT! Y'ALL CAN FIGHT ME ON THIS!!!!!!!!
-
-
We ran the same phishing email through two AI agents. The email had a prompt injection buried in the body and a customer's SSN and credit card number sitting in the text. -> Agent one, no guardrails, followed the planted instruction and echoed the SSN and card number straight back into its reply. -> Agent two, running CrowdStrike Falcon AI Detection and Response (AIDR), blocked the injection before the model ever read it, redacted 4 sensitive items (US_SSN and CREDIT_CARD), and still returned a usable draft. Same input. The gap between those two outcomes is the whole point. Underneath both, W&B Weave traced the runs end to end, so every scan, detection, and tool call was recorded and queryable. When something fires in production, a blocked request shows up in your security team's Slack, linked to the exact trace behind it. That is what it looks like to run internal AI on sensitive data without flying blind. Full walkthrough, with the code and both side-by-side runs, in the comments.
-
-
The talks were great. The hallway conversations were better. 🐝 We're still buzzing from AI Engineer's World Fair in SF. Four days at Moscone, a booth that stayed busy start to finish, and a steady stream of people who wanted to talk about models, evals, agents, and where all of this is actually heading. Alex Volkov and Wolfram Ravenwolf took the stage, and the conversations that spilled out afterward were the real highlight. CoreWeave ARIA was one of the things people kept bringing up. Watching folks get genuinely excited about an AI research and iteration agent that could do autoresearch made the whole week. Thank you to everyone who stopped by, said hi, and stuck around to talk shop. This community is a good one. We will definitely be back next year!