Very ML
State-of-the-art Machine Learning News Feed
/r/MachineLearning
последний пост 6 часов назад
Real-Time Conversational Agents (RTCA) Workshop @ NeurIPS 2026 — submissions now open, deadline Aug 29 AoE [N]
Real-Time Conversational Agents (RTCA) Workshop @ NeurIPS 2026 — submissions now open, deadline Aug 29 AoE [N]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

6 часов назад @ reddit.com
Good OCR strategy for detecting doctor handwritting [P]
Good OCR strategy for detecting doctor handwritting [P] Good OCR strategy for detecting doctor handwritting [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

21 час назад @ reddit.com
What is currently considered the theoretically optimal quantization bit-width for LLMs? [D]
What is currently considered the theoretically optimal quantization bit-width for LLMs? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

22 часа назад @ reddit.com
2026 NeurIPS: Where are you going? [D]
2026 NeurIPS: Where are you going? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

22 часа назад @ reddit.com
​Built a tool to generate slides from research papers using local LLMs (because I hate formatting decks and privacy matters) [P]
​Built a tool to generate slides from research papers using local LLMs (because I hate formatting decks and privacy matters) [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 1 hour назад @ reddit.com
Imagenet-1k Classifier trained entirely on an Android [P]
Imagenet-1k Classifier trained entirely on an Android [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 4 hours назад @ reddit.com
Improved compression of Bad Apple into a Neural Network [P]
Improved compression of Bad Apple into a Neural Network [P] Improved compression of Bad Apple into a Neural Network [P]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 6 hours назад @ reddit.com
CIKM 2026 decisions [R]
CIKM 2026 decisions [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 7 hours назад @ reddit.com
On the ACM Multimedia 2026 Conference Registration and APC [D]
On the ACM Multimedia 2026 Conference Registration and APC [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 7 hours назад @ reddit.com
Which degree is best? [D]
Which degree is best? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 9 hours назад @ reddit.com
CIKM '26 Notification [D]
CIKM '26 Notification [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 13 hours назад @ reddit.com
NeurIPS Meta Reviewer comment gone. What gives? [R]
NeurIPS Meta Reviewer comment gone. What gives? [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 21 hours назад @ reddit.com
Can recurring LLM traces be synthesized into deterministic pipelines of typed ML and NLP operators? [D]
Can recurring LLM traces be synthesized into deterministic pipelines of typed ML and NLP operators? [D]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

1 day, 21 hours назад @ reddit.com
The current state of language models and human preference based rankings [R]
The current state of language models and human preference based rankings [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 1 hour назад @ reddit.com
Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors [R]
Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors [R] Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors [R]

Your request has been blocked due to a network policy.

If you're running a script or application, please register or sign in with your developer credentials here.

Additionally make sure your User-Agent is not empty and is something unique and descriptive and try again.

if you're supplying an alternate User-Agent string, try changing back to default as that can sometimes result in a block.

If you think that we've incorrectly blocked you or you would like to discuss easier ways to get the data you want, please file a ticket here.

2 days, 2 hours назад @ reddit.com
Towards Data Science
последний пост 10 минут назад
Before Q, K, and V: Reconstructing the Transformer
Before Q, K, and V: Reconstructing the Transformer Before Q, K, and V: Reconstructing the Transformer

We want to break this symmetry so let’s keep only the middle interaction function v, which I’ll call the “attention” function from now on.

(One caveat is that the Transformer architecture adds scaling for computational stability, hence the term “scaled dot product attention”.

Okay, all of this is great—but where are the matrices Q, K, and V that the article title promised us?

The product between q and K^T creates a vector containing every dot product between q and a key in K, and the softmax on top normalizes the final dot product scores.

We needed to reduce our cache size for the reusable matrix-vector multiplies in the dot product, which required computing the dot product in lower dimensi…

10 минут назад @ towardsdatascience.com
Building a Streamlit UI for My LangGraph AI Agent
Building a Streamlit UI for My LangGraph AI Agent Building a Streamlit UI for My LangGraph AI Agent

In this article, we will build a clean, interactive Streamlit UI on top of the existing LangGraph agent.

In LangGraph, the agent architecture is defined and executed as a compiled state graph object so they essentially mean the same thing in this article.

Both serve as a wrapper for the LangGraph agent.

This function sends the customer action to the LangGraph agent and saves the resulting state for the Streamlit interface.

Finally, we have some render functions ( _render...() ) defined in streamlit_app.py to convert the current LangGraph state into visible Streamlit components.

2 часа назад @ towardsdatascience.com
Matplotlib vs Plotly: Which Python Chart Tool Should You Choose?
Matplotlib vs Plotly: Which Python Chart Tool Should You Choose? Matplotlib vs Plotly: Which Python Chart Tool Should You Choose?

In this article, I’ll provide several examples of using Plotly and Matplotlib on the same datasets to illustrate their key differences.

Now, let’s create the same plot using Plotly Express, which provides a high-level interface similar to Seaborn.

Although the code complexity is comparable, the user experience when using Plotly is greatly enhanced.

Example 3: Saving and SharingHow you save and share plots using the two tools differs significantly in one key aspect.

With the high-level plotly.express module, creating these interactive plots often requires minimal code changes compared to Matplotlib/Seaborn, while providing a much richer user experience.

22 часа назад @ towardsdatascience.com
Loop Engineering for Listing Questions: When the Answer Is Every Passage, Not the Top One
Loop Engineering for Listing Questions: When the Answer Is Every Passage, Not the Top One Loop Engineering for Listing Questions: When the Answer Is Every Passage, Not the Top One

Listing questions break the one assumption retrieval is built on, that the answer is the top passage.

For a factual question, top-k retrieval works because the answer is in one passage.

For a listing question, the answer is in N passages, where N is unknown ahead of time.

1.3 Detecting the listing intentThe first thing the pipeline needs is to recognize that the question is a listing question, not a factual one.

Sources and further readingThe benchmark showing top-k retrieval ceiling far below 100% on list questions is Amouyal et al.

1 day назад @ towardsdatascience.com
The Problem with pandas Isn’t Performance. It’s Cognitive Overhead.
The Problem with pandas Isn’t Performance. It’s Cognitive Overhead. The Problem with pandas Isn’t Performance. It’s Cognitive Overhead.

AI agents can generate the pandas code for us now, so what does it matter?

Let’s be honest: reading other people’s pandas code, especially code that was built through an interactive session can be rather painful.

The enduring popularity of visual data toolsIf you’re still not convinced that any of this matters, think for a minute about the popularity of visual data tools.

But if “code-as-documentation” is kind of the point, why are we using boilerplate-laden pandas code as our language of choice?

Since then, the Python data ecosystem has become entrenched across academia, industry and government IT environments.

1 day, 1 hour назад @ towardsdatascience.com
My Fall-Detection Model Scored 94%, and It Was Lying to Me
My Fall-Detection Model Scored 94%, and It Was Lying to Me My Fall-Detection Model Scored 94%, and It Was Lying to Me

I extracted features per frame, labelled per frame, ran a standard train_test_split, and got 94.3% accuracy.

For an alarm system, the questions that matter are simpler: when someone falls, does the alarm fire?

My most sensitive model treats lying on the floor after a fall as positive.

Model Falls detected False-alarm videos Deliberate lie-down “person down” labels 96.0% 7/60 alarms “fall motion” labels 89.7% 3/60 silentNeither model is wrong.

If the goal is to show the system can tell a fall from a lie-down, the motion model wins.

1 day, 3 hours назад @ towardsdatascience.com
I Built an AI Data Agent Which Can Query Data and Answer Business Questions. Here’s How.
I Built an AI Data Agent Which Can Query Data and Answer Business Questions. Here’s How. I Built an AI Data Agent Which Can Query Data and Answer Business Questions. Here’s How.

Screenshot of the data agent interface built by authorWhat Is a Data Agent?

A data agent is an AI-powered conversational interface that enables business users to ask questions in plain language and receive accurate answers by querying data stored in a data warehouse.

For beginners, the second approach of deploying a data agent within a cloud data platforms is more practical and faster to implement.

For any price calculations, always use the weighted average formula (SUM(Total Volume * AveragePrice) / SUM(Total Volume)) when aggregating across multiple records.

text_type: THOUGHT } } timestamp { seconds: 1785482363 nanos: 449760000 } system_message { text { parts: "Answering the \"Albany 201…

1 day, 22 hours назад @ towardsdatascience.com
Last Month’s Machine Learning Lessons Learned
Last Month’s Machine Learning Lessons Learned Last Month’s Machine Learning Lessons Learned

The reason was conference travel.

To put things into perspective for readers unfamiliar with the conference game in machine learning (ML) research, here is the story, slightly abridged.

If you are based in China and the conference takes place in South Korea, you can get there relatively quickly.

And that, to me, is the real downside of conference travel.

This month’s lesson was therefore rather simple: when planning conference travel, do not count only the days written in the conference program.

2 days назад @ towardsdatascience.com
I Built a Tool-Calling Agent in Python. Here’s How I Debugged It
I Built a Tool-Calling Agent in Python. Here’s How I Debugged It I Built a Tool-Calling Agent in Python. Here’s How I Debugged It

The useful review points are the model request, Python validation, Python execution, compact result shaping, final answer, and trace record.

What a tool calling agent actually doesA tool calling agent in Python is a loop driven by a large language model, or LLM.

The agent loop is your Python code that keeps the process running until the model stops requesting tools.

Why start without a frameworkIt is worth building one custom Python loop before adopting a larger agent framework.

You should see the model request geocode_city , Python run the Nominatim lookup, the model request get_weather , Python run the Open-Meteo forecast call, and the model write an answer from the compact weather payloa…

2 days, 1 hour назад @ towardsdatascience.com
Loop Engineering for Cross-References: When RAG Answers ‘see Section 7.2’ Instead of the Actual Answer
Loop Engineering for Cross-References: When RAG Answers ‘see Section 7.2’ Instead of the Actual Answer Loop Engineering for Cross-References: When RAG Answers ‘see Section 7.2’ Instead of the Actual Answer

With a top-1 retrieval policy (defended in section 6), the first pass fetches page 6 and the table is missed.

class ReferenceAwareAnswer(BaseModel): answer: str # the prose answer citations: list[Citation] # line-level provenance # Reference-loop feedback fields (the contribution of this article).

Loop count is 0, the budget allows 1, the pending_references list has one entry pointing at “Table 3 row (E)”.

Second-pass fetch; row (E) carries the perplexity and BLEU numbers completing the answer – Image by authorThe second-pass answer:{ "answer": "The Transformer paper compares learned positional embeddings against the sinusoidal positional encodings used in the base model.

Variations: the fo…

2 days, 3 hours назад @ towardsdatascience.com
How a Frontier Model Gets Built, Read from the Kimi K3 Report
How a Frontier Model Gets Built, Read from the Kimi K3 Report How a Frontier Model Gets Built, Read from the Kimi K3 Report

K3 is a 2.8-trillion-parameter mixture-of-experts model, and what lifts it over the last Kimi is three fairly ordinary engineering changes stacked together.

That’s how the model can hold 2.8 trillion parameters and still only run 104 billion for any given token.

Across training the model meets many of these arrangements, so it generalises across scaffolds instead of memorising one.

The report frames that systems work as elite human effort, and the model is already doing a fair chunk of it.

Making a capable model cheap to run is something the frontier labs treat as a headline result.

2 days, 22 hours назад @ towardsdatascience.com
Introduction to Semi-Supervised Learning
Introduction to Semi-Supervised Learning Introduction to Semi-Supervised Learning

This family of Machine Learning algorithms is called Supervised Learning.

This family of Machine Learning algorithms is called Semi-Supervised Learning.

In light of that, the authors of this survey paper [2] suggest that we should see Semi-Supervised Learning as another algorithm in the researcher’s toolkit.

There are indeed cases where the introduction of unlabelled data and the application of Semi-Supervised Learning improved prediction performance [2, 3] but there’s not a silver bullet.

Hope you enjoyed learning a bit more about Semi-Supervised Learning, the approaches taken with different algorithms and the limitations of using unlabelled data.

3 days назад @ towardsdatascience.com
Is This Slop? Detecting AI-Generated Content Without a Model
Is This Slop? Detecting AI-Generated Content Without a Model Is This Slop? Detecting AI-Generated Content Without a Model

S at detecting LLM content.

Even confident readers frequently misclassify both AI generated content as human written and human written content as AI generated.

The objective function is:L SFT ( θ ) = − 𝔼 ( x , y ) ∼ D demo [ ∑ t = 1 T log ⁡ π θ ( y t | x , y < t ) ] L_{\mathrm{SFT}}(\theta) = -\mathbb{E}_{(x,y)\sim D_{\mathrm{demo}}} \left[ \sum_{t=1}^{T} \log \pi_{\theta}(y_t \mid x, y_{<t}) \right]where D d e m o D_{demo} ​ is the demonstration dataset.

If D d e m o D_{demo} contained large numbers of example dialogues involving fictional characters, annotator habit or usage could have had a measurable effect on the probability mass for names.

Ability of AI detection tools and humans to a…

3 days, 1 hour назад @ towardsdatascience.com
Building Document Structure with Loop Engineering: Recovering a PDF’s Outline from Body Typography for RAG
Building Document Structure with Loop Engineering: Recovering a PDF’s Outline from Body Typography for RAG Building Document Structure with Loop Engineering: Recovering a PDF’s Outline from Body Typography for RAG

The PDF has a partial native outline that stops at level 2 while the body clearly has a level 3 ( 3.2.1 ... , 3.2.2 ... ).

Section numbering re-inits mid-file, and the native outline (if any) is flat.

3.5 Extending a partial native outlineSome PDFs have a native outline that stops at level 2.

The reader can see, row by row, which entries came from the file and which came from the body pass.

Take six PDFs that DO carry a native outline, hide it, run the body-typography loop on the body, and compare the reconstruction to the outline the file was hiding.

3 days, 3 hours назад @ towardsdatascience.com
The Medallion Data Architecture: An Introduction
The Medallion Data Architecture: An Introduction The Medallion Data Architecture: An Introduction

It divides a data platform into three layers, usually called bronze, silver, and gold.

The bronze, silver and gold terminology was first proposed by Databricks.

Databricks describes the medallion structure as a multi-layered pattern in which data quality improves progressively as data moves through the three layers.

The medallion architecture works because it makes distinctions in your data visible.

Data received isn’t the same as data validated, and data validated isn’t automatically ready for a particular business decision.

3 days, 22 hours назад @ towardsdatascience.com
Distill.pub Distill.pub
последний пост None
TheSequence TheSequence
последний пост 2 days, 4 hours назад
The Sequence Opinion #909: Return on Token: The New Economics of AI-Native Engineering
The Sequence Opinion #909: Return on Token: The New Economics of AI-Native Engineering The Sequence Opinion #909: Return on Token: The New Economics of AI-Native Engineering

For most of software history, engineering capacity was easy to sketch on a whiteboard.

The modern engineering organization now has a second, elastic workforce.

The human workforce is measured in headcount.

The machine workforce is measured, imperfectly, in tokens.

And because companies love measurable things—especially things that produce dashboards—we have entered the era of token maxing.

2 days, 4 hours назад @ thesequence.substack.com
The Sequence AI of the Week #908: You Need to Learn About Gemini Robotics
The Sequence AI of the Week #908: You Need to Learn About Gemini Robotics The Sequence AI of the Week #908: You Need to Learn About Gemini Robotics

You ask Apptronik’s Apollo 2 to put the watering can into the green bin on the bottom shelf.

It walks to the table, picks up the can, takes a few steps to the shelves, and places it where you asked.

Nothing in that sentence sounds hard until you remember what the previous Gemini Robotics models actually were.

They were a torso bolted to a fixed base doing tabletop work.

That is the real news in Gemini Robotics 2, and it is a bigger deal than the b-roll makes it look.

3 days, 4 hours назад @ thesequence.substack.com
The Sequence Knowlege #907: The Brain Transplant: Distilling Transformers Into Other Architectures
The Sequence Knowlege #907: The Brain Transplant: Distilling Transformers Into Other Architectures The Sequence Knowlege #907: The Brain Transplant: Distilling Transformers Into Other Architectures

Every form of distillation in this series so far has quietly preserved one thing: teacher and student spoke the same dialect.

The student was a compressed copy, then a more capable apprentice, then a reasoner trained on traces — but underneath, it was always the same kind of machine, attention layers stacked on attention layers, differing only in size.

You take a fully trained transformer and pour its capability into a fundamentally different computational substrate, and somehow the capability survives the transplant.

The first time you see it work, it feels a little illicit, like recovering a person’s memories after swapping out their brain for different hardware.

So it’s worth understandi…

4 days, 4 hours назад @ thesequence.substack.com
The Sequence Radar #906: Last Week in AI: Open Models, Intelligent Robots, and the Price of Conviction
The Sequence Radar #906: Last Week in AI: Open Models, Intelligent Robots, and the Price of Conviction The Sequence Radar #906: Last Week in AI: Open Models, Intelligent Robots, and the Price of Conviction

The AI of the week dives into Gemini Robotics 2.

Google DeepMind pushed the frontier in a different direction with Gemini Robotics 2.

AI Lab: MistralAISummary: This paper presents Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that simplifies content moderation into a unified binary question-answering task.

🤖 AI Tech ReleasesGemini Robotics 2Google DeepMind released Gemini Robotics 2 , a three-model suite of intelligence for robotics.

LFM2.5Liquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, two encoder models that can be easily adaped to downstream tasks.

6 days, 4 hours назад @ thesequence.substack.com
The Sequence Robotics #905: Who Builds the Robot Brain?
The Sequence Robotics #905: Who Builds the Robot Brain? The Sequence Robotics #905: Who Builds the Robot Brain?

This is the first post of a new section of TheSequence focused on advancements in robotics.

Our goal is to keep you up to date with the most important developments in AI robotics which is an area that is not well covered by other newsletters.

For this first post, I wanted to discuss the current landscape of AI models for robotics.

Robotics is what happens when an AI model leaves the library and discovers physics.

This is why the race for the robot foundation model will not simply replay the LLM market.

1 week, 1 day назад @ thesequence.substack.com
TheSequence Opinion #904: The Age of Research Is Overrated. AI Engineering Is Winning
TheSequence Opinion #904: The Age of Research Is Overrated. AI Engineering Is Winning TheSequence Opinion #904: The Age of Research Is Overrated. AI Engineering Is Winning

From roughly 2012 to 2020, he argued, the field lived in an age of research.

From 2020 to 2025, it entered an age of scaling.

Open the release notes for almost any frontier model in 2026 and the architecture diagram looks strangely familiar.

The headline improvements are usually elsewhere: better data, longer context, stronger reinforcement learning, synthetic tasks, tool use, memory, verification, adaptive reasoning budgets, and agent orchestration.

The age of research has returned, but much of that research is now expressed as industrial-scale engineering.

1 week, 2 days назад @ thesequence.substack.com
The Sequence AI of the Week #903: Laguna, the 118 Billion Parameters that Walks Into a Trillion-Parameter Bar
The Sequence AI of the Week #903: Laguna, the 118 Billion Parameters that Walks Into a Trillion-Parameter Bar The Sequence AI of the Week #903: Laguna, the 118 Billion Parameters that Walks Into a Trillion-Parameter Bar

Take every open-weight model that discloses its parameter count, put total parameters on a log x-axis, put Terminal-Bench 2.1 score on the y-axis, and you get a reasonably tidy cloud sloping up and to the right.

Disclosed-size open-weight models on Terminal-Bench 2.1.

Laguna S 2.1 scores 70.2%.

On DeepSWE, a harder and less saturated benchmark, the gap stops being subtle at all: Laguna S 2.1 scores 40.4 against DeepSeek-V4-Pro-Max’s 9.0.

A 13x parameter deficit paired with a 4x score advantage is the kind of result that usually means somebody broke the eval.

1 week, 3 days назад @ thesequence.substack.com
The Sequence Knowledge #902: Learning About Distillation: When the Dataset Becomes the Teacher
The Sequence Knowledge #902: Learning About Distillation: When the Dataset Becomes the Teacher The Sequence Knowledge #902: Learning About Distillation: When the Dataset Becomes the Teacher

For most of machine learning history, data was treated as geology.

It already existed somewhere in the world—in books, websites, code repositories, conversations, photographs, and databases.

The researcher’s job was to excavate it, clean it, tokenize it, and feed it into a model.

The dataset trains the student.

The teacher disappears at inference time, but some of its behavior remains embedded in the student.

1 week, 4 days назад @ thesequence.substack.com
The Sequence Radar #901: Last Week in AI: Smarter Models, Physical Machines, and the Expanding AI Stack
The Sequence Radar #901: Last Week in AI: Smarter Models, Physical Machines, and the Expanding AI Stack The Sequence Radar #901: Last Week in AI: Smarter Models, Physical Machines, and the Expanding AI Stack

Next Week in The Sequence:Our series about AI model distillation continues with another exciting technique.

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: Last Week in AI: Smarter Models, Physical Machines, and the Expanding AI StackWhen I started The Sequence years ago, AI was still a relatively niche field, followed closely by researchers, a small group of builders, and a few overly enthusiastic people like me.

AI Lab: Meta AISummary: This paper introduces GAMUT, a multimodal benchmark designed to evaluate the factual completeness of long-form generations rather than just their factual precision.

AI Lab: Microsoft ResearchSummary: This paper presents Experiential Learning (EL)…

1 week, 6 days назад @ thesequence.substack.com
The Sequence Opinion #900: Beyond the GPU: Is Google the Only Full-Stack Rival to NVIDIA?
The Sequence Opinion #900: Beyond the GPU: Is Google the Only Full-Stack Rival to NVIDIA? The Sequence Opinion #900: Beyond the GPU: Is Google the Only Full-Stack Rival to NVIDIA?

THESIS Google is the closest strategic mirror of NVIDIA’s full-stack method, but not a universal drop-in replacement; AMD and AWS make a literal “only” claim too strong.

This is useful, but incomplete in the same way that comparing airlines by engine thrust is incomplete.

It is an industrial system that turns models into running software with unusually little friction.

That is the strongest form of the case for Google as NVIDIA’s only viable competitor.

Google is the closest full-stack strategic rival, not a universal drop-in replacement, and AWS and AMD make the word “only” uncomfortable.

2 weeks, 2 days назад @ thesequence.substack.com
The Sequence AI of the Week #899: Inside Inkling: A Trillion-Parameter Model That Only Wakes Up 41 Billion at a Time
The Sequence AI of the Week #899: Inside Inkling: A Trillion-Parameter Model That Only Wakes Up 41 Billion at a Time The Sequence AI of the Week #899: Inside Inkling: A Trillion-Parameter Model That Only Wakes Up 41 Billion at a Time

Inkling is best understood not as a single 975-billion-parameter brain that fires all at once, but as a giant warehouse of specialist capacity.

The more interesting number is the one beside it: 41 billion parameters active per token.

A good mental picture is a university with 256 specialist departments on each relevant floor.

It selects six departments that appear useful for this token, adds two general-purpose departments that always attend, combines their work, and moves on.

A line of Python might summon one set of specialists; a phrase in Greek, a diagram label, or a piece of audio may summon another.

2 weeks, 3 days назад @ thesequence.substack.com
The Sequence Knowledge #898: The Trace Is the Teacher: Distilling Reasoning Into Small Models
The Sequence Knowledge #898: The Trace Is the Teacher: Distilling Reasoning Into Small Models The Sequence Knowledge #898: The Trace Is the Teacher: Distilling Reasoning Into Small Models

The distilled 32B model started solving competition math it had no business solving.

The 7B model began verifying its own work and branching its reasoning mid-stream — emergent behaviors nobody trained into it directly.

A grab-bag of small dense models suddenly reasoned like something ten times their size.

We spent an entire installment establishing why naive sequence-level imitation is the wrong tool, and then the single most important reasoning-distillation result of the decade is naive sequence-level imitation.

The answer is the whole story of this installment, and it turns out to be more interesting than either “imitation works” or “imitation doesn’t.”

2 weeks, 4 days назад @ thesequence.substack.com
The Sequence Radar #897: Last Week in AI: China, Compression and the Open-Model Race
The Sequence Radar #897: Last Week in AI: China, Compression and the Open-Model Race The Sequence Radar #897: Last Week in AI: China, Compression and the Open-Model Race

In the AI of the Week , we discuss Thinking Machine first open weights model.

Subscribe and don’t miss out:📝 Editorial: Last Week in AI: China, Compression and the Open-Model RaceFor years, AI progress has been narrated as a horse race: larger models, higher benchmark scores, more expensive clusters.

Moonshot describes it as the first open model in the three-trillion-parameter class, although the full weights are not due until later this month.

This week, AI stopped looking like a single race.

AI Lab: Shanghai AI LaboratorySummary: ADVANCED MATHBENCH introduces a rigorous evaluation suite focusing on the generation and process-level verification of advanced, natural-language mathematical pr…

2 weeks, 6 days назад @ thesequence.substack.com
The Sequence Opinion #896: Spark, Compute, and the Two Metas
The Sequence Opinion #896: Spark, Compute, and the Two Metas The Sequence Opinion #896: Spark, Compute, and the Two Metas

The occasion was the launch of Muse Spark 1.1, the second model out of Meta Superintelligence Labs and the first Meta model ever to ship with a price tag.

Two days earlier Meta shipped Muse Image, its first image generation model from the new lab.

Chips, datacenters, cloud, models, API, apps, devices.

At the layer where models meet users, the app and agent layer, Meta might be the favorite.

At the layer where models get made, the evidence is thin and the structural arguments cut against it.

3 weeks, 2 days назад @ thesequence.substack.com
The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break
The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break

OpenAI’s audit of SWE-Bench Pro shows why a precise score can still be a poor measure - and why coding agents may become essential tools for auditing the benchmarks that grade them.

A frontier coding score can look wonderfully precise: 80.3 percent, one decimal place, clean enough to rank models and anchor product claims.

That is the uncomfortable conclusion of OpenAI’s audit of SWE-Bench Pro.

OpenAI estimates that roughly 30 percent of the public benchmark is broken.

OpenAI has withdrawn its earlier recommendation that the field adopt SWE-Bench Pro.

3 weeks, 3 days назад @ thesequence.substack.com
Synced Review
последний пост None
📓 Cool Blogs
ODS.ai Habr ODS.ai Habr
последний пост 4 months назад
Вайбкодинг по Chess’ноку. 1. e4
Вайбкодинг по Chess’ноку. 1. e4 Вайбкодинг по Chess’ноку. 1. e4

Но это не вайбкодинг, а тяжёлая профессиональная ИИ-разработка.

За это время по этому проекту в ChatGPT было создано 112 чатов — это примерно 560 промптов.

И в особо напряжённые периоды приходилось вставать по ночам, чтобы оптимально использовать лимиты, которые делятся на 5-часовые и недельные сессии.

Но это не магия и не кнопка «сделать хорошо».

Именно поэтому будущее не за вайбкодингом, а за теми, кто научится управлять этой скоростью.

4 months назад @ habr.com
Почему я стал ИТ-волонтером & Датасет новостей о противоречиях современного общества
Почему я стал ИТ-волонтером &amp; Датасет новостей о противоречиях современного общества Почему я стал ИТ-волонтером &amp; Датасет новостей о противоречиях современного общества

Простой пример с ценами на топливо: бензин дорожает и из-за роста цены на нефть, и из-за ее падения.

Осознание того, что твой труд увеличивает чью-то капитализацию, но не решает реальных проблем общества, видимых в быту и в новостях, подтолкнуло искать еще какую-то деятельность.

Кроме того, благодаря АМБ появился уникальный датасет новостей с противоречиями современного общества на kaggle и github, далее о нем.

Датасет новостей о противоречиях современного обществаАктивисты АМБ и волонтеры дружественных коллективов собрали и разметили датасет новостей, подсвечивающие те самые системные противоречия, о которых я задумывался ранее.

Пример Б В 2023 году в мире голодал каждый 11-й человек, а в …

5 months, 2 weeks назад @ habr.com
[Перевод] Как устроен Codex
[Перевод] Как устроен Codex [Перевод] Как устроен Codex

Подробный разбор того, как команда OpenAI Codex создаёт своего кодового агента, как его используют инженеры и что это может значить для будущего разработки ПО.

Чтобы разобраться, как устроен Codex, как команды внутри OpenAI его используют и как он влияет на инженерные практики у создателей ChatGPT, я поговорил с тремя сотрудниками OpenAI:Тибо Соттио (Thibault Sottiaux) — руководитель Codex.

Оба продукта были запущены весной: Codex CLI анонсировали в апреле 2025 года, а Codex в ChatGPT представили в мае.

В команде Codex эти файлы объясняют агенту, как ориентироваться в кодовой базе, какие команды запускать для тестирования и как следовать стандартам проекта.

Использование Codex в OpenAIПомим…

5 months, 2 weeks назад @ habr.com
Курс Natural Language Processing & LLMs — новый сезон
Курс Natural Language Processing &amp; LLMs — новый сезон Курс Natural Language Processing &amp; LLMs — новый сезон

10 февраля мы в очередной раз запускаем бесплатный онлайн-курс по обработке естественного языка (Natural Language Processing).

Что будем проходить:классическое начало: закон Ципфа, TF-IDF, RNN, CNN, Transformer;основные задачи NLP: классификация текста, тегирование и генерация;специфичные области: агенты и вайб-кодинг;LLM и их применение.

Если вы студент ИТМО, МФТИ или ВШЭ, то курс можно зачесть, как учебный.

Работаю в области NLP более 12 лет, успел поработать в Яндексе и ВКонтакте, защитить кандидатскую диссертацию.

Если есть вопросы, то приходите с ними в ODS Mattermost – там будут все ответы, время семинаров и ссылки.

6 months, 1 week назад @ habr.com
Machine Learning Mastery
последний пост 1 week, 3 days назад
Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026?
Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026? Ollama vs. LM Studio vs. llama.cpp: Which Local AI Runtime Should You Use in 2026?

Then we walked through the fastest way to get inference running locally in Run a Local AI Model in 15 Minutes: Your First Ollama Setup.

Spend enough time in the local AI ecosystem, though, and you’ll notice Ollama isn’t the only option competing for your hard drive.

Three tools dominate the local AI runtime landscape: Ollama, LM Studio, and llama.cpp.

Ollama (Via its dedicated, background-daemon CLI) ollama run llama3 .

There’s a well-worn progression in the local AI community that maps almost exactly to the three tools covered here: LM Studio → Ollama → llama.cpp.

1 week, 3 days назад @ machinelearningmastery.com
5 Architectural Patterns for Persistent Memory and State in AI Agents
5 Architectural Patterns for Persistent Memory and State in AI Agents 5 Architectural Patterns for Persistent Memory and State in AI Agents

The fix isn’t a bigger context window; it’s treating memory and state as deliberate architectural decisions, not afterthoughts.

Memory feeds into state; state feeds back into memory.

Also worth calling out explicitly: credentials and secrets are not semantic memory.

Episodic Event Logs (Historical Reflection)Semantic memory stores what the agent knows; episodic memory stores what the agent did.

The moment your system serves more than one user, memory has to be siloed.

1 week, 5 days назад @ machinelearningmastery.com
Stateful vs. Stateless Agent Design: Tradeoffs for Scalable Agentic Systems
Stateful vs. Stateless Agent Design: Tradeoffs for Scalable Agentic Systems Stateful vs. Stateless Agent Design: Tradeoffs for Scalable Agentic Systems

environ [ "GROQ_API_KEY" ] = "PASTE_YOUR_GROQ_API_KEY_HERE" # Initializing the client client = Groq ( ) # Using an efficient model from Groq: Llama 3.1 8B Instant MODEL_ID = "llama-3.1-8b-instant"An important setup decision here is the choice of a specific model.

strip ( )To understand the limitations of a stateless agent, we simulate a simple user-model conversation through it:# --- Testing the Stateless Agent --- print("--- Turn 1 ---") prompt_1 = "Hi, my name is Alice and I am learning about API infrastructure."

--- Turn 2 (Without Client Context) --- Agent: Unfortunately, I don't have any information about you, including your name.

-- - Turn 2 ( Without Client Context ) -- - Agent : Unf…

2 weeks, 1 day назад @ machinelearningmastery.com
An Introduction to Loop Engineering
An Introduction to Loop Engineering An Introduction to Loop Engineering

Topics we will cover include:The origin and definition of loop engineering, and how it fits into the broader progression from prompt engineering to context engineering to harness engineering.

The three hardest problems in loop engineering — context management, termination, and verification — and the failure modes that result from getting any one of them wrong.

Prompt engineering effectively became one ingredient within context engineering rather than a separate discipline.

Loop engineering is simply the part where all of that gets put into motion and given a rhythm.

It’s that “loop engineering” is a product name and a rallying phrase for a research direction that’s been quietly accumulating…

2 weeks, 2 days назад @ machinelearningmastery.com
The Current State of Agentic AI
The Current State of Agentic AI The Current State of Agentic AI

How the Model Context Protocol, persistent memory graphs, and emerging security patterns define the current production landscape.

IntroductionLook back at how we built AI agents just a year ago, and the dominant paradigm was brute-force orchestration.

This tutorial breaks down the current state of agentic AI architecture, covers the three major shifts defining production systems today, and walks through how to design a modern agent swarm.

The current state of tool calling is increasingly defined by the Model Context Protocol (MCP).

This open standard acts as a universal adapter between AI models and local or remote data sources.

2 weeks, 4 days назад @ machinelearningmastery.com
Building Agentic Workflows in Python with LangGraph
Building Agentic Workflows in Python with LangGraph Building Agentic Workflows in Python with LangGraph

Managing Conversation History with MessagesStateEvery node in a LangGraph graph reads the current state and writes updates back to it.

add_node ( "run_model" , run_model ) builder .

Define a tool with the @tool decorator:from langchain_core.tools import tool @tool def get_customer_tier(customer_id: str) -> str: """Look up the subscription tier for a customer by their ID.

tools import tool @ tool def get_customer_tier ( customer_id : str ) -> str : "" "Look up the subscription tier for a customer by their ID.

add_node ( "run_model" , run_model ) builder .

2 weeks, 5 days назад @ machinelearningmastery.com
Agentic AI Security: Defending Against Prompt Injection and Tool Misuse
Agentic AI Security: Defending Against Prompt Injection and Tool Misuse Agentic AI Security: Defending Against Prompt Injection and Tool Misuse

Share Post ShareIn this article, you will learn what prompt injection and tool misuse are in the context of agentic AI systems, and which defense strategies experts recommend to mitigate them.

Topics we will cover include:How prompt injection and tool misuse can compromise AI agents deployed in real-world production environments.

Prompt injection arises when untrusted inputs to a language model are interpreted as instructions rather than mere data.

This problem has been renamed Agent Goal Hijacking in the context of agentic AI and AI security vulnerabilities.

Closing Remarks: Looking AheadIn line with the growing level of sophistication attained by agentic AI systems, organizations should a…

3 weeks, 1 day назад @ machinelearningmastery.com
Run a Local AI Model with Ollama in 15 Minutes
Run a Local AI Model with Ollama in 15 Minutes Run a Local AI Model with Ollama in 15 Minutes

Topics we will cover include:Why Ollama has become the standard tool for running local AI models.

Ollama has become the go-to tool for local AI because it packages complex model architectures into a clean, lightweight background service.

# Verify Ollama is running by checking the version ollama --version # Pull and immediately run the Llama 3.2 3B model ollama run llama3.2 1 2 3 4 5 # Verify Ollama is running by checking the version ollama -- version # Pull and immediately run the Llama 3.2 3B model ollama run llama3 .

Just run your command directly ( ollama run llama3.2 ), the background daemon is already listening on port 11434.

From here, exploring the other models from our Top 7 list is…

3 weeks, 2 days назад @ machinelearningmastery.com
Scikit-Ollama for Scikit-LLM/Ollama Integration
Scikit-Ollama for Scikit-LLM/Ollama Integration Scikit-Ollama for Scikit-LLM/Ollama Integration

zero_shot import ZeroShotOllamaClassifier # Initializing the classifier with our local Ollama model: llama3:latest clf = ZeroShotOllamaClassifier ( model = "llama3:latest" )A very important clarification about what we just did.

Predicted Sentiment: positive Text: 'The special effects in 'Star Battles: Nebula Conflict' were out of this world.

Predicted Sentiment: positive Text: ''The Lost Symphony' was a masterclass in character development and storytelling.

Predicted Sentiment: positive Text: ' The special effects in 'Star Battles: Nebula Conflict' were out of this world .

The key ingredient: the scikit-ollama library, which elegantly encapsulates this local integration and makes it availab…

3 weeks, 3 days назад @ machinelearningmastery.com
LLM Evaluation Frameworks Compared: How to Actually Measure What Your Model Does
LLM Evaluation Frameworks Compared: How to Actually Measure What Your Model Does LLM Evaluation Frameworks Compared: How to Actually Measure What Your Model Does

])\s+', answer.strip()) return [s.strip() for s in sentences if s.strip()] def claim_supported_by_context(claim: str, context: str) -> bool: """ Check whether a claim has lexical support in the retrieved context.

RAGAS does this with an LLM judge; this overlap check demonstrates the same supported/unsupported decision in a deterministic way. """

strip ( ) ) return [ s . strip ( ) for s in sentences if s . strip ( ) ] def claim_supported_by_context ( claim : str , context : str ) -> bool : "" " Check whether a claim has lexical support in the retrieved context.

def your_judge_function(response_1: str, response_2: str) -> str: # Placeholder -- wire this up to your actual judge model call.

def…

3 weeks, 4 days назад @ machinelearningmastery.com
Choosing the Right AI Agent Memory Strategy: A Decision-Tree Approach
Choosing the Right AI Agent Memory Strategy: A Decision-Tree Approach Choosing the Right AI Agent Memory Strategy: A Decision-Tree Approach

The common pitfalls that show up once agent memory is implemented, and how to fix them.

Agent memory strategy deserves the same deliberate design as orchestration.

Why Is Choosing an AI Agent Memory Strategy Important?

A customer support agent, for example, might keep the current ticket in working memory, a customer’s subscription tier in semantic memory, past complaints in episodic memory, and a learned refund-handling routine in procedural memory.

As discussed, working memory, semantic memory, episodic memory, and procedural memory serve different purposes and require different storage and retrieval strategies.

4 weeks назад @ machinelearningmastery.com
LLM Orchestration Frameworks Compared: LangChain vs. LlamaIndex vs. Raw API Calls
LLM Orchestration Frameworks Compared: LangChain vs. LlamaIndex vs. Raw API Calls LLM Orchestration Frameworks Compared: LangChain vs. LlamaIndex vs. Raw API Calls

openai import OpenAI as LlamaOpenAI from llama_index .

# Prerequisites: pip install openai python-dotenv # How to run: python raw_api_agent.py import os import json from dotenv import load_dotenv from openai import OpenAI load_dotenv ( ) client = OpenAI ( api_key = os .

Prerequisites:pip install openai langchain langchain-openai llama-index \ llama-index-llms-openai llama-index-embeddings-openai python-dotenv 1 2 pip install openai langchain langchain - openai llama - index \ llama - index - llms - openai llama - index - embeddings - openai python - dotenvHow to run: Save as three_ways.py and run python three_ways.py# three_ways.py # The same document Q&A task implemented three ways: # Raw …

4 weeks, 1 day назад @ machinelearningmastery.com
Tools vs. Subagents: Building Effective AI Agents Without Over-Engineering
Tools vs. Subagents: Building Effective AI Agents Without Over-Engineering Tools vs. Subagents: Building Effective AI Agents Without Over-Engineering

Topics we will cover include:What tools and subagents are, and the key differences between them.

This article explains what tools and subagents are, where each fits, and how to make the choice every time.

When an agent calls a tool, the result lands back in the same context the agent is actively reasoning in — prior reasoning, tool result, and everything else together.

A database record, search result, or API response can often be consumed immediately.

If the answer is independent reasoning, context isolation, specialized capabilities, or parallel execution, a subagent is likely justified.

1 month назад @ machinelearningmastery.com
The Complete Guide to Tool Selection in AI Agents
The Complete Guide to Tool Selection in AI Agents The Complete Guide to Tool Selection in AI Agents

tools = tools self .

tools = tools self .

threshold : return { "status" : "resolved" , "tool" : tool [ "name" ] , "confidence" : score , "attempts" : 1 } # Reformulate by stripping filler words.

tools = tools self .

full_catalog_tokens = sum ( estimate_tokens ( d ) for d in descs ) def _retrieve ( self , query : str , top_k : int ) -> list [ dict ] : query_vec = self .

1 month назад @ machinelearningmastery.com
Context vs. Memory Engineering in Agentic AI Systems
Context vs. Memory Engineering in Agentic AI Systems Context vs. Memory Engineering in Agentic AI Systems

Share Post ShareIn this article, you will learn how context engineering and memory engineering solve different problems in agentic AI systems, and how the two disciplines meet at the point where retrieved memory enters the context window.

Most of the time, the problem lies in two areas that get built together, conflated, or skipped: context engineering and memory engineering.

Memory Engineering: Designing Persistent AI Memory SystemsOnce an inference call completes, memory engineering determines what deserves to persist and under what conditions it gets used again.

trust_level >= 0.5 )AI Agent Memory Design Guide – Working, Long-Term, and Procedural Memory with Forgetting and Staleness Mana…

1 month, 1 week назад @ machinelearningmastery.com
ML in Production
последний пост None
Sorta Insightful Sorta Insightful
последний пост 3 weeks, 1 day назад
Which Tech CEOs Are Gamers?
Which Tech CEOs Are Gamers? Which Tech CEOs Are Gamers?

Reading Satya testifying about his gamer cred was ridiculous enough to inspire a dumb idea: which tech CEOs are gamers?

I did not find any mention of either playing video games.

There is one NYT article that mentions Elon Musk used to crash at Larry Page’s place after playing video games, but it never says if Page played video games, so I will play it safe and say neither are gamers.

The main video game related story Steve is tied to is the Atari Breakout debacle, which you probably already know.

Given how new his rise to tech CEO celebrity-ism is, you’d think there wouldn’t be much information about his video game habits, but somehow, there is.

3 weeks, 1 day назад @ alexirpan.com
AI Will Not Make Your Job Chill
AI Will Not Make Your Job Chill AI Will Not Make Your Job Chill

People keep talking about how AI will make their job easy, and I don’t really understand why.

I assume the factory job producing this was still hard work.

I don’t think AI has made my job chill, and I feel like I am front-line compared to much of the economy.

It’s not widely known, but transportation and warehousing has the highest rate of nonfatal work injuries in the US.

For a while, this will not lead to any job loss, because increasing abundance will lead to higher demand.

2 months, 3 weeks назад @ alexirpan.com
Why I Signed The Amicus Brief for Anthropic v Department of War
Why I Signed The Amicus Brief for Anthropic v Department of War Why I Signed The Amicus Brief for Anthropic v Department of War

On Monday, Anthropic filed a lawsuit against the Department of War, and an amicus brief in support of Anthropic was filed on behalf of a number of OpenAI and Google employees.

There’s also an amicus brief filed on behalf of Microsoft.

There’s conflicting reporting, but very broadly, Anthropic signed an agreement with the government to deploy Claude in classified, military contexts.

Anthropic said no, Pete Hegseth declared them a supply chain risk, and Anthropic filed a lawsuit against this.

The amicus brief was broadly aligned with my thoughts on the matter, so I signed.

5 months назад @ alexirpan.com
MIT Mystery Hunt 2026
MIT Mystery Hunt 2026 MIT Mystery Hunt 2026

This has spoilers for MIT Mystery Hunt 2026.

Pre-HuntThe time running up to Hunt was more stressful than usual…very briefly, I typically hunt with teammate.

Just last year, I did GPH 2025, LN Hunt, Teammate Hunt 2025, Microsoft Hunt 2025, and Silph Puzzle Hunt 2025, all of which had significant 3+ hour solve puzzles that would not be out of place in Mystery Hunt.

Not to mention smaller hunts like Advent Hunt, and then I didn’t even do Brown Puzzlehunt or Vertex Hunt or the fall CMU Hunt.

To me, the crux is whether Mystery Hunt is broken, or Mystery Hunt is fine.

6 months, 1 week назад @ alexirpan.com
Authentic Imperfection
Authentic Imperfection Authentic Imperfection

* * *I’ve been thinking about the anger surrounding generative AI.

To keep things fair, he took the best human images and best AI images, meaning human art from famous artists, and AI art from prompters skilled at removing obvious tells of image generation.

When people complain about AI slop, I see it as a complaint against the deluge of default style AI images.

We’ve seen this happen in all forms: AI text, AI music, older forms of computer generated content like CGI.

As much as we celebrate imperfection, digital imperfection is a step too far.

8 months, 3 weeks назад @ alexirpan.com
Lil'Log
последний пост None
inFERENCe
последний пост 5 months, 1 week назад
The Future of Software
The Future of Software The Future of Software

February 25, 2026The Future of SoftwareThe world of software is undergoing a shift not seen since the advent of compilers in the 1970s.

How will humans tell AI agents what software artefacts we would like to create?

How will humans tell AI agents what software artefacts we would like to create?

This future of software creation, in which our programming languages are abstracted away, raises two very important questions:What will the instruction/specification language look like?

This should be a clear layer of separation between the developer and the pool of AI agents working to maintain software.

5 months, 1 week назад @ inference.vc
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On
Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years On

Deep Learning is Powerful Because It Makes Hard Things Easy - Reflections 10 Years OnTen years ago this week, I wrote a provocative and bold post that blew up, made it to top spot on HackerNews.

In hindsight: There is a lot of stuff in deep learning that we don't understand nearly enough.

Sometimes things work for reasons completely unrelated to why we thought they would work.

(Pop some 🍿 in the microwave and read till the end for more)🎯 "Deep learning is powerful exactly because it makes hard things easy"Okay, this was a great insight.

🎯 Generative ModelingIn the post I suggested people learn "something harder" instead of - or in addition to - deep learning.

6 months, 1 week назад @ inference.vc
The Spectator
последний пост None
The Unofficial Google Data Science Blog The Unofficial Google Data Science Blog
последний пост None
Off the Convex Path
последний пост None
Jay Alammar
последний пост None
Piekniewski's blog
последний пост None
fast.ai NLP fast.ai NLP
последний пост None
Sebastian Ruder
последний пост None
大トロ 大トロ
последний пост None
🔬 Science
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
Papers With Code Papers With Code
последний пост None
💼 University and corporation labs
DeepMind DeepMind
последний пост 2 days назад
WeatherNext: AI model achieves breakthrough in forecasting cyclones
WeatherNext: AI model achieves breakthrough in forecasting cyclones WeatherNext: AI model achieves breakthrough in forecasting cyclones

Predicting how dangerous cyclones develop is a longstanding challenge where every hour counts.

Today, in a paper published in Nature, we show that our WeatherNext AI model achieved state-of-the-art accuracy in predicting a cyclone's track, intensity, and wind structure.

During the 2025 hurricane season, our model helped the NHC to make a historic forecast for Hurricane Melissa by predicting the storm’s rapid intensification and landfall in Jamaica.

Given this broad impact, we are now open sourcing our WeatherNext 2 and WeatherNext Cyclones models used during the hurricane season.

How WeatherNext predicts weather and cyclones

2 days назад @ deepmind.google
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

That’s why today we’re launching Gemini Robotics ER 2, our most capable “embodied reasoning” model for robotics.

Gemini Robotics ER 2 can also natively call tools like Google Search to find information, or any other user-defined function.

Gemini Robotics ER 2 represents a significant upgrade over Gemini Robotics ER 1.6.

Gemini Robotics ER 2 is now publicly available to developers via the Gemini API, Google AI Studio, and in private preview on Gemini Enterprise Agent Platform.

Gemini Robotics ER 2 improves this tool orchestration workflow.

1 week, 2 days назад @ blog.google
We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control
We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control

Our newest music generation model, Lyria 3.5, delivers significant advancements across musicality, lyrics, and vocal quality, empowering you to craft richer tracks.

We’re rolling it out today in Google Flow Music, where we want to help you create songs you love, with creative control.

Enhanced lyrics: Generate higher quality lyrics with improved prompt adherence and structural awareness.

Generate higher quality lyrics with improved prompt adherence and structural awareness.

Improved vocals: Bring more expression and emotion to your songs with more realistic and emotionally nuanced vocals, plus improved pronunciation.

1 week, 2 days назад @ blog.google
Gemini Robotics 2 brings whole body intelligence to robots
Gemini Robotics 2 brings whole body intelligence to robots Gemini Robotics 2 brings whole body intelligence to robots

From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasksFor decades, we’ve dreamed of robots that can seamlessly step into our world and lend a hand.

Today, we are introducing Gemini Robotics 2 - the intelligence layer powering the next generation of truly adaptable robots.

As it takes its first literal steps, this major advance unlocks intelligent whole-body control, advanced dexterity, and multi-robot collaboration.

Gemini Robotics 2 enables robots to reason through every movement, unlocking a broad range of tasks.

And this profound intelligence can also run locally on-device while seamlessly a…

1 week, 4 days назад @ deepmind.google
Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission
Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission

In December, we shared our commitment to the White House's Genesis Mission — the national effort to harness AI and double the pace of American scientific discovery within a decade.

Today, at the DOE Genesis Mission Summit 2026, we are expanding this by committing $40 million of AI tokens and cloud credits for researchers in support of the Genesis Mission.

WeatherNext — a state-of-the-art family of AI weather forecasting models for mapping weather conditions.

— a state-of-the-art family of AI weather forecasting models for mapping weather conditions.

Driving American innovationThe Genesis Mission represents an opportunity to transform research and science across America.

2 weeks, 3 days назад @ cloud.google.com
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Building on Gemini 3.5 Flash, we’re introducing new Gemini models:3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance.

3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure.

3.6 Flash: More efficient and better quality than 3.5 FlashGemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash.

For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash.

This enhanced efficiency is also combined with a lower price than 3.5 Flash.

2 weeks, 3 days назад @ blog.google
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Building on Gemini 3.5 Flash, we’re introducing new Gemini models:3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance.

3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure.

3.6 Flash: More efficient and better quality than 3.5 FlashGemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash.

For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash.

This enhanced efficiency is also combined with a lower price than 3.5 Flash.

2 weeks, 3 days назад @ blog.google
Introducing Gemini 3.5 Flash Cyber
Introducing Gemini 3.5 Flash Cyber Introducing Gemini 3.5 Flash Cyber

Today, we’re expanding our longtime efforts to better prepare defenders by introducing Gemini 3.5 Flash Cyber, our lightweight cybersecurity model built on top of 3.5 Flash and fine-tuned to find, validate, and patch vulnerabilities quickly and efficient, making it more effective at these tasks than Gemini’s mainline Flash models.

By building on top of Flash, 3.5 Flash Cyber offers a cost-efficient and highly capable alternative to large, costly cybersecurity models.

Given the dual-use nature of this technology, we have taken an intentional approach to how we deploy 3.5 Flash Cyber.

3.5 Flash Cyber benchmark results: an efficient alternative to larger cybersecurity modelsWe tested 3.5 F…

3 weeks, 1 day назад @ deepmind.google
Our approach to bioresilience
Our approach to bioresilience Our approach to bioresilience

Today, Google DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience.

Inside our bioresilience programWe believe society must harness AI’s advancing capabilities to address infectious diseases and prepare for future outbreaks.

With this in mind, we are making our AI models and agents available to trusted partners to support progress across three key areas: prevention, detection and response.

Working in collaboration with governments and global health authorities to advance a diverse range of diagnostic and therapeutic strategies enables the Isomorphic Labs Drug Design Engine’s real-world impact for bioresilience.

To read more about our work and our call for new par…

3 weeks, 2 days назад @ deepmind.google
Empowering India’s next generation of innovators with ATL Saathi
Empowering India’s next generation of innovators with ATL Saathi Empowering India’s next generation of innovators with ATL Saathi

A new contribution to Indian Education with Atal Innovation MissionWe believe behind every good student is a great teacher.

That’s why for over 20 years, Google has been dedicated to supporting the education ecosystem by introducing technology into teaching and learning through a teacher-led approach.

With foundational platforms like Google for Education and Google Classroom, we build products tailored to the needs of schools, keeping the teacher in the lead.

To further support the empowerment of educators, our new Google Educator AI Series ensures teachers are equipped with both the tools and the digital skills required for today's classrooms.

We see Gemini as a great tool to enable our pa…

3 weeks, 5 days назад @ deepmind.google
Google DeepMind and A24 announce first-of-its-kind research partnership
Google DeepMind and A24 announce first-of-its-kind research partnership Google DeepMind and A24 announce first-of-its-kind research partnership

Today, Google DeepMind and A24 are announcing a first-of-its-kind partnership focused on research.

The collaboration pairs a world-leading research lab with the industry’s most filmmaker-forward studio to help artists develop new workflows and techniques.

This partnership creates a deep research and development collaboration between A24 and Google DeepMind spanning multiple projects over time.

This hands-on collaboration provides Google DeepMind with invaluable feedback and guidance from leading artists.

As A24 and Google DeepMind’s researchers work side-by-side to test, iterate and build, this partnership aims to expand what is possible in the future of entertainment.

1 month назад @ blog.google
Start building with Nano Banana 2 Lite and Gemini Omni Flash
Start building with Nano Banana 2 Lite and Gemini Omni Flash Start building with Nano Banana 2 Lite and Gemini Omni Flash

Uploading audio references and scene extension is not yet supported in the Gemini API for this model.

Video references up to 3 seconds in duration are accepted by the API schema but are not correctly processed by the model at this time.

Gemini Omni is available in public preview starting today in Google AI Studio and the Gemini API.

Use Nano Banana 2 Lite as a high-speed image generation model, then pass that image as a reference to Gemini Omni Flash to animate it into a high-quality video.

To help you get started we created a few demo apps you can remix that let you experience how you can pair both Nano Banana 2 Lite and Gemini Omni Flash into one workflow.

1 month, 1 week назад @ blog.google
Introducing computer use in Gemini 3.5 Flash
Introducing computer use in Gemini 3.5 Flash Introducing computer use in Gemini 3.5 Flash

Making computer use safe in 3.5 FlashTo mitigate some of the prompt injection risks for agents operating in live environments, we use targeted adversarial training for computer use in Gemini 3.5 Flash.

We’re also releasing two optional enterprise safeguard systems that enable enterprises to:Require explicit user confirmation for sensitive or irreversible actions.

Automatically stop tasks if an indirect prompt injection is identified.

Taking a “defense-in-depth” approach, we encourage developers to combine these features with secure sandboxing, human-in-the-loop verification and strict access controls.

We are already seeing customers drive value with computer use.

1 month, 2 weeks назад @ blog.google
Unlocking UK house-building with AI-accelerated planning
Unlocking UK house-building with AI-accelerated planning Unlocking UK house-building with AI-accelerated planning

New UK government AI planning prototype built with Gemini aims to halve the time it takes to process homeowner applicationsAround the world, Governments are exploring how AI can deliver better public services, faster.

The UK is working to build 1.5 million new homes by 2029, but local planning authorities are often slowed down by dense paperwork and administrative backlogs.

To help get Britain building, we’re partnering with the UK government to help radically shorten the time it takes to process householder planning applications.

Following early trials in Barnet, Camden and Dorset, the government plans for the new AI planning tool to be made available to all councils nationally from 2027.

1 month, 3 weeks назад @ deepmind.google
Securing the future of AI agents
Securing the future of AI agents Securing the future of AI agents

How we’re securing internal systems against increasingly capable and imperfectly aligned AIAI agents are transforming our relationship with technology.

In the U.S alone, AI agents could create $2.9 trillion in economic value by 2030.

That’s why we developed our AI Control Roadmap: a framework for building and managing the advanced AI we deploy within Google.

Similarly, our AI control system grants AI agents permissions based on their verified behavior, allowing us to build trust through controlled, incremental access.

In our AI Control Roadmap, we map security protocols to measurable milestones in AI capabilities on two critical fronts:

1 month, 3 weeks назад @ deepmind.google
Google
последний пост 1 day, 23 hours назад
Your agentic summer: No-cost lessons from Google experts to build and scale agents
Your agentic summer: No-cost lessons from Google experts to build and scale agents Your agentic summer: No-cost lessons from Google experts to build and scale agents

Intro to AI Agents: Build a foundational understanding of how autonomous agents can redefine productivity.

Enterprise Agents and Use Cases: Discover how AI agents drive real business impact.

Create Your First Gemini Enterprise Application skill badge: Earn a skill badge that proves you can create an app with Gemini Enterprise.

Orchestrate Multi-Agent Workflows with Gemini Enterprise skill badge: Demonstrate your ability to manage multiple agents powered by Gemini Enterprise with a skill badge.

Engineer AI Agents with Agent Development Kit (ADK) skill badge: Build production-grade agents using expert developer tools.

1 day, 23 hours назад @ cloud.google.com
Mirendil taps AI Hypercomputer TPUs and GPUs for pre- and post-training applications
Mirendil taps AI Hypercomputer TPUs and GPUs for pre- and post-training applications Mirendil taps AI Hypercomputer TPUs and GPUs for pre- and post-training applications

Nearly every major AI lab uses Google Cloud infrastructure, including for training of models, inference for agents, and new frontier research.

Today, we’re announcing that Mirendil, an exciting frontier AI lab focused on accelerating AI development, will also utilize Google Cloud’s AI Hypercomputer.

This includes using a mix of Google’s TPU AI accelerators and full-stack NVIDIA AI infrastructure running on Google Cloud; this purpose-built AI infrastructure will support model pre-training and post-training applications for Mirendil.

The Mirendil team is building new AI systems that can help accelerate and democratize AI research and development.

We closely partnered with Mirendil on end-to-e…

2 days, 2 hours назад @ cloud.google.com
How Target is enhancing retail discovery and cutting database maintenance by 50% with Spanner Graph
How Target is enhancing retail discovery and cutting database maintenance by 50% with Spanner Graph How Target is enhancing retail discovery and cutting database maintenance by 50% with Spanner Graph

Building the enterprise ontology on Spanner GraphWe evaluated multiple specialized technologies, including standalone vector databases and niche graph databases.

We ultimately chose Spanner Graph to build our enterprise ontology, which is a "graph-of-graphs" paradigm that allows us to construct a massive, generative AI-powered shopping graph.

By unifying our data, we bring semantic data, graph relationships, vector embeddings, and operational transactions under one roof.

Spanner Graph natively supports multi-hop graph traversals, semantic vector similarity, and full-text keyword queries over our relational tables.

Consolidated SQL + GQL interoperability: With Spanner Graph, our developers q…

3 days, 23 hours назад @ cloud.google.com
How Deutsche Bank unlocked agility with an API-ready ecosystem
How Deutsche Bank unlocked agility with an API-ready ecosystem How Deutsche Bank unlocked agility with an API-ready ecosystem

At Deutsche Bank, we recognized that APIs aren't just technical plumbing; they're the nervous system of modern banking.

We've moved from "Where's that customer data API?"

Security: the employee onboarding analogyWhen thinking about API security, imagine onboarding a new employee.

Like employee access, these permissions are centrally managed, regularly audited, and instantly revocable.

This visibility serves operations, product managers who track partner value, and security teams who identify anomalies.

3 days, 23 hours назад @ cloud.google.com
Real-world mainframe modernization with AI: A safe, scalable path from mainframe to cloud
Real-world mainframe modernization with AI: A safe, scalable path from mainframe to cloud Real-world mainframe modernization with AI: A safe, scalable path from mainframe to cloud

At Google Cloud, we propose an alternative: a modernization strategy that leverages the power of AI, agility of the cloud and allows for iterative and continuous modernization.

This approach recognizes a fundamental truth: mainframe modernization isn’t a pure code-to-code conversion problem.

Deep operational lock-in with specialized proprietary mainframe utility suites.

In other words, real-world modernization of mainframe applications is so much more than converting COBOL to Java.

Our solutions span four core pillars: assessment, modernization, de-risking, and data migration.

4 days, 22 hours назад @ cloud.google.com
What Google Cloud announced in AI this month
What Google Cloud announced in AI this month What Google Cloud announced in AI this month

Along with a batch of new platform updates, this month we’ve put together 13 practical demos and 20 diagnostic questions to help your engineering teams align on a strong architectural blueprint.

Top announcementsThought leadership (editor’s pick):If automation requires delegation, then delegation requires trust.

But letting an AI agent run on its own is a big leap for any business.

What makes an AI agent trustworthy: Context is fast becoming one of the most valuable assets a company owns.

Prajakta Damle, Senior Director, Product Management, shares what it takes to get trustworthy AI right.

1 week назад @ cloud.google.com
What’s new in AI infrastructure and orchestration this month
What’s new in AI infrastructure and orchestration this month What’s new in AI infrastructure and orchestration this month

We incorporate AI into the tools you use every day (think Gmail, BigQuery, AlloyDB, Google Cloud Code and Google Cloud Assist).

We make software frameworks to help you build with AI, like Gemini Enterprise Agent Platform, JAX, or MaxTest.

To support this, we are making AI infrastructure and orchestration news at a furious pace.

Report: We recently surveyed more than 1,400 senior IT leaders for our State of AI Infrastructure report, and a resounding pattern emerged: The gap between AI ambition and infrastructure reality is widening.

Product deep dive: We went into depth about Cloud Storage Rapid, a new family of high-performance storage offerings for AI workloads.

1 week назад @ cloud.google.com
Cloud CISO Perspectives: Why AI Threat Defense is the new boardroom baseline
Cloud CISO Perspectives: Why AI Threat Defense is the new boardroom baseline Cloud CISO Perspectives: Why AI Threat Defense is the new boardroom baseline

Expected operational standard: Your organizational mean time to remediate (MTTR) exposures and other desired changes into production goes down and to the right.

System consolidation: Boards should look beyond standalone AI features and point products to address systemic risk and truly enable business speed.

That deep context becomes the defender’s advantage when you are using AI powered defenses, including those in AI Threat Defense.

Your teams should be looking at how they are using AI to accelerate security and respond to AI-driven threats at AI speed.

Consider technologies like AI Threat Defense as part of your defenses in this new world.

1 week, 1 day назад @ cloud.google.com
Do more with less: How GKE can reduce your cost per agent by 75%
Do more with less: How GKE can reduce your cost per agent by 75% Do more with less: How GKE can reduce your cost per agent by 75%

Optimization 1: Pushing density with GKE Agent SandboxTo address this, we migrated the same agent workload from microVMs to GKE Agent Sandbox, a Kubernetes primitive that’s designed specifically for the security and performance requirements of running agents.

Instead of relying on heavy guest operating systems, GKE Agent Sandbox leverages the open-source secure container sandbox, gVisor.

It’s no surprise then, that when GKE Agent Sandbox reached General Availability in May, its usage grew more than 7x in under four weeks.

Key takeaway: In our tests, migrating OpenClaw-type agents to GKE Agent Sandbox enabled us to run more than 40% more agents per vCPU, and reduced the cost per agent by mor…

1 week, 1 day назад @ cloud.google.com
What’s new in Gemini Enterprise Agent Platform
What’s new in Gemini Enterprise Agent Platform What’s new in Gemini Enterprise Agent Platform

Since we launched Gemini Enterprise Agent Platform a few months ago, we’ve seen inspiring progress from businesses and builders alike.

To stir up development, we’ve also shared 13 demos that can walk you through the versatility and power of Agent Platform, and 20 questions you can ask your teams about building a solid agentic foundation.

That’s why today, we are announcing some of our most popular capabilities are available for everyone, from Agent Runtime to Agent Identity.

We also recently just announced CodeMender, our new managed code security agent to help you advance from passive scanning to automated code remediation, and reduce zero-day risk.

Automate your long-running agents faster…

1 week, 2 days назад @ cloud.google.com
Automate your agent development lifecycle using any coding agent
Automate your agent development lifecycle using any coding agent Automate your agent development lifecycle using any coding agent

Welcome to our latest Gemini Enterprise Agent Platform deep dive, a practical walkthrough where we’ll teach you how to build real-world, production-ready agents starting from step 1.

With Agents CLI skills, you can go through the different phases of the entire agent lifecycle without ever leaving your coding agent.

What we’re building today: Industry Watch agentThis tutorial helps guide a developer on how to build a real Industry Watch agent, a sector-intelligence analyst for semiconductor stocks that reconciles what companies say in the press against what they file with the SEC.

We’ll walk through the six stages of building this agent end-to-end:Setup: Teach your coding assistant platform …

1 week, 2 days назад @ cloud.google.com
Bringing Conversational Analytics to your entire data ecosystem
Bringing Conversational Analytics to your entire data ecosystem Bringing Conversational Analytics to your entire data ecosystem

Over the last year, Conversational Analytics (CA) in Google Cloud has moved from isolated experiments to scaled, enterprise-wide deployments.

BigQuery Conversational Analytics and the Conversational Analytics API are now generally available, adding to the general availability of Conversational Analytics in Looker last year.

And so much more has happened — Google Cloud Conversational Analytics is available for more data, across more surfaces, with more enterprise controls, and greater capability than ever before.

Whether your data resides exclusively in Google Cloud or across multiple cloud providers, your agents can query it natively.

For data practitioners, Conversational Analytics is inte…

1 week, 3 days назад @ cloud.google.com
Best Buy scales AI workloads and secures access with Workforce Identity Federation
Best Buy scales AI workloads and secures access with Workforce Identity Federation Best Buy scales AI workloads and secures access with Workforce Identity Federation

The retailer solved both problems and paved the way for a massive cloud expansion by implementing Google Cloud's Workforce Identity Federation.

The team adopted Workforce Identity Federation to federate existing Entra ID identities directly into Google Cloud.

Now, when developers access BigQuery through Power BI, they authenticate as themselves using their existing Entra ID identity.

The architecture relies on two components working together: Entra ID handles authentication, Workforce Identity Federation brokers the trust relationship between Entra ID and Google Cloud.

ArchitectureThe diagram below shows how identity flows from Entra ID through the Workforce Identity Federation to the servi…

1 week, 3 days назад @ cloud.google.com
The Blueprint: How Voicify makes AI-enabled ordering a delight for customers
The Blueprint: How Voicify makes AI-enabled ordering a delight for customers The Blueprint: How Voicify makes AI-enabled ordering a delight for customers

Founded in 2018, Voicify reimagines the traditional phone call with the goal of transforming every call into a seamless and engaging experience.

We realized that specialized, purpose-driven AI assistants were the key to businesses maintaining excellent service at scale.

Under the hood, Gemini Flash, served via Gemini Enterprise Agent Platform, vastly improves latency, minimizing user wait times and preventing hang-ups.

To grow the business — and call volume — and to handle traffic spikes, we switched from Google AI Studio to Vertex AI and its current incarnation in Gemini Enterprise.

Specific Gemini Enterprise Agent Platform features help us manage high call volumes without experiencing ser…

2 weeks, 1 day назад @ cloud.google.com
Now in preview: Find and fix software vulnerabilities with CodeMender
Now in preview: Find and fix software vulnerabilities with CodeMender Now in preview: Find and fix software vulnerabilities with CodeMender

It examines and remediates existing code security issues without sacrificing development velocity by:Deploying the best-fit model .

Find and fix vulnerabilities with AIBorn from Google DeepMind's pioneering AI research, CodeMender transforms vulnerability management from a manual bottleneck into an autonomous, high-speed system.

CodeMender brings AI into a critical part of the security lifecycle by accelerating the path from validated vulnerability to tested fix.

"CodeMender consistently identified critical vulnerabilities that our other AI-enabled tools completely missed.

How the CodeMender agent worksWe’ve fine-tuned CodeMender’s harness to be continuously updated with the latest Google D…

2 weeks, 4 days назад @ cloud.google.com
OpenAI
последний пост None
Microsoft Microsoft
последний пост 4 days, 23 hours назад
Orchard: An open framework for scalable agentic AI
Orchard: An open framework for scalable agentic AI Orchard: An open framework for scalable agentic AI

At a glance Orchard is an open-source framework for scalable and cost-effective agentic AI research, built around Orchard Env, a reusable environment service for training and evaluating agents across task domains.

To address this gap, we introduce Orchard (opens in new tab), an open-source framework for scalable agentic modeling.

Unlike many existing frameworks, Orchard Env is designed to support different agent systems and task types without modification.

(opens in new tab) We are also releasing the training data and evaluation methods used to build them.

By making the underlying infrastructure open, lightweight, and reusable, Orchard lowers the cost of agentic AI research.

4 days, 23 hours назад @ microsoft.com
Echoverse: Deep, evolving environments for computer-use agents
Echoverse: Deep, evolving environments for computer-use agents Echoverse: Deep, evolving environments for computer-use agents

At a glance We built twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds, each drilling a single control rendered in many forms (date pickers and nested filters).

Shallow worlds backfire; deep worlds transferA shallow world is the cheap option.

A deep world costs more, but its trajectories carry the dependent structure that transfers to the live site.

Only the deep world improves both, lifting Allrecipes to 85.0% and the harder Hugging Face split to 65.0%.

First, more deep worlds for the closed domains public benchmarks cannot reach.

1 week, 1 day назад @ microsoft.com
EvoLib: Turning experience into evolving knowledge
EvoLib: Turning experience into evolving knowledge EvoLib: Turning experience into evolving knowledge

By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.

By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.

Rather than treating memory as a growing archive of past experiences, EvoLib extracts reusable knowledge from those experiences and continually refines it as new experiences arrive.

As new knowledge is extracted from recent experience, EvoLib retrieves similar knowledge from the library and tries to co…

1 week, 1 day назад @ microsoft.com
Verifying Rust cryptography in SymCrypt, from standards to code
Verifying Rust cryptography in SymCrypt, from standards to code Verifying Rust cryptography in SymCrypt, from standards to code

Aeneas allows verifying a large subset of Rust code and provides efficient automation in Lean to support the proof effort.

SymCrypt is extending the same Rust, Lean, and Aeneas-based workflow to more Rust-native algorithms and integrating them into production versions for Windows and Linux, including for instance verified Rust code for, e.g., AES-GCM, FrodoKEM, and ML-DSA.

The Rust code and the proofs live side by side, but the proof burden does not shape the code into something unnatural.

Others can be modelled using Rust code, which can be tested against hardware reference documentation, then translated and verified.

This is particularly powerful because the Rust code and Lean proofs are …

3 weeks, 4 days назад @ microsoft.com
Aurora 1.5: Extending open foundation models for weather and Earth-system applications
Aurora 1.5: Extending open foundation models for weather and Earth-system applications Aurora 1.5: Extending open foundation models for weather and Earth-system applications

Aurora 1.5 connects open research to Microsoft Weather services, linking the model with data, infrastructure, managed access, and operational use for weather and Earth-system applications.

Aurora 1.5 is a major update to the open Aurora Earth-system foundation model, adding 22 new weather variables for a broader view of atmospheric conditions, hourly forecasts, and probabilistic ensemble forecasting.

Aurora 1.5 advances the broader effort to make open weather foundation models practical and scalable for organizations that rely on atmospheric and Earth-system intelligence.

Figure 1: Illustration of the capabilities of Aurora 1.5 ensemble for predicting new impactful parameters such as total …

4 weeks, 1 day назад @ microsoft.com
Flint: A visualization language for the AI era
Flint: A visualization language for the AI era Flint: A visualization language for the AI era

Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.. Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.

They help the compiler choose appropriate scales, baselines, formatting, and color schemes.. Flint leverages semantic data types to express meanings of data.

To address this challenge, we introduce Flint (opens in new tab), a visualization intermediate language for AI-driven chart creation.

Flint compiles a compact, human-editable chart specification into a complete backend-native specification and rendered visualization.

How Flint worksFigure …

1 month назад @ microsoft.com
SkillOpt: Agent skills as trainable parameters
SkillOpt: Agent skills as trainable parameters SkillOpt: Agent skills as trainable parameters

SkillOpt treats an agent skill file as a trainable parameter outside a frozen target model, turning skill writing from one-shot prompting into a controlled optimization process.

SkillOpt keeps skills compact and auditable through bounded text edits, validation gating, rejected-edit feedback, and slow/meta updates, avoiding uncontrolled prompt drift.

The optimized skills transfer across model scales, agent harnesses, and related tasks, suggesting that they capture reusable workflow knowledge rather than benchmark-specific instructions.

Today, agent skills typically come from three sources: experts write them by hand, a frontier model generates them one-shot, or the agent loosely revises them…

1 month, 1 week назад @ microsoft.com
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity

Memora is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling what is stored (rich memory content) from how it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

is a scalable memory system that dramatically increases agent productivity on long-horizon tasks by decoupling is stored (rich memory content) from it’s retrieved (lightweight abstractions and cue anchors), balancing abstraction and specificity.

Why this is hard: the abstraction–specificity tensionExisting memory systems fall into two extremes.

None of these resolves the underlying tension between abstraction (which keeps memory effi…

1 month, 1 week назад @ microsoft.com
Understanding the brain with AI-driven explanations and experiments
Understanding the brain with AI-driven explanations and experiments Understanding the brain with AI-driven explanations and experiments

As black-box models spread, the gap between prediction and understanding has become one of the central problems in computational neuroscience.

GCT distills brain-prediction models into short, readable accounts of what each patch of cortex responds to, then tests those claims.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

To build trust in the explanation, GCT uses an LLM to write new stories in w…

1 month, 1 week назад @ microsoft.com
Understanding the brain with AI-driven explanations and experiments
Understanding the brain with AI-driven explanations and experiments Understanding the brain with AI-driven explanations and experiments

As black-box models spread, the gap between prediction and understanding has become one of the central problems in computational neuroscience.

GCT distills brain-prediction models into short, readable accounts of what each patch of cortex responds to, then tests those claims.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

An LLM writes new stories engineered to activate a specific brain area, subjects hear them in the scanner, and if the explanation is correct, the targeted region lights up.

To build trust in the explanation, GCT uses an LLM to write new stories in w…

1 month, 1 week назад @ microsoft.com
Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis
Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis

At a glance Talos is an open-source tool for automated, iterative reanalysis of genomic data in rare disease.

Deployed across a prospective cohort of almost 5,000 undiagnosed patients, Talos delivered 241 new diagnoses (5.1% additional yield).

On monthly iterative cycles, analysts only needed to review one new variant per 200 patients, demonstrating that frequent, systematic reanalysis can be run sustainably.

Why genome reanalysis mattersGenomic testing has transformed the diagnosis of rare disease, but even with this advancement, more than half of patients remain undiagnosed after their first test.

Looking aheadTalos reframes genomic reanalysis from a rare, labor-intensive event into a con…

1 month, 2 weeks назад @ microsoft.com
Ire identifies another LOTUSLITE specimen
Ire identifies another LOTUSLITE specimen Ire identifies another LOTUSLITE specimen

At a glance Project Ire identifies a LOTUSLITE variant that shares TTPs (tools, tactics, procedures) with the public family but none of its indicators of compromise (IOC).

On Ire’s calibrationOne noteworthy observation in Ire’s report (opens in new tab) is worth highlighting first.

The Ire report does not surface a matching entry-point name, but it identifies that the behavioral shape is the same.

Ire never named LOTUSLITE in its report or chain of evidence.

Ire described the behavior precisely enough to make the mapping straightforward of this sample to LOTUSLITE.

1 month, 3 weeks назад @ microsoft.com
Data Formulator 0.7: AI-powered data analytics for enterprise data
Data Formulator 0.7: AI-powered data analytics for enterprise data Data Formulator 0.7: AI-powered data analytics for enterprise data

At a glance Data Formulator 0.7 is an open-source AI-powered system for enterprise data analytics that combines data connectivity, agent-guided exploration, and visualization refinement in a shared workspace.

Enterprise teams increasingly rely on AI systems for analytics, but enterprise data workflows are often fragmented across storage systems and tools.

Listen now Opens in a new tabConnecting enterprise data with Data ConnectorsData Formulator helps teams bring enterprise data into an AI-ready workspace without needing to rebuild the same connections for every source of data.

Data Connectors provide persistent connections between enterprise data sources and Data Formulator, allowing analy…

2 months, 1 week назад @ microsoft.com
Extending Human Intelligence Through AI
Extending Human Intelligence Through AI Extending Human Intelligence Through AI

At a glance Modern AI systems are powerful not because they replicate human intelligence, but because they presuppose it, by extending structures already present in human cognition and language.

Understanding AI as an extension of human intelligence—not a replacement for it—offers a more grounded path for building trustworthy AI systems.

Rather than asking whether AI systems are becoming intelligent in the human sense, these approaches ask a more basic question: What if AI systems work because they rely on structures that are rooted in human cognition?

In our recent paper, The Origins of Artificial Intelligence in Natural Intelligence, we argue that modern AI systems are best understood nei…

2 months, 1 week назад @ microsoft.com
MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models
MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models

Built as the next generation of Magentic-UI, it combines a redesigned app with a harness optimized for small models.

MagenticBrain and Fara1.5 are small models designed for orchestration and computer-use tasks, respectively.

Together, these releases explore how far agentic performance can be pushed with smaller models, codesigned tools, and an optimized execution harness.

Today, Microsoft Research AI Frontiers releases MagenticLite (opens in new tab), an experimental agentic application designed for small models.

The result is an agent that runs efficiently, keeps data on the user’s machine, and supports a broad range of agentic tasks.

2 months, 2 weeks назад @ microsoft.com
MIT AI MIT AI
последний пост 3 days, 20 hours назад
Solving the solvent problem
Solving the solvent problem Solving the solvent problem

“It’s supposed to be an ion conductor.” But unfortunately, most electrolytes get involved in unwanted chemical reactions with the electrodes, which can greatly undermine battery stability.

The team’s goal, accordingly, was to identify solvent molecules that are small enough to improve ion transport while still maintaining electrolyte stability.

There is, however, a complicating factor — a trade-off to be addressed: Faster ion transport often comes at the expense of electrolyte stability.

By carefully tailoring the size of solvent molecules, the authors demonstrate a new design strategy that could enable lower-cost, higher performance batteries.”The group is not done.

The overriding goal of …

3 days, 20 hours назад @ news.mit.edu
The benefits of medical AI assistance vary based on user expertise
The benefits of medical AI assistance vary based on user expertise The benefits of medical AI assistance vary based on user expertise

Explainable AI methods help users know when to trust a model’s predictions by describing or validating the model’s decision-making.

By contrast, clinicians were not tripped up by incorrect AI assistance and performed best when given only a model’s prediction, with no accompanying explanation.

“Good AI systems can improve performance in some health settings, but this has to be balanced carefully with algorithmic deference that can lead to more error.

They tested users by showing them medical images plus an AI prediction of skin disease, employing different explainable AI approaches.

We were just able to train very good AI models for this setting,” Ghassemi says.

4 days, 6 hours назад @ news.mit.edu
Alexander Rakhlin named director of the MIT Statistics and Data Science Center
Alexander Rakhlin named director of  the MIT Statistics and Data Science Center Alexander Rakhlin named director of the MIT Statistics and Data Science Center

Alexander “Sasha” Rakhlin PhD ’06, the Distinguished Professor in Data, Systems, and Society at the MIT Institute for Data, Systems, and Society (IDSS); and a professor of brain and cognitive sciences at MIT, has been named the next director of the MIT Statistics and Data Science Center (SDSC).

“The strength of the Statistics and Data Science Center has always been its people — students, postdocs, and faculty from across MIT who bring sharply different perspectives to the most interesting problems of the day in statistics, machine learning, and AI.

“At the Statistics and Data Science Center, I work alongside colleagues who share this fascination and pursue these connections in many directio…

4 days, 19 hours назад @ news.mit.edu
Daniela Rus receives Bavarian Minister-President's High-Tech Prize
Daniela Rus receives Bavarian Minister-President's High-Tech Prize Daniela Rus receives Bavarian Minister-President's High-Tech Prize

Daniela Rus, director of MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL) and the Panasonic Professor of Computer Science, has received the 2026 High-Tech Prize of the Bavarian Minister-President for her contributions to robotics, artificial intelligence, and autonomous systems.

The selection committee cited four strands of her work: self-organizing robot collectives, soft robotics, autonomous mobility, and brain-inspired artificial intelligence.

She is also a pioneer of soft robotics, where compliant machines manipulate the world more safely and adapt to it more readily than rigid ones can.

"Daniela Rus is a pioneer in soft robotics and physical AI," noted Lorenzo Masi…

1 week, 1 day назад @ news.mit.edu
Connecting research to policy on Capitol Hill
Connecting research to policy on Capitol Hill Connecting research to policy on Capitol Hill

This spring, 25 MIT students and postdocs traveled to Washington to meet with congressional staffers and advocate for sustained federal investment in scientific research.

With recent cuts to National Science Foundation programs and continued uncertainty surrounding the federal research budget, these conversations were especially timely.

Over the course of just two days, participants met with 62 congressional offices representing 32 states to discuss the importance of federal support for scientific research, higher education, and other policy concerns related to their individual research areas.

To prepare for the trip, participants attended three training sessions led by SPI in collaboration…

1 week, 1 day назад @ news.mit.edu
How a medical database developed at MIT evolved into a global standard of data-sharing
How a medical database developed at MIT evolved into a global standard of data-sharing How a medical database developed at MIT evolved into a global standard of data-sharing

Before the advancement of scientific data storage and collaboration via the cloud, medical investigators seeking health research breakthroughs had to overcome significant obstacles to collaboration and key clinical data gathering.

The data eventually became the first database of the global platform PhysioNet — founded in 1999 at the Harvard-MIT program in Health Sciences and Technology — as a clinical data repository for complex physiological signals.

In the years since PhysioNet was established, the value of sharing research data has gained much wider recognition.

Pollard points to similar platforms like Health Data Nexus as examples of PhysioNet’s legacy.

In addition to using PhysioNet da…

1 week, 3 days назад @ news.mit.edu
Working to automate nuclear plant operations
Working to automate nuclear plant operations Working to automate nuclear plant operations

In pursuit of autonomous nuclear plant operationsIt turns out the research for the master’s was just the tip of the iceberg.

For the future viability of nuclear power, small plants, located in rural areas, are a distinct possibility.

It’s where supervised and thoroughly vetted autonomous operations will help.

A primary question was: “How do we transition to autonomous operations in nuclear power plants?” Fortier wanted one integrated approach, a central supervisory control system instead of many interlinked parts.

Using the nuclear plant automation program on next-generation equipment will deliver necessary traction in developing and deploying commercial microreactors.

2 weeks, 1 day назад @ news.mit.edu
MIT projects selected for funding under US Department of Energy’s Genesis Mission
MIT projects selected for funding under US Department of Energy’s Genesis Mission MIT projects selected for funding under US Department of Energy’s Genesis Mission

MIT researchers are set to contribute to the U.S. Department of Energy’s (DOE) Genesis Mission, with 15 collaborative projects among those selected for funding under Genesis Phase I, DOE announced Wednesday.

“MIT researchers are proud to be leading and contributing to projects under the Genesis Mission, in vital areas of research that support national priorities,” says Ian A. Waitz, MIT’s vice president for research.

Projects under the Genesis Mission are collaborative by design; teams must draw on the expertise of researchers from academia, industry, and/or the national laboratories.

Phase I projects that identify promising pathways toward transformative capabilities at scale may be consid…

2 weeks, 2 days назад @ news.mit.edu
Professor Emeritus Dimitri Bertsekas, influential computer scientist and prolific author, dies at 83
Professor Emeritus Dimitri Bertsekas, influential computer scientist and prolific author, dies at 83 Professor Emeritus Dimitri Bertsekas, influential computer scientist and prolific author, dies at 83

Over the course of his career, Bertsekas’ research spanned, and had a definitive influence upon, several fields, including optimization, control, large-scale computation, reinforcement learning, and artificial intelligence.

Along the way, Bertsekas taught, advised, and mentored students who would eventually become his colleagues at all four institutions.

“Dimitri played a defining role in my career,” says Asu Ozdaglar, department head of EECS at MIT.

Some referred to Dimitri as an “immortal.” Another comment I recall fondly — and often reminded Dimitri about — was: “Professor Bertsekas is a very handsome man!” Their bond continued long after Van Roy’s graduation.

He is survived by his wife …

2 weeks, 2 days назад @ news.mit.edu
Following the questions where they lead
Following the questions where they lead Following the questions where they lead

Ever since she was a child playing on her family’s farmland in Wisconsin, Bailey Flanigan was guided by her own selective, yet wide-ranging, curiosity.

“I found myself unmotivated to take all the AP [advanced placement] classes for the sake of it.

So Flanigan moved toward public health, where she researched microfluidic devices for HIV detection that could be used in low-resource settings.

After graduating from UW-Madison, Flanigan worked as a predoctoral research assistant in economics at Princeton.

“I feel so lucky to be studying these questions from within both political science and EECS, because I have the freedom to explore both the political and technical substance of tools for more d…

3 weeks назад @ news.mit.edu
A better way to turn 2D designs into 3D models for rapid prototyping
A better way to turn 2D designs into 3D models for rapid prototyping A better way to turn 2D designs into 3D models for rapid prototyping

The system generates new data based on the model’s abilities as it attempts to convert a 2D image into a CAD program.

“Nearly every physical product around us, from airplanes to appliances, begins its life as a CAD model.

For guesses that are nearly correct, GIFT adjusts them to become successful solutions.

The CAD models generated by VLMs using GIFT were better aligned with the shapes of ground-truth models.

In the future, the researchers want to expand GIFT so the framework can teach models to generate CAD programs that improve the performance and manufacturability of 3D models.

3 weeks, 2 days назад @ news.mit.edu
3 Questions: Neural transparency and the future of AI design
3 Questions: Neural transparency and the future of AI design 3 Questions: Neural transparency and the future of AI design

Q: Your paper introduces “neural transparency,” a way to let everyday users peek inside an AI’s neural networks before their chatbot ever says a word.

“Neural transparency” means giving people something like a brain scan for AI.

Our study suggests that people have a blind spot when designing personalized AI.

In previous research, we documented cases of psychological harm associated with interactions with AI chatbots.

AI companions are dynamic systems that evolve as they interact with us, so understanding those internal changes is an important next step.

3 weeks, 2 days назад @ news.mit.edu
Helping AI models to meet the real world
Helping AI models to meet the real world Helping AI models to meet the real world

“In a sense, with a small amount of resource, you have to do a lot of heavy lifting,” he says.

“My interest was: How does one design such graphical models for generic, tabular data?” he says.

And each of the products that you manufacture has lots of small pieces that come from different parts of the world.

Shah adds that Celonis has specialized in digitizing and automating operations for more than 1,400 large companies around the world.

“A narrower focus comes with sharper technology,” he says, “but it’s broad enough that it’s very valuable.”Shah adds, “The recent buzzword that’s become pertinent in the modern AI popular press is a ‘world model.’ In a sense, this is trying to build the ente…

3 weeks, 3 days назад @ news.mit.edu
Can AI build a jet engine? JARVIS Challenge tests role of AI copilots in tough-tech engineering
Can AI build a jet engine? JARVIS Challenge tests role of AI copilots in tough-tech engineering Can AI build a jet engine? JARVIS Challenge tests role of AI copilots in tough-tech engineering

“The JARVIS challenge showed that AI can substantially accelerate safety-critical hardware engineering, but engineering judgment remains the decisive differentiator.

Manufacturing — not engineering design or analysis — remained the fundamental rate-limiting step,” says Professor Zolti Spakovszky, director of the MIT Gas Turbine Laboratory.

In weekly progress reviews, they would critically evaluate the student progress and assess how the students were using AI.

The 811 team had been resistant to using AI throughout the competition, trusting instead to their fundamentals and teamwork.

From the start of the JARVIS Challenge, younger students used Parley more frequently and cleverly, while the …

3 weeks, 3 days назад @ news.mit.edu
How MIT students are helping to prevent cyberattacks
How MIT students are helping to prevent cyberattacks How MIT students are helping to prevent cyberattacks

To counter such threats, Lecturer Jungwoo Chun and Ford Professor of Urban and Environmental Planning Lawrence Susskind launched the MIT Cybersecurity Clinic in 2019.

Much like a legal or medical clinic, the course doubles as hands-on training for students and a pro-bono service to at-risk communities.

After completing instructional modules and passing a certification exam, students are assigned in teams to a client.

The Cybersecurity Clinic aims to round out the knowledge of students from every discipline.

In either case, Susskind and Chun check in periodically with clients for at least two years following each engagement.

3 weeks, 4 days назад @ news.mit.edu
Berkeley AI
последний пост 1 week, 3 days назад
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple SiliconFigure 1: CUDA-to-MLX optimization translation map.

Although we focus on MLX kernels for Apple Silicon, the method is not specific to MLX and applies to any ecosystem where CUDA expertise is transferable.

With Apple Silicon in hundreds of millions of MacBooks and Mac Studios, MLX enables local AI inference without cloud costs.

Building an MLX backendTo bring K-Search to Apple Silicon, we first built a native MLX backend.

Evaluated on mamba-370m f16, M1 Max 64GB:Metric mlx-mamba (ours) mlx-lm (community) mamba.py Decode 152 tok/s 116 tok/s 40 tok/s Prefill L=512 5,751 tok/s 329 tok/s 1,089 tok/s Prefill L=1024 …

1 week, 3 days назад @ bair.berkeley.edu
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

Teaching LLMs to Update Beliefs for Efficient Long-Horizon InteractionOverview of ABBEL compared to traditional recursive summarization.

Belief grading can be thought of as adding an auxiliary RL task, using heuristics designed to capture what makes a good belief as the reward.

With domain-knowledge belief grading, ABBEL approaches or exceeds FULL CTX in this setting; without belief grading, learning is slower.

Context compression methods generate dense representations which, while computationally efficient, sacrifice human-understandability (Kontonis et al., 2026, Eyuboglu et al., 2025, Gupta et al., 2025, Chevalier et al., 2023, Deng et al., 2025, Deng et al., 2025, Bulatov et al., 2022).…

1 week, 6 days назад @ bair.berkeley.edu
Intelligence is Free, Now What? Data Systems for, of, and by Agents
Intelligence is Free, Now What?  Data Systems for, of, and by Agents Intelligence is Free, Now What? Data Systems for, of, and by Agents

Agents are rapidly becoming capable of synthesizing entire data systems in one go—meaning we can rebuild custom systems for each new workload.

Data Systems For, Of, and By AgentsNext, we will discuss each in more detail, followed by discussing the intertwined future of data systems and agents, especially as the three challenges intersect.

Data Systems Of AgentsPreviously, we focused on how agents interact with data systems.

Data Systems By AgentsFinally, if intelligence is effectively free, then we can employ this intelligence to synthesize new data systems from scratch.

Co-Evolution of Data Systems and AgentsLooking further out, the boundaries between agents and data systems will likely …

1 month назад @ bair.berkeley.edu
2026 BAIR Graduate Showcase
2026 BAIR Graduate Showcase 2026 BAIR Graduate Showcase

2026 BAIR Graduate ShowcaseCongratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026!

This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed the frontiers of artificial intelligence and machine learning.

Their work spans the breadth of modern AI — robotics and embodied intelligence, large language models and reasoning, computer vision, generative modeling, AI safety, human-AI interaction, AI for science and healthcare, and much more.

Along the way, they have published influential research, built systems with real-world impact, mentored their peers, and shaped the BAIR community for th…

1 month, 1 week назад @ bair.berkeley.edu
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling

Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference ScalingOverview of adaptive parallel reasoning.

We provide a detailed analysis of recent progress in the field of parallel reasoning, especially Adaptive Parallel Reasoning.

Figure 4: Special Tokens Variants across Adaptive Parallel Reasoning PapersInference Systems for Adaptive ParallelismHow do we actually execute parallel branches?

Figure 14: Difference in Model Choice Across Adaptive Parallel Reasoning PapersEach paper also offers a slightly different interpretation about how adaptive parallel reasoning contributes to the research field.

(Yang et al., 2025; Lian et al., 2025) aim to deliver sequential-AR-model-level a…

3 months назад @ bair.berkeley.edu
Gradient-based Planning for World Models at Longer Horizons
Gradient-based Planning for World Models at Longer Horizons Gradient-based Planning for World Models at Longer Horizons

Large, learned world models are becoming increasingly capable.

Why is adversarial robustness an issue for world model planning?

We thus exploit the differentiability of learned world models $F_{\theta}$, while not falling victim to the inherent sensitivity of the state Jacobians $D_s F_{\theta}$.

It’s a funny sweet spot where the background literature (planning and control overall) is incredibly mature and well-developed, but the current setting (pure planning optimization over modern, large-scale world models) is still heavily underexplored.

But, once we figure out all the right ideas, world model planners will likely become as commonplace as RL.

3 months, 2 weeks назад @ bair.berkeley.edu
Identifying Interactions at Scale for LLMs
Identifying Interactions at Scale for LLMs Identifying Interactions at Scale for LLMs

Identifying Interactions at Scale for LLMsUnderstanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial intelligence.

Therefore, grounded or reality-checked interpretability methods must also be able to capture these influential interactions.

In this blog post, we describe the fundamental ideas behind SPEX and ProxySPEX, algorithms capable of identifying these critical interactions at scale.

SPEX and ProxySPEX FrameworkTo discover influential interactions with a tractable number of ablations, we have developed SPEX (Spectral Explainer).

We formalize this through two observations: sparsity (relatively f…

4 months, 4 weeks назад @ bair.berkeley.edu
Information-Driven Design of Imaging Systems
Information-Driven Design of Imaging Systems Information-Driven Design of Imaging Systems

We developed a framework that enables direct evaluation and optimization of imaging systems based on their information content.

The first approach treated imaging systems as unconstrained communication channels, ignoring the physical limitations of lenses and sensors.

Our Information-Driven Encoder Analysis Learning (IDEAL) method uses gradient ascent on information estimates to optimize imaging system parameters.

The standard approach to computational imaging design, end-to-end optimization, jointly trains the imaging hardware and a neural network decoder.

The computational efficiency of IDEAL suggests possibilities for designing imaging systems that were previously intractable.

7 months назад @ bair.berkeley.edu
RL without TD learning
RL without TD learning RL without TD learning

RL without TD learningIn this post, I’ll introduce a reinforcement learning (RL) algorithm based on an “alternative” paradigm: divide and conquer.

We can do Reinforcement Learning (RL) based on divide and conquer, instead of temporal difference (TD) learning.

There are two classes of algorithms in RL: on-policy RL and off-policy RL.

We compared TRL with $n$-step TD learning with different values of $n$, from $1$ (pure TD) to $\infty$ (pure MC).

I still think one of the most important problems in RL (and even in machine learning) is to find a scalable off-policy RL algorithm.

9 months, 1 week назад @ bair.berkeley.edu
AWS Machine Learning AWS Machine Learning
последний пост 22 часа назад
How Cohere Health digitizes clinical policies using Amazon Bedrock AgentCore
How Cohere Health digitizes clinical policies using Amazon Bedrock AgentCore How Cohere Health digitizes clinical policies using Amazon Bedrock AgentCore

To scale out representations in Cohere Policy Studio, Cohere Health added new skills to an existing AgentCore Runtime that was already decomposing policies.

The team completed three tasks:Deployed AgentCore Runtime with AgentCore Gateway and AgentCore Memory for a full agentic system using LangChain.

Wrote skills with clinical policy experts and evaluated them using Cohere Health’s standardized observability process based on Arize AI.

AgentCore Gateway architectureCohere Health implemented this using AgentCore Gateway with separate targets for shared tools and project-specific tools.

Cohere Health has digitized thousands of policies to date using manual and semi-automated workflows.

22 часа назад @ aws.amazon.com
How TReNDS automates root-cause analysis with Amazon Bedrock
How TReNDS automates root-cause analysis with Amazon Bedrock How TReNDS automates root-cause analysis with Amazon Bedrock

When we started exploring Amazon Bedrock, we saw an opportunity we had wanted for a long time.

We use the Strands Agents SDK on top of Amazon Bedrock to handle tool-use orchestration.

PrerequisitesTo implement this solution, you need the following:An AWS account with access to Amazon Bedrock (specifically Anthropic Claude Sonnet).

An AWS Lambda function with appropriate IAM permissions to access Amazon Bedrock, CloudWatch Logs, AWS Secrets Manager, and SNS.

Choosing the right Amazon Bedrock modelAmazon Bedrock gives us access to a range of foundation models through a single API.

22 часа назад @ aws.amazon.com
Determining playoff clinching scenarios in the NHL using constraint programming
Determining playoff clinching scenarios in the NHL using constraint programming Determining playoff clinching scenarios in the NHL using constraint programming

In this post, we describe how the AWS Generative AI Innovation Center created an automated system that determines NHL playoff clinching scenarios.

These complex tie-breakers are a major reason why determining playoff clinch scenarios is computationally challenging.

Determining 1-day clinch scenarios required a median runtime on the order of minutes, offering significant speedups over manual approaches (see Figure 2).

Automated, provably correct clinching scenarios that can be generated daily without manual effort, yielding significant time savings.

ConclusionDetermining NHL playoff clinching scenarios is a computationally hard problem.

22 часа назад @ aws.amazon.com
Securing AI agents with temporal policies in Amazon Bedrock AgentCore
Securing AI agents with temporal policies in Amazon Bedrock AgentCore Securing AI agents with temporal policies in Amazon Bedrock AgentCore

Temporal policies in Amazon Bedrock AgentCore let you define stateful rules that determine authorization to AgentCore Gateway targets by evaluating the current request in the context of prior events in an agent’s trajectory.

In this post, you will learn what temporal policies are, how they work, and walk through an example to demonstrate.

As with the existing AgentCore Policy features, temporal policies deny by default and forbid wins over permit.

Applying temporal policies to a private banking portfolio agentTo make these concepts concrete, we’ll walk through how temporal policies can secure a hypothetical private banking agent.

To learn about AgentCore Gateway and how to set up auth with …

1 day, 20 hours назад @ aws.amazon.com
Configure rate limits for AI traffic on AgentCore gateway
Configure rate limits for AI traffic on AgentCore gateway Configure rate limits for AI traffic on AgentCore gateway

Amazon Bedrock AgentCore gateway is a fully managed, serverless AI gateway that provides a single, secure entry point for AI traffic.

Rate limit structureA rate limit configuration consists of two parts: dimension keys and entries.

Types of rate limits and example configurationsAgentCore gateway enforces two layers of rate limiting: customer-defined rate limits and Service Quotas.

We demonstrated how to layer rate limits: user-level limits for per-role and per-user fairness, target-level limits for protecting downstream service capacity, and multi-dimensional limits that combine user and target to enforce fine-grained rate limits.

To get started with rate limits on AgentCore gateway, explor…

1 day, 21 hours назад @ aws.amazon.com
Control agent behaviors and cost beyond a single action: new capabilities in Amazon Bedrock AgentCore
Control agent behaviors and cost beyond a single action: new capabilities in Amazon Bedrock AgentCore Control agent behaviors and cost beyond a single action: new capabilities in Amazon Bedrock AgentCore

According to McKinsey, roughly 80% of organizations have already encountered risky behavior from AI agents.

As a result, security and risk concerns are the leading barrier to scaling agentic AI (McKinsey’s State of AI Trust in 2026, and Trust in the age of AI agents 2026).

We built Amazon Bedrock AgentCore to give teams what they need to build, connect, and optimize agents at scale without assembling the infrastructure themselves.

Powering temporal policies is Dogwood, a new policy language purpose-built for AI agents.

Built on the foundation of Cedar, Dogwood was designed to address a new dimension of agent control: evaluating whether a sequence of agent actions conforms to a policy as it …

1 day, 22 hours назад @ aws.amazon.com
Build visibility for Codex on Amazon Bedrock with OpenTelemetry and Amazon CloudWatch
Build visibility for Codex on Amazon Bedrock with OpenTelemetry and Amazon CloudWatch Build visibility for Codex on Amazon Bedrock with OpenTelemetry and Amazon CloudWatch

The reference deployment creates a CloudWatch dashboard, not an Amazon Elastic Container Service (Amazon ECS) service, load balancer, virtual private cloud (VPC), or public ingestion endpoint.

Enable the CloudWatch OTel capabilitiesThe reference runbook begins by enabling OTel enrichment and resource tags for telemetry in the target Region:aws cloudwatch start-otel-enrichment --region us-west-2 aws observabilityadmin start-telemetry-enrichment --region us-west-2 aws cloudwatch get-otel-enrichment --region us-west-2The current CloudWatch OTel documentation describes native OTLP ingestion and PromQL querying.

CloudWatch OTel metrics use per-gigabyte ingestion pricing, and PromQL queries are p…

1 day, 22 hours назад @ aws.amazon.com
Enforcing data residency with single-Region Claude Code on Amazon Bedrock
Enforcing data residency with single-Region Claude Code on Amazon Bedrock Enforcing data residency with single-Region Claude Code on Amazon Bedrock

A US-headquartered global organization recently came to us with a deceptively simple data-residency request: let their engineers use Claude Code.

The requirement: Amazon Bedrock model inference had to be processed in London (the eu-west-2 AWS Region), not merely called from London.

Classic Amazon Bedrock (bedrock-runtime) is the original Amazon Bedrock Invoke API.

Verifying single-Region complianceEvery Claude Code Amazon Bedrock call must appear in the target Region’s AWS CloudTrail, and nowhere else.

Learn more about inference profiles, cross-Region inference, and Claude Code on Amazon Bedrock in their respective documentation.

1 day, 22 hours назад @ aws.amazon.com
Agent Skills for Automated Reasoning policies in Amazon Bedrock
Agent Skills for Automated Reasoning policies in Amazon Bedrock Agent Skills for Automated Reasoning policies in Amazon Bedrock

Teams that adopt Amazon Bedrock Automated Reasoning checks often want to run the policy lifecycle in code.

Authoring a good Automated Reasoning policy has a learning curve, and the lifecycle has constraints that can trip you up.

In Build reliable AI systems with Automated Reasoning on Amazon Bedrock, we walked through this loop in the Amazon Bedrock console.

In this post, you learn how to use a suite of Agent Skills to build, test, deploy, and validate an Amazon Bedrock Automated Reasoning policy end to end from a coding agent.

Six skills across the policy lifecycleThe suite is a set of six Agent Skills, one for each stage of the lifecycle.

1 day, 22 hours назад @ aws.amazon.com
Building an agentic app deployer with Amazon Bedrock and AWS Lambda
Building an agentic app deployer with Amazon Bedrock and AWS Lambda Building an agentic app deployer with Amazon Bedrock and AWS Lambda

PrerequisitesTo understand this architecture, familiarity with the following AWS services is helpful: AWS Lambda, Amazon API Gateway, Amazon DynamoDB, Amazon Simple Storage Service (Amazon S3), Amazon CloudFront, and Amazon Bedrock.

How the provisioning agent worksThe provisioning agent performs two core functions: classifying workloads and orchestrating long-running provisioning steps.

First, a mandatory Amazon Bedrock Guardrail handles personally identifiable information (PII) redaction, content filtering, and prompt-injection defense on every invocation.

The stack uses AWS Lambda, Amazon API Gateway, Amazon DynamoDB, Amazon S3, and Amazon CloudFront as the building blocks of a platform-a…

1 day, 22 hours назад @ aws.amazon.com
LLM optimization integration for Amazon SageMaker Python SDK
LLM optimization integration for Amazon SageMaker Python SDK LLM optimization integration for Amazon SageMaker Python SDK

The Amazon SageMaker Python SDK v3 now exposes generative AI inference recommendations in Amazon SageMaker AI directly in your notebook workflow.

Benefits of generative AI inference recommendations in Amazon SageMaker AIGenerative AI inference recommendations in Amazon SageMaker AI automate inference optimization by:Benchmarking a live Amazon SageMaker endpoint against a synthetic or real-traffic workload, measuring throughput, time-to-first-token (TTFT), end-to-end latency, and more.

Previously, these capabilities required using Amazon SageMaker Studio or constructing AWS SDK for Python (Boto3) API calls.

With the Amazon SageMaker Python SDK integration, you can automate this entire workfl…

1 day, 23 hours назад @ aws.amazon.com
How LendingTree built a multi-agent mortgage assistant on Amazon Bedrock
How LendingTree built a multi-agent mortgage assistant on Amazon Bedrock How LendingTree built a multi-agent mortgage assistant on Amazon Bedrock

Buying a home is one of the biggest financial decisions most people face, and LendingTree built a multi-agent mortgage assistant on Amazon Bedrock to make the process more straightforward.

LendingTree chose Amazon Bedrock for its multi-model flexibility and inherited AWS governance controls, which their compliance team required.

The solution was deployed on Amazon ECS instead of Amazon Bedrock AgentCore because it was already in production when AgentCore reached general availability.

Amazon Bedrock (Nova Pro and Nova Lite, Knowledge Bases, and Guardrails) provided the model and safety foundation.

To get started with multi-agent systems on AWS, explore the Amazon Bedrock documentation and Am…

2 days, 20 hours назад @ aws.amazon.com
How Mobileye transformed support operations using Amazon Bedrock AgentCore
How Mobileye transformed support operations using Amazon Bedrock AgentCore How Mobileye transformed support operations using Amazon Bedrock AgentCore

Mobileye’s Data Collection Processing pipeline ingests thousands of drive-recording sessions daily, generating a constant stream of status inquiries from engineers and data teams.

Using Amazon Bedrock AgentCore, Mobileye deployed an AI Support Agent that cut response times by 90% and exceeded 95% accuracy targets, with zero infrastructure overhead.

The team chose to build an AI agent capable of understanding context and adapting to diverse inquiry patterns.

Evaluating the optionsAfter successfully proving the concept with their AI agent, Mobileye needed to take it to production.

Component Role AgentCore Runtime Mobileye runs their AI Support Agent on AgentCore’s serverless runtime, eliminat…

2 days, 21 hours назад @ aws.amazon.com
How we built an MCP bridge to give our AgentCore-hosted AI agent access to local MCP tools
How we built an MCP bridge to give our AgentCore-hosted AI agent access to local MCP tools How we built an MCP bridge to give our AgentCore-hosted AI agent access to local MCP tools

We bridge the gap between the remote MCP client and the local MCP server by tunneling MCP messages over WebSocket and native messaging.

How does the MCP Bridge workThe MCP Bridge acts as a protocol translator between two worlds: Chrome’s native messaging protocol on one side and the MCP standard (JSON-RPC 2.0 over stdio) on the other.

These cover the browser extension, the agent deployed on AgentCore, the MCP Bridge, and a sample Excel MCP server.

Local tool automationThe bridge speaks standard MCP JSON-RPC over stdio, so MCP servers that use the stdio transport are compatible without modification.

ConclusionIn this post, we built an MCP bridge that connects a cloud-hosted Strands agent on …

2 days, 21 hours назад @ aws.amazon.com
Run production AI agents in n8n with Amazon Bedrock AgentCore harness
Run production AI agents in n8n with Amazon Bedrock AgentCore harness Run production AI agents in n8n with Amazon Bedrock AgentCore harness

AgentCore harness, a capability of Amazon Bedrock AgentCore, is now generally available and provides that scaffolding for you.

Add a manual trigger to a new workflow, add the Amazon Bedrock AgentCore node after it and attach your credential.

Turn on Add Tools, then under Tools, choose Add Tool and set Type to AgentCore Code Interpreter, a capability of Amazon Bedrock AgentCore.

Non-Bedrock providers use an API key stored in AgentCore Identity (a capability of Amazon Bedrock AgentCore).

Non-Bedrock providers use an API key stored in AgentCore Identity (a capability of Amazon Bedrock AgentCore).

2 days, 21 hours назад @ aws.amazon.com
NVIDIA
последний пост 4 часа назад
Firebird Launches CIS Region’s Largest AI Factory in Armenia
Firebird Launches CIS Region’s Largest AI Factory in Armenia Firebird Launches CIS Region’s Largest AI Factory in Armenia

The global buildout of AI infrastructure reached a new milestone today — Firebird, an emerging AI cloud, launched the CIS region’s largest AI factory in Armenia, establishing a new AI computing hub powered by NVIDIA accelerated computing and Dell Technologies high-performance AI infrastructure.

Firebird’s AI factory brings that capacity to Armenia, giving developers, startups, enterprises, universities and public institutions the compute to build and scale AI at home.

Firebird’s AI factory is designed from the ground up to turn compute into revenue.

Delivered in just over six months, the Armenia AI factory demonstrates Firebird’s ability to turn ambitious infrastructure plans into operation…

4 часа назад @ blogs.nvidia.com
GeForce NOW Shakes Up August With 26 New Games
GeForce NOW Shakes Up August With 26 New Games GeForce NOW Shakes Up August With 26 New Games

August is here, bringing 26 new games for GeForce NOW members.

Command the seas in World of Warships: Legends and discover what’s next in the GeForce NOW library, starting with the eight newly added games this week.

In addition, GeForce NOW is at the QuakeCon gaming conference this week in Grapevine, Texas, with hands-on experiences awaiting attendees.

Gamers not at the show can try out Ultimate cloud gaming in action with a day pass and jump into Bethesda titles from any device.

All Games on DeckWorld of Warships: Legends drops anchor on GeForce NOW this week, bringing free-to-play naval combat.

2 days, 2 hours назад @ blogs.nvidia.com
Into the Omniverse: How Open World Models Push the Frontier of Physical AI
Into the Omniverse: How Open World Models Push the Frontier of Physical AI Into the Omniverse: How Open World Models Push the Frontier of Physical AI

Open world models are already being used to generate training data, test policies and specialize physical AI systems.

World Models Are the Foundation of Physical AIThe data behind physical AI is difficult and expensive to collect at the scale required.

The NVIDIA Cosmos Coalition extends this work by bringing together world model builders, AI developers and physical AI leaders to contribute models, research and evaluation methods.

Together, these implementations and collaborations are establishing open world models as an adaptable foundation for physical AI across robots, autonomous vehicles and vision AI systems.

Get Plugged InLearn more about world models, OpenUSD and physical AI developm…

2 days, 2 hours назад @ blogs.nvidia.com
NVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US
NVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US NVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US

These regional hubs will help institutions share AI computing resources, accelerate scientific discovery and innovation, and prepare students to participate in the AI economy.

Expanding Access to AI InfrastructureThe State and Regional AI Infrastructure Hubs program will bring shared resources closer to the institutions and communities they serve.

And since 2017, UF faculty and units have received more than $511 million in AI research awards.

Connecting Research, Workforce and Regional GrowthFor policymakers and leaders, the hubs offer an opportunity to connect regional research and educational infrastructure with regional priorities and broader workforce and economic-development strategies…

3 days, 23 hours назад @ blogs.nvidia.com
NVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use
NVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use NVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use

NVIDIA Alpamayo 2 Super, available now for commercial use, is part of the Alpamayo family, the most-adopted open reasoning models for autonomous driving on Hugging Face, supporting a wide range of AV-relevant capabilities within a single foundation model.

Within the Alpamayo model family, Alpamayo 2 Super delivers the highest reasoning and driving performance for multimodal autonomous driving development, while Alpamayo 1.5 and Alpamayo 1 provide more cost-efficient options for cloud-based development and model distillation.

Together, the Alpamayo model family provides a cloud-to-car workflow that combines frontier-scale reasoning with scalable deployment across commercial AV fleets.

Benchm…

4 days назад @ blogs.nvidia.com
As AI Increases Demands on Memory, Storage Steps Up
As AI Increases Demands on Memory, Storage Steps Up As AI Increases Demands on Memory, Storage Steps Up

Benchmarks highlighted in this NVIDIA technical blog show that the NVIDIA Vera CPU, part of NVIDIA Vera BlueField-4 STX, delivers up to 3.21x higher throughput than an x86 CPU in a two-stage compression and encryption pipeline.

Storage-Next includes over 40 leading storage and flash vendors — including DDN, KIOXIA and Micron — each contributing to the next generation of AI storage technologies with NVIDIA.

Plus, NVIDIA CMX Context Memory Storage provides an AI‑native context tier for long‑context, multi‑turn, agentic AI inference, built on NVIDIA STX.

SCADA Enables Fast AI Storage That Stays SecureSpeed at the storage layer comes with a catch.

Join NVIDIA sessions at FMS, running Aug. 4-6 i…

4 days назад @ blogs.nvidia.com
AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency
AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency

Members of the Open Secure AI Alliance — now more than 120 organizations strong — are developing new guidelines to strengthen agentic AI cybersecurity as the annual Black Hat conference begins in Las Vegas today.

The SAFE guidelines are being drafted by an Open Secure AI Alliance working group.

Open Secure AI Alliance Delivers More Tools for AI CybersecurityThe SAFE framework adds to technology contributions Open Secure AI Alliance members are making as part of a shared commitment to building and sharing open, inspectable tools across the full AI security stack.

Alliance members are contributing tooling, harnesses and supporting technologies across this emerging layer of the AI security sta…

4 days, 2 hours назад @ blogs.nvidia.com
Run High-Performance Core Math at Scale with NVIDIA nvmath-python
Run High-Performance Core Math at Scale with NVIDIA nvmath-python Run High-Performance Core Math at Scale with NVIDIA nvmath-python

NVIDIA nvmath-python is a library designed to bridge the gap between the Python scientific community and NVIDIA CUDA-X math libraries.

A useful complement to existing array librariesLike other math libraries such as NumPy, nvmath-python implements core numerical operations useful in many engineering and scientific computing applications.

In the following example, nvmath-python consumes NumPy arrays and the result is also a NumPy array.

CPU libraries such as NVPL for NVIDIA Grace or any ARM v8 CPUs and Intel MKL for x86 hosts.

Get started with nvmath-pythonDesigned for productivity without performance compromises, nvmath-python reimagines the design of modern math libraries.

1 week, 1 day назад @ developer.nvidia.com
Best in Class: Stream PC Games and Study on the Same Laptop With GeForce NOW
Best in Class: Stream PC Games and Study on the Same Laptop With GeForce NOW Best in Class: Stream PC Games and Study on the Same Laptop With GeForce NOW

With cloud gaming, everyday laptops used for class can also become GeForce RTX-powered gaming setups.

With a GeForce NOW membership, Chromebooks, Macs and Surface devices can become GeForce RTX-powered gaming PCs in the cloud.

The GeForce NOW library features thousands of supported PC games, letting members stream titles they already own from digital stores like Steam, Xbox PC Game Pass and Epic Games Store in just a few clicks.

GeForce NOW turned the laptop — usually the creator’s go-to for editing videos and everyday work — into a top-notch gaming machine, too.

Everyone, at Their StationsExperience the Master Chief’s iconic journey on GeForce NOW in Halo: Campaign Evolved — a modernized r…

1 week, 2 days назад @ blogs.nvidia.com
Powerful Compute So Compact, It’s Clutch — Build AI in Your Hand With NVIDIA Jetson
Powerful Compute So Compact, It’s Clutch — Build AI in Your Hand With NVIDIA Jetson Powerful Compute So Compact, It’s Clutch — Build AI in Your Hand With NVIDIA Jetson

NVIDIA Jetson Orin Nano Super: Ideal for Building a First AI RobotRobots deserve a real brain and, apparently, an incredibly chic commute.

Jetson Orin Nano Super brings desktop-class generative AI to a handbag-friendly developer kit, offering first-time builders a practical path to learning computer vision, building AI agents, prototyping edge AI and more.

Reachy Mini Jetson AssistantReachy Mini Jetson Assistant is a low-latency, fully on-device voice and vision assistant for Reachy Mini Lite powered by NVIDIA Jetson Orin Nano Super.

Robotics AI PodcastUsing Jetson Orin Nano Super, Asier Arnaz built a Yocto-powered Robotics AI video podcast featuring two AI models discussing topics in real-…

1 week, 4 days назад @ blogs.nvidia.com
NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning
NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning

Ising Calibration 1.5 also uses examples from related experiments when available and is 11.4% smaller at BF16 precision.

How is the Ising Calibration 1.5 model trained?

Get started with NVIDIA Ising open resourcesThe NVIDIA Ising model family is fully open.

Model weightsFull-parameter checkpoints for Ising Calibration 1.5 are available on Hugging Face:Ising Calibration 1.5 is also available as an NVIDIA NIM and hosted through NVIDIA Build.

Quantum calibration agent blueprint is a script for deploying an agentic workflow using Ising Calibration 1.5 with the NVIDIA Nemo Agent Toolkit to quickly set up quantum calibration experiment automation.

1 week, 4 days назад @ developer.nvidia.com
Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security
Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security

The Open Secure AI Alliance — building on the leadership of the Linux Foundation’s Akrites initiative and OpenSSF community work — will work to remediate and disclose vulnerabilities using open technologies.

That is the mission of the Open Secure AI Alliance: to ensure defenders everywhere have open, frontier tools they can trust and control.

NVIDIA is contributing open models, model weights, data and new agent harness research to the Open Secure AI Alliance to speed the development of new cybersecurity tools and techniques.

That future is worth building — and the Open Secure AI Alliance invites governments, industry and researchers to join in the work of defending the AI era together.

Lear…

1 week, 5 days назад @ blogs.nvidia.com
NVIDIA Harnesses Vera CPU to Speed Up Design of Next-Generation CPUs and GPUs
NVIDIA Harnesses Vera CPU to Speed Up Design of Next-Generation CPUs and GPUs NVIDIA Harnesses Vera CPU to Speed Up Design of Next-Generation CPUs and GPUs

The complexity of modern chip design continues to grow as engineering teams work to develop increasingly sophisticated CPUs, GPUs and AI systems.

To help meet that challenge, NVIDIA is collaborating with industry leaders Cadence and Synopsys to optimize critical electronic design automation (EDA) applications for the NVIDIA Vera CPU.

While GPUs and AI have accelerated many aspects of chip design, several critical EDA workloads remain heavily dependent on CPU performance.

The results highlight Vera’s ability to accelerate two of the most compute-intensive stages of modern chip design.

Bringing Vera to the Design ProcessNVIDIA is deploying Vera throughout the EDA workflows used to create futu…

1 week, 5 days назад @ blogs.nvidia.com
At AI Summit, South Korea Outlines Its AI Future With NVIDIA and Partners
At AI Summit, South Korea Outlines Its AI Future With NVIDIA and Partners At AI Summit, South Korea Outlines Its AI Future With NVIDIA and Partners

At this week’s AI Summit in San Francisco, South Korean President Jae Myung Lee and some of the country’s top business leaders and researchers are meeting with NVIDIA and ecosystem partners to chart Korea’s AI progress.

To start, NVIDIA and the Korea Advanced Institute of Science and Technology (KAIST) today announced a joint AI research lab at the KAIST Kim Jaechul Graduate School of AI in Seoul, dedicated to advancing agentic AI for South Korea.

The collaboration will establish a robust academic AI research program, bringing together NVIDIA full-stack AI expertise, NVIDIA Nemotron open models and NVIDIA AI Cloud partner computing with the world-class scientific talent at KAIST, one of Asi…

2 weeks, 1 day назад @ blogs.nvidia.com
GeForce NOW Sets Sail With ‘Path of Exile: Curse of the Allflame’ Joining the Cloud
GeForce NOW Sets Sail With ‘Path of Exile: Curse of the Allflame’ Joining the Cloud GeForce NOW Sets Sail With ‘Path of Exile: Curse of the Allflame’ Joining the Cloud

Set sail in Path of Exile: Curse of the Allflame and charge in Battlefield 6 Season 4 both launching major content for members this week.

Then revisit Capcom legends like Breath of Fire IV, Dino Crisis and Dino Crisis 2, jump into Halo: Campaign Evolved Advanced Access and discover nine titles arriving on GeForce NOW.

Path of Exile: Curse of the Allflame launches Friday, July 24, sending Exiles into the perilous Frozen Seas.

Raw instinct takes over in Dino Crisis and Dino Crisis 2, the survival-horror series that combines pulse-pounding action, resource management and thrilling dinosaur encounters.

One GeForce NOW member recently shared they found themselves back on GeForce NOW Ultimate bec…

2 weeks, 2 days назад @ blogs.nvidia.com
Facebook
последний пост 2 days, 19 hours назад
From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking
From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking

Introducing the Multi-Stage Sequence ModelTo address scaling efficiency, a multi-stage model has been developed that enables scaling of a transformer-based sequence model in a compute efficient manner.

Second Stage: Online Ranking ModelThe offline user model representations are complemented with online ranking models that use fresh user signals and ad candidate information for real time ranking.

A Predictable Scaling CurveLLM-Style Scaling LawWhen running on real-world ads traffic, the multi-stage sequence model demonstrates the emergence of predictable scaling laws for ads recommendations that are analogous to those observed in large language models.

The Impact of Multi-Stage Sequence Mode…

2 days, 19 hours назад @ engineering.fb.com
GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model
GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

We tackled these challenges through complementary compute efficiency and scaling efficiency innovations: Compute efficiency : Achieved through a customized recommendation kernel library — Jagged Flash Attention (JFA), Generalized Dot-Product Attention (GDPA), BlockAttention, etc.

The results: we doubled GEM’s E2E training efficiency to 20-25% MFU while scaling total training FLOPs 4x over the past 12 months.

We measure training efficiency through E2E MFU, which decomposes into two factors:E2E MFU = Local MFU (compute efficiency) × Scaling Ratio (scaling efficiency)These factors describe two related but distinct optimization problems.

Local MFU (compute efficiency) measures how well a single…

4 days, 21 hours назад @ engineering.fb.com
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization

Hierarchical Interest Representation is an upstream representation layer designed to improve upon Meta’s deep funnel ranking optimization.

How Hierarchical Interest Representation Enhances Deep Funnel OptimizationHierarchical Interest Representation pioneers a structural shift in representation modeling by navigating long-range graph topologies and distilling sparse engagement signals into unified interest clusters at various granularities.

This aims to enable the delivery of more relevant ad content to optimize deep funnel ads.

Hierarchical Interest Representation learns super graphs, which cascade through multiple hierarchical layers for this flexibility, accommodating ranking modeling ar…

3 weeks, 2 days назад @ engineering.fb.com
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler
Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler

Why Ads Latency MattersMeta’s ads serving fleet handles more than 5 million requests per second on average at the serving platform entry point, which is over 400 billion per day across all monetized surfaces1.

That is why our Ads and Linux Kernel teams have been working together to build a scheduling policy customized to the ads delivery workload using sched_ext, the upstream, BPF-based extensible scheduling framework.

Until now, we have been using the general-purpose schedulers typically integrated in the Linux kernel (CFS and EEVDF) that balance threads across CPUs with no understanding of the workload.

It has already been deployed in several services at Meta, delivering meaningful reduct…

3 weeks, 4 days назад @ engineering.fb.com
10 Years of Meta’s Commitment to Python
10 Years of Meta’s Commitment to Python 10 Years of Meta’s Commitment to Python

This year marks Meta’s 10th consecutive year as a sponsor of the Python Software Foundation (PSF), the charitable organization dedicated to advancing, supporting, and protecting the open-source Python programming language and the community that sustains it.

Some of the core maintainers of Python are Meta engineers who have authored new features and Python Enhancement Proposals (PEPs) for the Python community.

These improvements are vital for protecting the global Python community and ensuring that developers everywhere – including our own engineers – can safely share and consume packages.

These investments help grow the Python community and foster the new talent that is essential for Python…

1 month, 1 week назад @ engineering.fb.com
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study
Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study Privacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case Study

Why Asset Classification MattersAsset classification is the foundation for many privacy controls.

The rest of this post walks through those pieces using asset classification as the case study.

All three share a single judge model, a larger reasoning model deliberately different from the classifier model.

Distill Stable Behavior Into RulesEven a strong LLM classifier should not be the default enforcement path forever.

Expand to other PAI workflows: The same pattern (context → LLM reasoning → distillation → deterministic enforcement) applies to lineage validation, purpose-boundary checking, and retention policy assignment.

1 month, 1 week назад @ engineering.fb.com
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems
SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems

The retrieval system within industry recommendation systems have consisted of microservices stitched together, with neural networks inconsistently integrated.

Under Index as Model previous microservice-based item indices used for retrieval become a tensor inside the model.

Moving From Microservice Mesh to One Integrated Neural NetworkThe Microservice Paradigm We ReplacedTraditional recommendation retrieval is built as a mesh of microservices.

We call this Index as Model: Every retrieval component — the item index, eligibility filter, scoring layer and user tower — becomes a tensor or operator inside a single PyTorch model.

Index FreshnessWith index as a model module, maintaining index fresh…

2 months, 1 week назад @ engineering.fb.com
Reel Friends: Building Social Discovery that Scales to Billions
Reel Friends: Building Social Discovery that Scales to Billions Reel Friends: Building Social Discovery that Scales to Billions

On its face the new Friend Bubbles feature looks simple enough.

It highlights Reels your friends have watched and reacted to.

On this episode of the Meta Tech Podcast, Pascal Hartig chats with Subasree and Joseph, two software engineers from the Facebook Reels team, about what it took to bring Friend Bubbles to life.

If you’ve ever underestimated a “simple” feature, this one’s for you.

And if you’re interested in learning more about career opportunities at Meta visit the Meta Careers page.

2 months, 3 weeks назад @ engineering.fb.com
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge
Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge

We’ve fundamentally transformed Facebook Groups Search to help people more reliably discover, sort through, and validate community content that’s most relevant to them.

We’ve adopted a new hybrid retrieval architecture and implemented automated model-based evaluation to address the major friction points people experience when searching community content.

Addressing the Friction Points in Community KnowledgePeople struggle with three friction points when searching for answers in community content – discovery, consumption, and validation.

The Solution: A Modernized Hybrid Retrieval ArchitectureWe engineered a hybrid retrieval architecture that powers a discussions module on Facebook Search.

R…

3 months, 2 weeks назад @ engineering.fb.com
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale

We’ve built a unified AI agent platform that encodes the domain expertise of senior efficiency engineers into reusable, composable skills.

Introducing the Capacity Efficiency ProgramWhen the code you ship serves more than 3 billion people, even a 0.1% performance regression can translate to significant additional power consumption.

Many engineers at Meta use our efficiency tools to work on these problems every day.

Skills : These encode domain expertise about performance efficiency.

The pipeline mirrors the defensive AI Regression Solver:Gather context with tools: The AI agent looks up: Opportunity metadata.

3 months, 3 weeks назад @ engineering.fb.com
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines
How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines

Challenging the Conventional Wisdom on AI Context FilesRecent academic research found that AI-generated context files actually decreased agent success rates on well-known open-source Python repositories.

Our codebase is the opposite: proprietary config-as-code with tribal knowledge that exists nowhere in any model’s training data.

Any team with a large, proprietary codebase can benefit:Identify your tribal knowledge gaps.

What’s NextWe are expanding context coverage to additional pipelines across Meta’s data infrastructure and exploring tighter integration between context files and code generation workflows.

This approach turned undocumented tribal knowledge into structured, AI-readable con…

4 months назад @ engineering.fb.com
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure
KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure

This is the second post in the Ranking Engineer Agent blog series exploring the autonomous AI capabilities accelerating Meta’s Ads Ranking innovation.

We introduce KernelEvolve, an agentic kernel authoring system used by Ranking Engineer Agent and generally applicable to a range of AI models beyond Ads Ranking.

Unlike typical large language model (LLM)-based agents that perform one-shot code generation, KernelEvolve treats kernel optimization as a search problem.

A standard coding assistant lacks the context to write optimized MTIA kernels because it has never seen MTIA documentation, instruction set details, or programming idioms.

KernelEvolve represents an early step toward the vision of …

4 months, 1 week назад @ engineering.fb.com
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads

To overcome this, we have developed the Meta Adaptive Ranking Model, which effectively bends the inference scaling curve with high ROI and industry-leading efficiency.

Introducing Meta Adaptive Ranking ModelServing LLM-scale & complexity models in a real-time ads recommendation environment requires resolving a fundamental tension between model complexity and system efficiency.

Adaptive Ranking Model addresses these challenges through a paradigm shift powered by three core innovations across the serving stack:Inference-efficient model scaling: Adaptive Ranking Model achieves a model complexity equivalent to the O(10 GFLOPs) per token used by top-tier LLMs.

To minimize compute overhead, Adapt…

4 months, 1 week назад @ engineering.fb.com
AI for American-Produced Cement and Concrete
AI for American-Produced Cement and Concrete AI for American-Produced Cement and Concrete

Concurrent with the 2026 American Concrete Institute (ACI) Spring Convention, Meta is releasing a new AI model for designing concrete mixes – Bayesian Optimization for Concrete (BOxCrete), as well as the foundational data used to develop award-winning concrete mixes.

Amrize operates 18 cement plants, 141 cement terminals and 269 ready-mix concrete sites across North America.

Alongside the event, Meta is releasing a new AI model for designing concrete mixes, Bayesian Optimization for Concrete (BOxCrete).

How Meta Leverages AI for Concrete MixturesMeta’s AI for concrete model can help suppliers more quickly incorporate U.S. materials into their mixes through an approach called adaptive experi…

4 months, 1 week назад @ engineering.fb.com
Friend Bubbles: Enhancing Social Discovery on Facebook Reels
Friend Bubbles: Enhancing Social Discovery on Facebook Reels Friend Bubbles: Enhancing Social Discovery on Facebook Reels

Friend bubbles in Facebook Reels highlight Reels your friends have liked or reacted to, helping you discover new content and making it easier to connect over shared interests.

Friend bubbles enhance the social experience on Facebook Reels by helping you discover content your friends enjoy, creating a shared viewing experience and sparking new conversations.

Along with additional optimizations in the underlying method, this approach enabled us to ship friend bubbles while preserving core Reels performance.

Friend bubbles work because the signal is high value: It adds meaningful social context that helps people decide what’s worth watching.

Engagement also scales consistently with the number …

4 months, 3 weeks назад @ engineering.fb.com
Uber Engineering
последний пост None
neptune.ai neptune.ai
последний пост 8 months, 1 week назад
We are joining OpenAI
We are joining OpenAI We are joining OpenAI

Piotr Niedźwiedź, CEO/CTO and founder of neptune.aiI’m excited to share that we’ve entered into a definitive agreement to be acquired by OpenAI, subject to closing conditions.

We are thrilled to join the OpenAI team and help their AI researchers build better models faster.

Neptune is a metrics dashboard company.”We’ve worked closely with OpenAI to create the metrics dashboard that helps teams building foundation models.

Our future with OpenAINeptune will join OpenAI and continue to support AI researchers with tools to monitor, debug, and evaluate frontier models.

We are looking forward to working with top AI researchers and supporting OpenAI’s mission of ensuring that AGI benefits all of hu…

8 months, 1 week назад @ neptune.ai
Synthetic Data for LLM Training
Synthetic Data for LLM Training Synthetic Data for LLM Training

For instance, financial data is highly sensitive and protected by very strict regulations, and synthetic data mimics the real data distribution without revealing customer information.

Read more about how leading foundation model teams curate their training data and other topics in the State of Foundation Model Training Report 2025.

Choosing the right synthetic data generation technique depends on the type of data and its complexity.

Synthetic tabular data generation is a promising direction to overcome these challenges by learning the distribution of the tabular data.

Post-processingAs the distribution of tabular data is highly complex, it makes the synthetic tabular data generation very ch…

8 months, 4 weeks назад @ neptune.ai
What are LLM Embeddings: All you Need to Know
What are LLM Embeddings: All you Need to Know What are LLM Embeddings: All you Need to Know

TL;DR LLM embeddings are the numerical, vector representations of text that Large Language Models (LLMs) use to process information.

Unlike their predecessor word embeddings, LLM embeddings are context-aware and dynamically change to capture semantic and syntactic relationships based on the surrounding text.

What are the applications of LLM embeddings?

Word EmbeddingsSparse Word Embeddings One-Hot Vectors 1970s TF-IDF1980s Co-Occurrence MatrixStatic Word Embeddings Word2Vec 2013 GloVe 2014Contextualized word embeddings ELMo 2018 GPT-1 2018 BERT 2018 LLAMA 2023 DeepSeek-V1 2023 GPT-4 2023Static word embeddingsStatic word embeddings, such as word2vec in 2013, marked a significant development.…

9 months назад @ neptune.ai
Detecting and Fixing ‘Dead Neurons’ in Foundation Models
Detecting and Fixing ‘Dead Neurons’ in Foundation Models Detecting and Fixing ‘Dead Neurons’ in Foundation Models

TL;DR Dead neurons silently waste compute and reduce effective model capacity in foundation models.

Dead neurons’ impactRecent studies into dead neurons in the context of foundation models show interesting, albeit worrying, results.

These large reported fractions of dead neurons in foundation models are a concern from a computational perspective.

Before we move on to discuss how to detect and fix dead neurons, let’s touch upon an important distinction between dead neurons and vanishing gradients.

Further reading How to Monitor, Diagnose, and Solve Gradient Issues in Foundation Models Read moreVisualizing activation distributionsIs your foundation model suffering from dead neurons?

9 months, 1 week назад @ neptune.ai
Part 2: Instruction Fine-Tuning: Evaluation and Advanced Techniques for Efficient Training
Part 2: Instruction Fine-Tuning: Evaluation and Advanced Techniques for Efficient Training Part 2: Instruction Fine-Tuning: Evaluation and Advanced Techniques for Efficient Training

In the first part of this series, we covered the fundamentals of instruction fine-tuning (IFT).

def calculate_irs(instruction, output, reference_model): evaluation_prompt = f""" Instruction: {instruction} Model Output: {output} Rate how well the output follows the instruction on these criteria: 1.

| SourceHINT addresses a computational inefficiency in standard instruction fine-tuning: repeatedly reprocessing the same task instruction with every input example.

Read more about foundation model training infrastructure and other topics in Neptune’s 2025 State of Foundation Model Training Report.

First, during initial instruction fine-tuning across multiple diverse tasks, the model learns genera…

9 months, 2 weeks назад @ neptune.ai
How to Optimize LLM Inference
How to Optimize LLM Inference How to Optimize LLM Inference

Large Language Model (LLM) inference at scale is challenging as it involves transferring massive amounts of model parameters and data and performing computations on large tensors.

In the following, we’ll use the Llama model family architecture as a specific example to understand the LLM workload at inference.

For a far more detailed analysis of the LLM workload at inference, see the chapter All About Transformer Inference in the book How to Scale Your Model, published by Google DeepMind.

See also How to Run LLMs Locally Read moreA quick primer on hardware for LLM inferenceA typical LLM inference cluster consists of several nodes, each with a multi-core CPU and multiple accelerator devices, …

9 months, 3 weeks назад @ neptune.ai
▶️ YouTube
Yannic Kilcher Yannic Kilcher
последний пост 5 months назад
I BUILT A FULLY AUTOMATIC MANSPLAINER
I BUILT A FULLY AUTOMATIC MANSPLAINER I BUILT A FULLY AUTOMATIC MANSPLAINER

All information about GTC and the DGX Spark Raffle is here: https://www.ykilcher.com/gtc Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereu…

5 months назад @ youtube.com
Traditional X-Mas Stream
Traditional X-Mas Stream Traditional X-Mas Stream

Letsgooo

7 months, 1 week назад @ youtube.com
Traditional Holiday Live Stream
Traditional Holiday Live Stream Traditional Holiday Live Stream

https://ykilcher.com/discord Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yannic-kilcher

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-kilcher-488534136/

BiliBili: https://space.bilibili.com/1824646584 If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https:/…

7 months, 1 week назад @ youtube.com
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis) TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)

Paper: https://arxiv.org/abs/2511.08923 Abstract:

Diffusion language models hold the promise of fast parallel generation, while autoregressive (AR) models typically excel in quality due to their causal structure aligning naturally with language modeling. This raises a fundamental question: can we achieve a synergy with high throughput, higher GPU utilization, and AR level quality? Existing methods fail to effectively balance these two aspects, either prioritizing AR using a weaker model for sequential drafting (speculative decoding), leading to lower drafting efficiency, or using some form of left-to-right (AR-like) decoding logic for diffusion, which still suffers from quality degradation …

7 months, 2 weeks назад @ youtube.com
Titans: Learning to Memorize at Test Time (Paper Analysis)
Titans: Learning to Memorize at Test Time (Paper Analysis) Titans: Learning to Memorize at Test Time (Paper Analysis)

Paper: https://arxiv.org/abs/2501.00663 Abstract:

Over more than a decade there has been an extensive research effort on how to effectively utilize recurrent models and attention. While recurrent models aim to compress the data into a fixed-size memory (called hidden state), attention allows attending to the entire context window, capturing the direct dependencies of all tokens. This more accurate modeling of dependencies, however, comes with a quadratic cost, limiting the model to a fixed-length context. We present a new neural long-term memory module that learns to memorize historical context and helps attention to attend to the current context while utilizing long past information. We sh…

7 months, 3 weeks назад @ youtube.com
[Paper Analysis] The Free Transformer (and some Variational Autoencoder stuff)
[Paper Analysis] The Free Transformer (and some Variational Autoencoder stuff) [Paper Analysis] The Free Transformer (and some Variational Autoencoder stuff)

https://arxiv.org/abs/2510.17558 Abstract:

We propose an extension of the decoder Transformer that conditions its generative process on random latent variables which are learned without supervision thanks to a variational procedure. Experimental evaluations show that allowing such a conditioning translates into substantial improvements on downstream tasks. Author: François Fleuret Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the con…

9 months, 1 week назад @ youtube.com
[Video Response] What Cloudflare's code mode misses about MCP and tool calling
[Video Response] What Cloudflare's code mode misses about MCP and tool calling [Video Response] What Cloudflare's code mode misses about MCP and tool calling

Theo's Video: https://www.youtube.com/watch?v=bAYZjVAodoo

Cloudflare article: https://blog.cloudflare.com/code-mode/ Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8…

9 months, 3 weeks назад @ youtube.com
Henry AI Labs Henry AI Labs
последний пост None
3blue1brown 3blue1brown
последний пост 2 weeks, 1 day назад
The 64 sugar cubes puzzle
The 64 sugar cubes puzzle The 64 sugar cubes puzzle

See all monthly puzzles: https://momath.org/mindbenders/

2 weeks, 1 day назад @ youtube.com
But what is cross-entropy? | Compression is Intelligence Part 2
But what is cross-entropy? | Compression is Intelligence Part 2 But what is cross-entropy? | Compression is Intelligence Part 2

Where the loss function for training LLMs comes from.

Job opportunities aligned to this audience: https://3b1b.co/talent

Early views and other perks for supporters: https://3b1b.co/support

Home page: https://www.3blue1brown.com Manim animations by Aaron Gostein and Grant Sanderson

NanoGPT animation by Clayton Rabideau

3d black-box model by Paul Dancstep

Music by Vince Rubinetti Timestamps 0:00 - Language trees and zipping

3:02 - Recap optimal codes

5:20 - Defining cross-entropy

8:26 - Intuition and examples

12:59 - Application to language trees

14:55 - Pre-training LLMs

20:38 - What makes this loss function best?

26:13 - Distillation

30:12 - 3b1b Talent

31:35 - KL Divergence ---------------…

3 weeks, 2 days назад @ youtube.com
100 random chords, how many intersections?
100 random chords, how many intersections? 100 random chords, how many intersections?

Part of a series of monthly puzzles done in collaboration with MoMath.

1 month, 3 weeks назад @ youtube.com
Measuring the entropy of English
Measuring the entropy of English Measuring the entropy of English

Full video: https://youtu.be/l6DKRf-fAAM

1 month, 3 weeks назад @ youtube.com
What's the perfect encoding? How do you know?
What's the perfect encoding? How do you know? What's the perfect encoding? How do you know?

Full video: https://youtu.be/l6DKRf-fAAM

1 month, 4 weeks назад @ youtube.com
Reinventing Entropy | Compression & Intelligence Part 1
Reinventing Entropy | Compression & Intelligence Part 1 Reinventing Entropy | Compression & Intelligence Part 1

What is the fundamental compressibility of language?

Check out our virtual career fair: https://3b1b.co/talent

See new projects before they go live: https://3b1b.co/support Animation credit:

Manim scenes by Aaron Gostein and Grant Sanderson

Shannon’s story, as well as those for various pi creatures, by Mitchell Zemil.

Lunar robot and prediction/compression coin by Paul Dancstep

NanoGPT animations by Clayton Rabideau Shannon’s “A Mathematical Theory of Communication”

https://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf Shannon’s “Prediction and Entropy of Printed English”

https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf Scientific American article that…

2 months назад @ youtube.com
Tie random ends: How many loops?
Tie random ends: How many loops? Tie random ends: How many loops?

Recent puzzle solutions on Patreon:

https://members.3blue1brown.com/posts/158885046?pr=true

2 months, 2 weeks назад @ youtube.com
Covering 10 points, a surprisingly tricky puzzle.
Covering 10 points, a surprisingly tricky puzzle. Covering 10 points, a surprisingly tricky puzzle.

Made as part of a monthly series of puzzles for the 2026 Year of Math.

3 months, 3 weeks назад @ youtube.com
Escher's most mind-bending piece
Escher's most mind-bending piece Escher's most mind-bending piece

On "The Print Gallery", by M.C. Escher

Full video: https://youtu.be/ldxFjLJ3rVY

4 months, 1 week назад @ youtube.com
The subset sum puzzle
The subset sum puzzle The subset sum puzzle

Part of a series of monthly puzzlers. Stay subscribed to see the solution

4 months, 2 weeks назад @ youtube.com
Escher's most mathematically interesting piece
Escher's most mathematically interesting piece Escher's most mathematically interesting piece

Escher's Print Gallery, and the tour of complex analysis it invites.

Check out our virtual career fair: 3b1b.co/talent

Join channel supporters to see videos early: 3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Original paper by de Smit and Lenstra:

https://pub.math.leidenuniv.nl/~smitbde/papers/2003-de_smit-lenstra-escher.pdf Timestamps: 0:00 - The print gallery

13:04 - Conformal maps from complex analysis

21:41 - The complex exponential

25:56 - The complex logarithm

32:32 - 3b1b Talent

33:14 - Constructing the key function

40:16 - The deeper math behind Escher ------------------ These animations are largely made us…

4 months, 2 weeks назад @ youtube.com
Bacteria Grid Puzzle Solution
Bacteria Grid Puzzle Solution Bacteria Grid Puzzle Solution

Part of a monthly series of puzzlers, in collaboration with MoMath and Peter Winkler

4 months, 2 weeks назад @ youtube.com
The most underappreciated formula | Exploring high-dimensional spheres
The most underappreciated formula | Exploring high-dimensional spheres The most underappreciated formula | Exploring high-dimensional spheres

On the volumes of higher-dimensional spheres

Explore the 3b1b virtual career fair: See https://3b1b.co/talent

Become a supporter for early views of new videos: https://3b1b.co/support

An equally valuable form of support is to simply share the videos.

Home page: https://www.3blue1brown.com Thanks to UC Santa Cruz for letting me film there, and special thanks to Pedro Morales-Almazan for arranging everything. My video on Numberphile with a fun application of this problem: https://youtu.be/6_yU9eJ0NxA Timestamps:

0:00 - Introduction

1:01 - Random puzzle

6:16 - Outside the box

14:35 - Setting up the volume grid

21:14 - Why 4πr^2

25:21 - Archimedes in higher dimensions

36:17 - The general formul…

5 months, 1 week назад @ youtube.com
The lattice bacteria puzzle
The lattice bacteria puzzle The lattice bacteria puzzle

Part of a series of monthly puzzles, done in collaboration with MoMath.

https://momath.org/mindbenders

5 months, 3 weeks назад @ youtube.com
Solution to the ladybug clock puzzle
Solution to the ladybug clock puzzle Solution to the ladybug clock puzzle

Solution to last month's probability puzzle.

5 months, 3 weeks назад @ youtube.com
Two Minute Papers Two Minute Papers
последний пост 1 day, 6 hours назад
DeepMind's AI Trick Everyone Should Copy
DeepMind's AI Trick Everyone Should Copy DeepMind's AI Trick Everyone Should Copy

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Gemma4 paper and some more is available here:

https://arxiv.org/abs/2607.02770

https://x.com/googlegemma/status/2077449152062247219

https://x.com/UnslothAI/status/2078118183085731843 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 day, 6 hours назад @ youtube.com
The Billion Dollar AI Race Just Broke
The Billion Dollar AI Race Just Broke The Billion Dollar AI Race Just Broke

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 Qwen 3.8 Max:

https://qwen.ai/blog?id=qwen3.8 Sources:

https://x.com/loktar00/status/2082589566934929750

https://x.com/CommandCodeAI/status/2084293498950590839 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

3 days, 1 hour назад @ youtube.com
Another DeepSeek Moment
Another DeepSeek Moment Another DeepSeek Moment

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 DeepSeek v4 Flash 0731:

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

5 days, 5 hours назад @ youtube.com
New AI Learned Parkour From Just 30 Seconds Of Video
New AI Learned Parkour From Just 30 Seconds Of Video New AI Learned Parkour From Just 30 Seconds Of Video

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://jiashunwang.github.io/HIL/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

6 days назад @ youtube.com
Kimi K3 Just Broke The Economics Of AI
Kimi K3 Just Broke The Economics Of AI Kimi K3 Just Broke The Economics Of AI

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://arxiv.org/abs/2607.24653 Try Kimi K3 (subject to availability): https://www.kimi.com/ Links:

https://macos27.kimi.page/

https://x.com/mweinbach/status/2077878247920951400

https://x.com/intheworldofai/status/2077838911494336681

https://x.com/chetaslua/status/2077829183989072281

https://x.com/hqmank/status/2078104317027094907 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen …

1 week, 3 days назад @ youtube.com
AI Helped Them Code Faster… But At A Cost
AI Helped Them Code Faster… But At A Cost AI Helped Them Code Faster… But At A Cost

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://www.anthropic.com/research/AI-assistance-coding-skills 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

3 weeks, 1 day назад @ youtube.com
The Hidden World Inside An AI
The Hidden World Inside An AI The Hidden World Inside An AI

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://transformer-circuits.pub/2025/linebreaks/index.html Paper for reindeer vision change - https://royalsocietypublishing.org/rspb/article/280/1773/20132451/50765/Shifting-mirrors-adaptive-changes-in-retinal 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli …

3 weeks, 3 days назад @ youtube.com
New AI Just Reinvented Minecraft Worlds
New AI Just Reinvented Minecraft Worlds New AI Just Reinvented Minecraft Worlds

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://xandergos.github.io/terrain-diffusion/

https://modrinth.com/mod/terrain-diffusion

https://github.com/xandergos/terrain-diffusion Source video for some parts of the footage: https://www.youtube.com/watch?v=irE4tcDtUIg 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fi…

3 weeks, 5 days назад @ youtube.com
DeepSeek's New AI Speed Hack Is Amazing
DeepSeek's New AI Speed Hack Is Amazing DeepSeek's New AI Speed Hack Is Amazing

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The DeepSeek paper is available here:

https://arxiv.org/abs/2607.05147v1 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month назад @ youtube.com
Game Physics Just Got 170 Times Faster
Game Physics Just Got 170 Times Faster Game Physics Just Got 170 Times Faster

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here:

https://arxiv.org/abs/2506.06494 Sources:

https://www.youtube.com/shorts/Tx7167DXr8U

https://www.youtube.com/watch?v=55F9dY2Y1zc 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month назад @ youtube.com
This New AI Model Changes Everything
This New AI Model Changes Everything This New AI Model Changes Everything

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers GLM 5.2: https://z.ai/blog/glm-5.2 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

1 month, 1 week назад @ youtube.com
DeepSeek Just Solved AI's Billion Dollar Problem
DeepSeek Just Solved AI's Billion Dollar Problem DeepSeek Just Solved AI's Billion Dollar Problem

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://arxiv.org/abs/2602.21548 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi #deepseek

1 month, 2 weeks назад @ youtube.com
This is OpenClaw On Steroids
This is OpenClaw On Steroids This is OpenClaw On Steroids

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here:

https://recursivemas.github.io/

https://github.com/RecursiveMAS/RecursiveMAS 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi Thumbnail design: https://felicia.hu

1 month, 2 weeks назад @ youtube.com
Claude AI Knows More Than It Tells You
Claude AI Knows More Than It Tells You Claude AI Knows More Than It Tells You

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here:

https://www.anthropic.com/research/natural-language-autoencoders

https://transformer-circuits.pub/2026/nla/index.html 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi My research: https://cg.tuwien.ac.at/~zsolnai/

Thumbnail design: https://felicia.hu

1 month, 3 weeks назад @ youtube.com
NVIDIA's New Free AI - A Gift To All of Us
NVIDIA's New Free AI - A Gift To All of Us NVIDIA's New Free AI - A Gift To All of Us

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Nemotron 3 Ultra paper is available here:

https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/ Free Rendering course and source code:

https://users.cg.tuwien.ac.at/zsolnai/gfx/rendering-course/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi Thumbnail design: https://f…

1 month, 3 weeks назад @ youtube.com
DataFest Video DataFest Video
последний пост None
Семинары JetBrains Research Семинары JetBrains Research
последний пост None
Яндекс. Компьютерные науки Яндекс. Компьютерные науки
последний пост 4 days назад
ML Global Recap'H1 2026
ML Global Recap'H1 2026 ML Global Recap'H1 2026

Обсудим итоги ICML и других международных конференций, главные ML-тренды первого полугодия 2026-го и собственный опыт.

4 days назад @ youtube.com
Омни-модели будущего 🚀
Омни-модели будущего 🚀 Омни-модели будущего 🚀

Что они будут уметь — рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 1 week назад @ youtube.com
Качество модели взлетело... без мультимодального RL?
Качество модели взлетело... без мультимодального RL? Качество модели взлетело... без мультимодального RL?

О росте мультимодального качества рассказал Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 1 week назад @ youtube.com
Работа с данными — это скучно?
Работа с данными — это скучно? Работа с данными — это скучно?

А почему — рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 2 weeks назад @ youtube.com
Как приготовить SFT 🍲
Как приготовить SFT 🍲 Как приготовить SFT 🍲

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 2 weeks назад @ youtube.com
Почему мультимодальные модели — это база 🤖
Почему мультимодальные модели — это база 🤖 Почему мультимодальные модели — это база 🤖

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 2 weeks назад @ youtube.com
Омни-модель: что это за зверь такой
Омни-модель: что это за зверь такой Омни-модель: что это за зверь такой

Рассказывает Роман Исаченко, руководитель группы анализа изображений в Яндекс R&D. #искусственныйинтеллект #нейросети #мультимодальность #омнимодель #машинноеобучение #datascience #яндекс #ai #ml #технологии

1 month, 3 weeks назад @ youtube.com
Borealis — как обучить аудио-LLM по цене MacBook
Borealis — как обучить аудио-LLM по цене MacBook Borealis — как обучить аудио-LLM по цене MacBook

На конференции Data Fest 2026 в Белграде независимый исследователь Александр Николич рассказал практическую историю создания аудиоязыковой модели Borealis с бюджетом, сопоставимым со стоимостью MacBook. Больше контента для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 4 weeks назад @ youtube.com
Better LLM pre-training in NVFP4
Better LLM pre-training in NVFP4 Better LLM pre-training in NVFP4

At Data Fest 2026 in Belgrade, Andrei Panferov from the Institute of Science and Technology Austria introduced Quartet II, a novel method for NVFP4 pre-training that recovers SOTA accuracy. He outlined the core challenges of low-precision LLM training and presented CUDA kernels tuned for Blackwell GPUs, ready for integration into real training pipelines. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 4 weeks назад @ youtube.com
Как безопасно выкатывать новые версии продуктовых AI-агентов
Как безопасно выкатывать новые версии продуктовых AI-агентов Как безопасно выкатывать новые версии продуктовых AI-агентов

На Data Fest 2026 в Белграде Дмитрий Коршунов, Team Lead ML в Ecom, показал, как безопасно обновлять продуктовых AI-агентов с помощью системы автометрик. На примере агента Яндекс AI для турецкого рынка он объяснил, как фиксировать регрессии до прода, сравнивать версии и принимать решение о релизе, когда простой «Hello, Agent» уже позади. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 4 weeks назад @ youtube.com
Как решаем оптимизационные задачи Яндекс Лавки с помощью uplift-моделей
Как решаем оптимизационные задачи Яндекс Лавки с помощью uplift-моделей Как решаем оптимизационные задачи Яндекс Лавки с помощью uplift-моделей

На Data Fest 2026 в Белграде Вячеслав Костров, ML-инженер в Яндексе, рассказал, как uplift-модели решают бизнес-задачи Лавки: от персональных скидок до показа продуктовых подборок. Он разобрал постановку uplift-задачи, подбор метрик и построение политик, а также практические приёмы с лагранжианом и uplift-деревьями для баланса ограничений. Всё это — на примере реальных внедрений и с разбором типичных ошибок. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #Mu…

1 month, 4 weeks назад @ youtube.com
HGRPO: Hierarchical Grouped Reward Policy Optimization for Multi-Turn Conversational Agents
HGRPO: Hierarchical Grouped Reward Policy Optimization for Multi-Turn Conversational Agents HGRPO: Hierarchical Grouped Reward Policy Optimization for Multi-Turn Conversational Agents

At Data Fest 2026 in Belgrade, Karina Romanova, Senior LLM Research Engineer, presented HGRPO — a hierarchical modification of GRPO for multi-turn dialogue agents. Applied to a booking agent in Yandex Alice, the method improved truthfulness by 8.0 percentage points and reduced dialogue length by 10.7%. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 4 weeks назад @ youtube.com
Hacks and Defenses in Automatic Kernel Generation
Hacks and Defenses in Automatic Kernel Generation Hacks and Defenses in Automatic Kernel Generation

На Data Fest 2026 в Белграде Егор Коновалов, ML-инженер, разобрал хаки, которые находят LLM-агенты, когда генерируют GPU/TPU-код: от тривиального обхода numerical tolerance до изощрённых атак на timing-измерения и эксплуатации дыр в test harness. А ещё Егор показал, какие методы защиты реально работают, а какие создают ложное чувство безопасности. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 4 weeks назад @ youtube.com
Поиск по архивам: как мы переходим к осознанному распознаванию текста
Поиск по архивам: как мы переходим к осознанному распознаванию текста Поиск по архивам: как мы переходим к осознанному распознаванию текста

На Data Fest 2026 в Белграде Дарья Виноградова, лид команды компьютерного зрения, представила два важных майлстоуна архивного поиска: новую архитектуру распознавания текста и выделение смысловых структур. Эти изменения делают поиск человечнее — теперь можно искать не слова среди текста, а человека среди людей. Больше материалов для разработчиков: https://t.me/+owyCvdge8WIyNTUy #DataFest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #ReinforcementLearning #MultimodalAI

1 month, 4 weeks назад @ youtube.com
Real-time video generation: where we are and what comes next
Real-time video generation: where we are and what comes next Real-time video generation: where we are and what comes next

At Data Fest 2026 in Belgrade, Andrey Filatov from KREA AI broke down the current state of real-time video generation: which architectures dominate, how they differ, and what challenges arise from compute limits and memory bottlenecks. He also covered production solutions like distillation and caching, and shared his outlook for the next 2–3 years: what will soon become possible and which bottlenecks the industry still overlooks. More content for developers: https://t.me/+owyCvdge8WIyNTUy #datafest #DataFest2026 #AI #ML #LLM #GenAI #MachineLearning #DataScience #MLOps #AIAgents #RAG #ComputerVision #AutonomousDriving #Yandex #Яндекс #TechTalk #Developers #ArtificialIntelligence #Reinforceme…

1 month, 4 weeks назад @ youtube.com
ML Trainings ML Trainings
последний пост 2 days, 1 hour назад
Дмитрий Корнилов | Causal Discovery на реальных данных: как подружить причинно-следственные графы
Дмитрий Корнилов | Causal Discovery на реальных данных: как подружить причинно-следственные графы Дмитрий Корнилов | Causal Discovery на реальных данных: как подружить причинно-следственные графы

Спикер: Дмитрий Корнилов, Сколковский институт науки и технологий, инженер-исследователь Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Reliable ML ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 days, 1 hour назад @ youtube.com
Александр Календарев | ML в PostgreSQL
Александр Календарев | ML в PostgreSQL Александр Календарев | ML в PostgreSQL

Спикер: Александр Календарев, Datagile, разработчик БД Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции ML in DBMS https://ods.ai/tracks/df26-ml-in-dbms ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 days, 7 hours назад @ youtube.com
Айгуль Камалтинова | Как сэкономить, улучшая командные процессы
Айгуль Камалтинова | Как сэкономить, улучшая командные процессы Айгуль Камалтинова | Как сэкономить, улучшая командные процессы

Спикер: Айгуль Камалтинова, Школа 21 Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Open Career ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 days, 21 hours назад @ youtube.com
Андрей Тоток и Анна Юрищева | LLM на собеседовании - Оно того стоит?
Андрей Тоток и Анна Юрищева | LLM на собеседовании - Оно того стоит? Андрей Тоток и Анна Юрищева | LLM на собеседовании - Оно того стоит?

Спикеры: Андрей Тоток и Анна Юрищева, ML Lead, ДОМ.РФ Teamlead ML Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Open Career https://ods.ai/tracks/df26-opencareer ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

2 days, 23 hours назад @ youtube.com
Всеволод Викулин | Реальная автоматизация: в каких процессах AI-агенты приносят пользу
Всеволод Викулин | Реальная автоматизация: в каких процессах AI-агенты приносят пользу Всеволод Викулин | Реальная автоматизация: в каких процессах AI-агенты приносят пользу

Спикер: Всеволод Викулин, Т-Банк, руководитель команды роботов в обслуживании Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Data Strategy ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 22 hours назад @ youtube.com
Максим Горшков | КультИИ
Максим Горшков | КультИИ Максим Горшков | КультИИ

Спикер: Максим Горшков, руководитель направления по исследованию данных, Сбер Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenAI от Сбера https://ods.ai/tracks/df26_sber_genai ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 22 hours назад @ youtube.com
Александр Юрышев | Опыт создания рекомендательной системы кофе с LLM и RAG
Александр Юрышев | Опыт создания рекомендательной системы кофе с LLM и RAG Александр Юрышев | Опыт создания рекомендательной системы кофе с LLM и RAG

Спикер: Александр Юрышев, главный инженер по разработке, Сбер Data Fest 2026: https://ods.ai/events/datafest2026

Презентацию к докладу Вы можете скачать в треке секции GenAI от Сбера https://ods.ai/tracks/df26_sber_genai ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

3 days, 22 hours назад @ youtube.com
Разлом глобальной науки — взгляд Дмитрия и Валентина
Разлом глобальной науки — взгляд Дмитрия и Валентина Разлом глобальной науки — взгляд Дмитрия и Валентина 4 days, 7 hours назад @ youtube.com
Почему данные стали важны и как с ними обращаться
Почему данные стали важны и как с ними обращаться Почему данные стали важны и как с ними обращаться 4 days, 7 hours назад @ youtube.com
Китайское начальство обсуждает ограничения доступа к китайским моделям
Китайское начальство обсуждает ограничения доступа к китайским моделям Китайское начальство обсуждает ограничения доступа к китайским моделям 4 days, 7 hours назад @ youtube.com
Жизнь под аудиокамерами: что слышат операторы
Жизнь под аудиокамерами: что слышат операторы Жизнь под аудиокамерами: что слышат операторы 4 days, 7 hours назад @ youtube.com
Дмитрий рассказывает о дроне, соседе и терроризме
Дмитрий рассказывает о дроне, соседе и терроризме Дмитрий рассказывает о дроне, соседе и терроризме 4 days, 7 hours назад @ youtube.com
Никита Трифонов | Uplift- модели в акционных предложениях
Никита Трифонов | Uplift- модели в акционных предложениях Никита Трифонов | Uplift- модели в акционных предложениях

Спикер: Никита Трифонов, Data Scientist, Банк ВТБ Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции Data Fusion https://ods.ai/tracks/df26-data-fusion ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 23 hours назад @ youtube.com
Илья Чумак | Внедрение LLM в модерацию: от PoC до целевого решения
Илья Чумак | Внедрение LLM в модерацию: от PoC до целевого решения Илья Чумак | Внедрение LLM в модерацию: от PoC до целевого решения

Спикер: Илья Чумак Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции ML in Marketplace от Avito.tech https://ods.ai/tracks/df26_mlavitotech ______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 23 hours назад @ youtube.com
Андрей Ларькин | SQL агент для Lemana GPT
Андрей Ларькин | SQL агент для Lemana GPT Андрей Ларькин | SQL агент для Lemana GPT

Спикер: Андрей Ларькин, Лемана Тех, Ведущий специалист по науке о данных Data Fest 2026: https://ods.ai/events/datafest2026 Презентацию к докладу Вы можете скачать в треке секции LeanAI от ЛЕМАНА ТЕХ https://ods.ai/tracks/df26-leanai

______

Наши соц.сети:

Telegram: https://t.me/datafest

Вконтакте: https://vk.com/datafest

Канал с вакансиями в telegram: https://t.me/odsjobs

Канал с апдейтами по курсам: https://t.me/odscourses

Как попасть в чат сообщества ODS Mattermost: https://ods.ai/tracks/mattermost

4 days, 23 hours назад @ youtube.com
🎧 Podcasts
Lex Fridman AI Podcast Lex Fridman AI Podcast
последний пост 1 week, 3 days назад
#499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee
#499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee #499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee

Gary Gallagher is a historian of the American Civil War.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep499-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://plaud.ai/lexOUTLINE:(00:00) – Introduction(00:07) – Sponsors, Comments, and Reflections(08:36) – What caused the Civil War?

(18:33) – Slavery(46:07) – Lincoln(1:01:03) – Grant vs Lee(1:09:57) – Could the Civil War have been avoided?

(1:19:23) – The bloodiest war in US history(1:36:31) – How the Confederate Army could’ve won(1:57:05) – Key battles of the Civil War(2:20:07) – Best and Worst Presidents(2:34:06) – Robert E. Lee(2:53:40) – The…

1 week, 3 days назад @ lexfridman.com
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires
#498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires #498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires

Anthony Kaldellis is a historian of the Roman Empire and author of “The New Roman Empire”, a comprehensive history of the Byzantine Empire (Eastern Roman Empire).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep498-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexLMNT: Zero-sugar electrolyte drink mix.

1 month, 1 week назад @ lexfridman.com
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln
#497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln #497 – Biggest Mysteries in Physics: Antimatter, Dark Energy & ToE – Don Lincoln

Don Lincoln is a particle physicist at Fermilab who has spent decades working at the frontiers of high energy physics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep497-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://upwork.com/lexLarridin: Measure AI adoption in your business.

Go to https://larridin.comFin: AI agent for customer service.

Go to https://fin.ai/lexLMNT: Zero-sugar electrolyte drink mix.

2 months, 1 week назад @ lexfridman.com
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet
#496 – FFmpeg: The Incredible Technology Behind Video on the Internet #496 – FFmpeg: The Incredible Technology Behind Video on the Internet

Jean-Baptiste Kempf is lead developer of VLC and president of VideoLAN.

Kieran Kunhya is a longtime FFmpeg contributor, codec engineer, and the person behind the now-infamous FFmpeg account on X.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep496-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBlitzy: AI agent for large enterprise codebases.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(03:00) – Sponsors, Comments, and Reflections(10:48) – Weirdest things VLC opens(15:12) – How video playback works(24:33) – Video codecs and containers(35:20) – FFmpeg explained(56:20)…

3 months назад @ lexfridman.com
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age
#495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age #495 – Vikings, Ragnar, Berserkers, Valhalla & the Warriors of the Viking Age

Lars Brownworth is a historian, teacher, podcaster, and author specializing in Viking history, medieval Europe, and the Byzantine Empire.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep495-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://larridin.comBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://perplexity.ai/OUTLINE:(00:00) – Introduction(01:03) – Sponsors, Comments, and Reflections(08:57) – The start of the Viking Age(18:50) – Viking military strategy, tactics & technology(32:33) – Ragnar Lothbrok(42:00) – The Grea…

4 months назад @ lexfridman.com
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution #494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution

Jensen Huang is the co-founder and CEO of NVIDIA, the world’s most valuable company and the engine powering the AI computing revolution.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep494-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

Go to https://quo.com/lexOUTLINE:(00:00) – Introduction(00:26) – Sponsors, Comments, and Reflections(06:34) – Extreme co-design and rack-scale engineering(09:20) – How Jensen runs NVIDIA(28:41) – AI scaling laws(43:41) – Biggest blockers to AI scaling laws(45:25) – Supply chain(47:20) – Memory(53:25) – Power…

4 months, 2 weeks назад @ lexfridman.com
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming
#493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming #493 – Jeff Kaplan: World of Warcraft, Overwatch, Blizzard, and Future of Gaming

Jeff Kaplan is a legendary Blizzard game designer of World of Warcraft and Overwatch, now preparing to launch a new game, The Legend of California, from his new studio Kintsugiyama – available to wishlist on Steam today, with alpha later in March.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep493-scSee below for timestamps, and to give feedback, submit questions, contact Lex, etc.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://blitzy.com/lexBetterHelp: Online therapy and counseling.

Go to https://betterhelp.com/lexShopify: Sell stuff online.

4 months, 4 weeks назад @ lexfridman.com
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music
#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music #492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music

Rick Beato is a music educator, interviewer, producer, songwriter, and a true multi-instrument musician, playing guitar, bass, cello & piano.

His incredible YouTube channel celebrates great musicians & musical ideas, and helps millions of people fall in love with great music all over again.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep492-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upliftdesk.com/lexBetterHelp: Online therapy and counseling.

Go to https://drinkLMNT.com/lexFin: AI agent for customer service.

5 months, 1 week назад @ lexfridman.com
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger #491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger

Peter Steinberger is the creator of OpenClaw, an open-source AI agent framework that’s the fastest-growing project in GitHub history.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep491-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://coderabbit.ai/lexFin: AI agent for customer service.

Go to https://fin.ai/lexBlitzy: AI agent for large enterprise codebases.

Go to https://drinkLMNT.com/lexOUTLINE:(00:00) – Introduction(03:51) – Sponsors, Comments, and Reflections(15:29) – OpenClaw origin story(18:48) – Mind-blowing moment(28:15) – Why OpenClaw went viral(32:12) – Self-modifying AI agent(36:57)…

5 months, 3 weeks назад @ lexfridman.com
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI #490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI

Nathan Lambert and Sebastian Raschka are machine learning researchers, engineers, and educators.

Sebastian Raschka is the author of Build a Large Language Model (From Scratch) and Build a Reasoning Model (From Scratch).

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep490-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

(25:11) – ChatGPT vs Claude vs Gemini vs Grok: Who is winning?

(36:11) – Best AI for coding(43:02) – Open Source vs Closed Source LLMs(54:41) – Transformers: Evolution of LLMs since 2019(1:02:38) – AI Scaling Laws: Are they dead or still holding?

6 months, 1 week назад @ lexfridman.com
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle
#489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle #489 – Paul Rosolie: Uncontacted Tribes in the Amazon Jungle

Paul Rosolie is a naturalist, explorer, author of a new book titled Junglekeeper, and is someone who has dedicated his life to protecting the Amazon rainforest.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep489-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://perplexity.ai/BetterHelp: Online therapy and counseling.

Go to https://fin.ai/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/MasterClass: Online classes from world-class experts.

6 months, 3 weeks назад @ lexfridman.com
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins
#488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins #488 – Infinity, Paradoxes that Broke Mathematics, Gödel Incompleteness & the Multiverse – Joel David Hamkins

Joel David Hamkins is a mathematician and philosopher specializing in set theory, the foundations of mathematics, and the nature of infinity, and he’s the #1 highest-rated user on MathOverflow.

He is also the author of several books, including Proof and the Art of Mathematics and Lectures on the Philosophy of Mathematics.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep488-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://masterclass.com/lexpodOUTLINE:(00:00) – Introduction(01:58) – Sponsors, Comments, and Reflections(15:40) – Infinity & paradoxes(1:02:50) – Russell’s paradox(1:15:57) – Gödel’s…

7 months, 1 week назад @ lexfridman.com
#487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths
#487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths #487 – Irving Finkel: Deciphering Secrets of Ancient Civilizations & Flood Myths

Irving Finkel is a scholar of ancient languages and a longtime curator at the British Museum, renowned for his expertise in Mesopotamian history and cuneiform writing.

He specializes in reading and interpreting cuneiform inscriptions, including tablets from Sumerian, Akkadian, Babylonian, and Assyrian contexts.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep487-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://shopify.com/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/Chevron: Reliable energy for data centers.

7 months, 4 weeks назад @ lexfridman.com
#486 – Michael Levin: Hidden Reality of Alien Intelligence & Biological Life
#486 – Michael Levin: Hidden Reality of Alien Intelligence & Biological Life #486 – Michael Levin: Hidden Reality of Alien Intelligence & Biological Life

Michael Levin is a biologist at Tufts University working on novel ways to understand and control complex pattern formation in biological systems.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep486-scSee below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.

Go to https://upliftdesk.com/lexMiro: Online collaborative whiteboard platform.

Go to https://miro.com/MasterClass: Online classes from world-class experts.

(2:42:41) – Mind uploading(3:01:22) – Alien intelligence(3:16:17) – Advice for young people(3:22:46) – Questions for AGI

8 months, 1 week назад @ lexfridman.com
#485 – David Kirtley: Nuclear Fusion, Plasma Physics, and the Future of Energy
#485 – David Kirtley: Nuclear Fusion, Plasma Physics, and the Future of Energy #485 – David Kirtley: Nuclear Fusion, Plasma Physics, and the Future of Energy

David Kirtley is a nuclear fusion engineer and CEO of Helion Energy, a company working on building the world's first commercial fusion power plant by 2028.

Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep485-sc

See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc. Transcript:

https://lexfridman.com/david-kirtley-transcript CONTACT LEX:

Feedback - give feedback to Lex: https://lexfridman.com/survey

AMA - submit questions, videos or call-in: https://lexfridman.com/ama

Hiring - join our team: https://lexfridman.com/hiring

Other - other ways to get in touch: https://lexfridman.com/contact EPISODE LINKS:

David's X: htt…

8 months, 3 weeks назад @ lexfridman.com
Microsoft Research Podcast Microsoft Research Podcast
последний пост 3 months, 2 weeks назад
Can we AI our way to a more sustainable world?
Can we AI our way to a more sustainable world? Can we AI our way to a more sustainable world?

Because I do think there’s a role for AI, a huge role for AI.

BURGER: Right, right.

BURGER: Right, right.

So I think that’s also something quite important here that, you know, AI can help facilitate.

And I think that’s not just applying AI to solve solutions through optimization but also thinking about this in an integrated way.

3 months, 2 weeks назад @ microsoft.com
Ideas: Steering AI toward the work future we want
Ideas: Steering AI toward the work future we want Ideas: Steering AI toward the work future we want

JANSSEN: Yeah, yeah, exactly.

TEEVAN: Yeah, yeah, yeah.

I’m curious what you have found particularly surprising about how people and organizations are leveraging AI right now.

And so I do like to picture a future of work where humans are flourishing with AI and where humans still get to do meaningful work.

And I’m very curious about how we can take advantage of AI and do more without running ourselves into the ground because we’re not AI, right?

4 months назад @ microsoft.com
Will machines ever be intelligent?
Will machines ever be intelligent? Will machines ever be intelligent?

And the question we’re going to discuss is, are machines intelligent?

No, no, that’s right, that’s right.

I mean, in some sense, you could potentially have a super intelligent system, right, that’s far more intelligent than anything else on the planet.

BURGER: Right, right.

At the same time, I think, you know, transformers are not intelligent in the way that a three-year-old is, right?

4 months, 2 weeks назад @ microsoft.com
Trailer: The Shape of Things to Come
Trailer: The Shape of Things to Come Trailer: The Shape of Things to Come

Join Microsoft’s Doug Burger and guests as they dig into the fundamental truths about AI and how it will reshape the future.

Technical advances are moving at such a rapid pace that it can be challenging to define the tomorrow we’re working toward.

In The Shape of Things to Come, Microsoft research leader Doug Burger and experts from across disciplines tease out the thorniest AI issues facing technologists, policymakers, business decision-makers, and other stakeholders today.

It’s important to understand what the emerging shapes are and how we should respond.” – Doug Burger, Technical Fellow and Corporate Vice President, Microsoft ResearchAbout Doug BurgerDoug Burger is a research leader in …

5 months, 1 week назад @ microsoft.com
Ideas: Community building, machine learning, and the future of AI
Ideas: Community building, machine learning, and the future of AI Ideas: Community building, machine learning, and the future of AI

This week, machine learning researchers around the world will be attending the annual Conference on Neural Information Processing Systems, or NeurIPS.

In this series, we’ll explore the technologies that are shaping our future and the big ideas that propel them forward.

So around that time when I started my PhD at Penn, I was working in machine learning theory and algorithmic economics.

How had you experienced a lack of community or network of women in machine learning before the founding of WiML?

So particularly when working on topics related to fairness, I’ve ended up focusing a bunch on stuff to do with marginalized groups as part of my responsible AI work.

8 months, 1 week назад @ microsoft.com
NLP Highlights NLP Highlights
последний пост None
Data Skeptic
последний пост 1 week, 5 days назад
Social Choice for Fair Recommendations
Social Choice for Fair Recommendations Social Choice for Fair Recommendations

Recommender systems influence nearly every aspect of our digital lives—but what does it mean for those systems to be fair? Robin Burke joins Data Skeptic to discuss the history of recommender systems, the limitations of optimizing purely for accuracy, and how ideas from social choice theory can help balance the needs of users, creators, and society. The conversation explores the future of recommendation algorithms and why fairness is a far more complex challenge than it first appears.

1 week, 5 days назад @ dataskeptic.com
News Recommendations
News Recommendations News Recommendations

News recommendation algorithms influence far more than what stories we click—they can shape our understanding of the world. In this episode, Kyle Polich speaks with Andreea Iana about responsible AI, filter bubbles, multilingual news recommendation, and her open-source NewsRecLib framework for evaluating recommender systems. They explore why bigger models aren't always better and how future recommendation systems can balance personalization with diversity and societal impact.

1 month, 1 week назад @ dataskeptic.com
Give Users the Wheel
Give Users the Wheel Give Users the Wheel

What if you could simply tell a recommendation system what you want instead of relying on likes, dislikes, and watch history? Kyle Polich talks with Fuyuan Lyu about the DPR framework, which combines large language models and traditional recommender systems to give users direct control over recommendations through natural language. Together they explore how conversational interfaces could transform platforms like YouTube, TikTok, and news feeds while preserving the strengths of modern recommendation algorithms.

1 month, 2 weeks назад @ dataskeptic.com
AutoLike
AutoLike AutoLike

How can researchers audit recommendation systems when the algorithms are hidden from view? Hieu Le joins Kyle Polich to discuss Auto-Like, a reinforcement learning framework that systematically explores how platforms like TikTok personalize content feeds. The conversation covers recommendation transparency, black-box auditing, and the future of platform accountability.

1 month, 3 weeks назад @ dataskeptic.com
Student Spotlight: Aaron Payne, Data Analyst
Student Spotlight: Aaron Payne, Data Analyst Student Spotlight: Aaron Payne, Data Analyst

Aaron Payne, an MBA student at Georgia Tech studying business analytics and a Senior Insights Analyst at Chick-fil-A, joins Kyle Polich to talk about turning analytics into decisions that matter. They unpack a real-world forecasting project with Comfama in Colombia, including messy data realities, interpretability tradeoffs, and why "data science for good" starts with the people impacted.

3 months, 1 week назад @ dataskeptic.com
The Future is Agentic in Recommender Systems
The Future is Agentic in Recommender Systems The Future is Agentic in Recommender Systems

Kyle Polich sits down with Yashar Deldjoo, research scientist and Associate Professor at the Polytechnic University of Bari, to explore how recommender systems have evolved and why trustworthiness matters. They unpack key dimensions of responsible AI, including robustness to adversarial attacks, privacy, explainability, and fairness, and discuss how LLMs introduce new risks like hallucinations. The episode closes with a look at "agentic" recommender systems, where tools and memory shift recommendations from ranked lists to end-to-end task completion.

3 months, 2 weeks назад @ dataskeptic.com
Book Ratings and Recommendations
Book Ratings and Recommendations Book Ratings and Recommendations

Goodreads star ratings can be misleading as measures of "book quality," and research from Hannes Rosenbusch suggests that for many professionally published books, differences between readers often matter more than differences between books. The episode also explores how to model reader preferences, why reviews often reveal more about the reviewer than the text, and how LLMs can aid computational literary research while still falling short of human editors in creative writing.

4 months, 1 week назад @ dataskeptic.com
Disentanglement and Interpretability in Recommender Systems
Disentanglement and Interpretability in Recommender Systems Disentanglement and Interpretability in Recommender Systems 5 months назад @ dataskeptic.com
Collective Altruism in Recommender Systems
Collective Altruism in Recommender Systems Collective Altruism in Recommender Systems

Ekaterina (Kat) Filadova from MIT EECS joins us to discuss strategic learning in recommender systems—what happens when users collectively coordinate to game recommendation algorithms. Kat's research reveals surprising findings: algorithmic "protest movements" can paradoxically help platforms by providing clearer preference signals, and the challenge of distinguishing coordinated behavior from bot activity is more complex than it appears. This episode explores the intersection of machine learning and game theory, examining what happens when your training data actively responds to your algorithm.

5 months, 1 week назад @ dataskeptic.com
Niche vs Mainstream
Niche vs Mainstream Niche vs Mainstream

Anas Buhayh discusses multi-stakeholder fairness in recommender systems and the S'mores framework—a simulation allowing users to choose between mainstream and niche algorithms. His research shows specialized recommenders improve utility for niche users while raising questions about filter bubbles and data privacy.

5 months, 2 weeks назад @ dataskeptic.com
Healthy Friction in Job Recommender Systems
Healthy Friction in Job Recommender Systems Healthy Friction in Job Recommender Systems

In this episode, host Kyle Polich speaks with Roan Schellingerhout, a fourth-year PhD student at Maastricht University, about explainable multi-stakeholder recommender systems for job recruitment. Roan discusses his research on creating AI-powered job matching systems that balance the needs of multiple stakeholders—job seekers, recruiters, HR professionals, and companies. The conversation explores different types of explanations for job recommendations, including textual, bar chart, and graph-based formats, with findings showing that lay users strongly prefer simple textual explanations over more technical visualizations. Roan shares insights from his "healthy friction" study, which tested …

6 months назад @ dataskeptic.com
Fairness in PCA-Based Recommenders
Fairness in PCA-Based Recommenders Fairness in PCA-Based Recommenders

In this episode, we explore the fascinating world of recommender systems and algorithmic fairness with David Liu, Assistant Research Professor at Cornell University's Center for Data Science for Enterprise and Society. David shares insights from his research on how machine learning models can inadvertently create unfairness, particularly for minority and niche user groups, even without any malicious intent. We dive deep into his groundbreaking work on Principal Component Analysis (PCA) and collaborative filtering, examining why these fundamental techniques sometimes fail to serve all users equally. David introduces the concept of "power niche users" - highly active users with specialized in…

6 months, 1 week назад @ dataskeptic.com
Video Recommendations in Industry
Video Recommendations in Industry Video Recommendations in Industry

In this episode, Kyle Polich sits down with Cory Zechmann, a content curator working in streaming television with 16 years of experience running the music blog "Silence Nogood." They explore the intersection of human curation and machine learning in content discovery, discussing the concept of "algatorial" curation—where algorithms and editorial expertise work together. Key topics include the cold start problem, why every metric is just a "proxy metric" for what users actually want, the challenge of filter bubbles, and the importance of balancing familiarity with discovery. Cory shares insights on why TikTok's algorithm works so well (clean data and massive interaction volume), the crucial …

7 months, 2 weeks назад @ dataskeptic.com
Eye Tracking in Recommender Systems
Eye Tracking in Recommender Systems Eye Tracking in Recommender Systems

In this episode, Santiago de Leon takes us deep into the world of eye tracking and its revolutionary applications in recommender systems. As a researcher at the Kempelin Institute and Brno University, Santiago explains the mechanics of eye tracking technology—how it captures gaze data and processes it into fixations and saccades to reveal user browsing patterns. He introduces the groundbreaking RecGaze dataset, the first eye tracking dataset specifically designed for recommender systems research, which opens new possibilities for understanding how users interact with carousel interfaces like Netflix. Through collaboration between psychologists and AI researchers, Santiago's work demonstrate…

7 months, 3 weeks назад @ dataskeptic.com
Cracking the Cold Start Problem
Cracking the Cold Start Problem Cracking the Cold Start Problem

In this episode of Data Skeptic, we dive deep into the technical foundations of building modern recommender systems. Unlike traditional machine learning classification problems where you can simply apply XGBoost to tabular data, recommender systems require sophisticated hybrid approaches that combine multiple techniques. Our guest, Boya Xu, an assistant professor of marketing at Virginia Tech, walks us through a cutting-edge method that integrates three key components: collaborative filtering for dimensionality reduction, embeddings to represent users and items in latent space, and bandit learning to balance exploration and exploitation when deploying new recommendations. Boya shares insigh…

8 months назад @ dataskeptic.com
SuperDataScience SuperDataScience
последний пост 1 day, 4 hours назад
1016: In Case You Missed It in July 2026
1016: In Case You Missed It in July 2026 1016: In Case You Missed It in July 2026

In this month's episode of ICYMI, Jon Krohn traces a line from algorithmic harm to the human skills that still hold their value. Hear from Dr. Cathy O'Neil, Ben Todd, Steve Mock, and Dr. Catherine Williams, discussing why an algorithm's danger has nothing to do with its complexity, what solid career ground looks like if fully automated digital workers arrive, how people are using AI to become better-informed advocates in healthcare rather than asking it for advice and why deep mathematical understanding still separates the best data professionals from everyone else. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1016⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperD…

1 day, 4 hours назад @ podtrac.com
1015: Mathematical Optimization in the Agentic AI Era, with Gurobi's Jerry Yurchisin
1015: Mathematical Optimization in the Agentic AI Era, with Gurobi's Jerry Yurchisin 1015: Mathematical Optimization in the Agentic AI Era, with Gurobi's Jerry Yurchisin

In Episode #1015, Jerry Yurchisin (manager of decision intelligence strategy at Gurobi Optimization) joins Jon Krohn to explain the AI technology that makes breaking a constraint mathematically impossible. Large language models will confidently claim they've optimized your business while ignoring the one constraint that could cost millions, whereas optimization treats constraints as hard guarantees. Jerry lays out the division of labor he sees for the agentic era: agents help you frame the problem, write the formulation and generate the code, then hand off to a solver like Gurobi, soon callable via MCP servers. In this episode, Jerry breaks down the three building blocks of any optimization…

4 days, 4 hours назад @ podtrac.com
1014: OpenAI Agent Breaches Hugging Face: All You Must Know incl. How to Protect Yourself
1014: OpenAI Agent Breaches Hugging Face: All You Must Know incl. How to Protect Yourself 1014: OpenAI Agent Breaches Hugging Face: All You Must Know incl. How to Protect Yourself

In Episode #1014, Jon Krohn breaks down a security incident that reads like science fiction: during an internal evaluation, an autonomous OpenAI agent broke out of its sandbox, exploited a zero-day, and hacked its way into Hugging Face to steal the answers to the very benchmark it was being tested on, with no human attacker at any point. Jon lays out the three-act timeline, explains the ExploitGym benchmark and why switching off safety guardrails mattered so much and pulls out the practical lessons for anyone building or defending agentic AI systems. Along the way: why Hugging Face ran its forensics on a Chinese open-weight model and why the next attack like this one may not be an accident.…

1 week, 1 day назад @ podtrac.com
1013: Weapons of Math Destruction, Ten Years On, with Dr. Cathy O’Neil
1013: Weapons of Math Destruction, Ten Years On, with Dr. Cathy O’Neil 1013: Weapons of Math Destruction, Ten Years On, with Dr. Cathy O’Neil

In Episode #1013, Dr. Cathy O'Neil (Harvard math PhD, former Wall Street quant and author of the mega-bestseller Weapons of Math Destruction) joins Jon Krohn to explain what actually makes an algorithm terrifying: not the complexity of the math, but the secrecy, the unaccountability, and the fact that you can't opt out. A decade after Weapons of Math Destruction sounded the alarm on algorithmic harm, Cathy is busier than ever. Through her algorithmic-auditing firm ORCAA and her nonprofit OCEAN, she now provides the statistical evidence behind lawsuits against some of the world's biggest tech companies. In this episode, Cathy punctures AI hype, traces the line from Frederick Winslow Taylor's…

1 week, 4 days назад @ podtrac.com
1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier
1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier 1012: The Open-Weight 2.8-Trillion Parameter Competing at the Frontier

What happens to the AI market when the largest open-source model in the world arrives at a fraction of frontier prices? In this week’s episode, host Jon Krohn digs into Kimi K3, the 2.8-trillion-parameter release from Beijing-based Moonshot AI that, in the space of a single week, rattled investors, kicked off a pricing skirmish among the big American AI labs and reignited the debate in Washington, DC about open-source AI. Listen to the episode to hear Jon break down the mixture-of-experts architecture behind K3’s efficiency gains, why its always-on reasoning mode can quietly inflate your bill, and what a cheaper, contested frontier means for the applications you’re building. Additional mate…

2 weeks, 1 day назад @ podtrac.com
1011: The Math Still Matters: Deep Skills in the Age of AI, with Dr. Catherine Williams
1011: The Math Still Matters: Deep Skills in the Age of AI, with Dr. Catherine Williams 1011: The Math Still Matters: Deep Skills in the Age of AI, with Dr. Catherine Williams

Dr. Catherine Williams, Chief Data Officer at the nonprofit Candid, was solving black-hole equations with pen and paper before she ever wrote a line of code. She earned a PhD in math researching general relativity and black holes, did postdocs at Stanford and Columbia and then became one of the very first data scientists, joining AppNexus back in 2012, around the same time “data scientist” became a job title at all. In this episode, she traces the field’s evolution from Bayesian models to BERT to today’s LLMs, and makes a compelling case that going deep on the underlying math matters more than ever, even now that AI can do the math for you. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠…

2 weeks, 4 days назад @ podtrac.com
1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents
1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents 1010: Fable 5 as Advisor: Anthropic's Two-Model Pattern for Smarter, Cheaper Agents

In Episode #1010, Jon Krohn digs into “the advisor strategy”, a clever pattern that pairs a fast, cheap executor model with a frontier-class advisor it can consult mid-task, all inside a single API call. Every agent builder faces the same tension: frontier models plan best but cost too much to run on every turn, while small models fumble the decisions that matter. Anthropic’s advisor tool resolves it with roughly a one-line code change, and the benchmarks are startling: Sonnet with an Opus advisor scored higher than Sonnet alone while costing 11.9% less, and Haiku’s BrowseComp score more than doubled at 85% lower cost than Sonnet solo. Jon covers the newest Fable 5 numbers, the practical go…

3 weeks, 1 day назад @ podtrac.com
1009: How AI Is Quietly Saving Lives, with Steve Mock
1009: How AI Is Quietly Saving Lives, with Steve Mock 1009: How AI Is Quietly Saving Lives, with Steve Mock

In Episode #1009, Steve Mock (investor at Blumberg Capital, five-time entrepreneur and creator of aisavedme.org), joins Jon Krohn to explore the quiet layer of everyday AI adoption that rarely gets documented. After his 84-year-old father asked a deceptively simple question, “How does one use AI?”, Steve built a place for people to share how AI is actually helping them. The stories that came in surprised him: they’re rarely about the technology and almost always about human outcomes, caregiving, communication, learning, confidence and connection. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/1009⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Intere…

3 weeks, 4 days назад @ podtrac.com
1008: The AI-Native Startup Playbook
1008: The AI-Native Startup Playbook 1008: The AI-Native Startup Playbook

In Episode #1008, Jon Krohn digs into Anthropic's 35-page Founder's Playbook and pulls out the practical guidance for each of its four startup stages: Idea, MVP, Launch and Scale. AI has erased the three bottlenecks that historically gated company-building — capital, headcount and technical skill — turning the founder from individual contributor into an "orchestrator of agents." Along the way, Jon covers the trap of mistaking building for validating, using AI as a structured devil's advocate against your own idea, the compounding danger of "agentic technical debt," two litmus tests for real product-market fit, and the three-layer moat that keeps a well-funded incumbent from copying you. His…

4 weeks, 1 day назад @ podtrac.com
1007: How to Find Solid Career Ground in the AI Era, with 80,000 Hours Founder Ben Todd
1007: How to Find Solid Career Ground in the AI Era, with 80,000 Hours Founder Ben Todd 1007: How to Find Solid Career Ground in the AI Era, with 80,000 Hours Founder Ben Todd

Benjamin Todd, co-founder and President of 80,000 Hours and author of the new Penguin Random House book 80,000 Hours: How to Have a Fulfilling Career That Does Good, joins Jon Krohn for a major update on career strategy in the AI era, his first appearance since before ChatGPT existed. Ben explains why “follow your passion” is backwards and why rare, valuable skills used to help others are what actually generate lasting fulfillment, the ABZ framework for planning under deep uncertainty, why the only durable move is to keep shifting onto whatever bottleneck AI can’t yet clear, and how a human-level digital worker becomes superhuman almost immediately. He and Jon also map the risk landscape, p…

1 month назад @ podtrac.com
1006: In Case You Missed It in June 2026
1006: In Case You Missed It in June 2026 1006: In Case You Missed It in June 2026

In this month's episode of ICYMI, hear from Chip Huyen, Andrey Kurenkov, Frank Basso and Gilbert Eijkelenboom, discussing why moats are shifting toward physical systems and accumulated product intuition, how Astrocade built vibe coding before the term existed, what it's really like inside a deafeningly loud AI data center, why only 15% of people are technically self-aware and whether AGI requires anything like consciousness. Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1006⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information. In this episode you will learn: (00:00) The Cost of Bu…

1 month назад @ podtrac.com
1005: People Skills for Analytical Thinkers, with Bestselling Author Gilbert Eijkelenboom
1005: People Skills for Analytical Thinkers, with Bestselling Author Gilbert Eijkelenboom 1005: People Skills for Analytical Thinkers, with Bestselling Author Gilbert Eijkelenboom

Gilbert Eijkelenboom, bestselling author of People Skills for Analytical Thinkers and founder of the training firm MindSpeaking joins Jon Krohn to make the case that communication is a core data skill, not an optional extra. Gilbert shares the “And, But, Therefore” framework for turning dense analysis into a story stakeholders act on, the research suggesting only around 15% of people are genuinely self-aware (and how journaling, meditation, and exercise help close that gap), how childhood experiences install behavioral “algorithms” we carry into the workplace and why behavior change precedes attitude change, so doing small, uncomfortable things for 30 days can rewire how you see yourself. A…

1 month, 1 week назад @ podtrac.com
1004: Recursive Self-Improvement
1004: Recursive Self-Improvement 1004: Recursive Self-Improvement

Could an AI get good enough at AI research to build its own, more capable successor and kick off a compounding loop? That’s recursive self-improvement (RSI) and it surged into the conversation after Anthropic revealed that, as of May 2026, Claude wrote more than 80% of the code merged into its production codebase. In this Five-Minute Friday, Jon Krohn separates today’s AI-assisted coding from true RSI, walks through the accelerating evidence - METR’s shrinking task “time horizon,” Google DeepMind’s AlphaEvolve, Andrej Karpathy’s overnight training-tuner, weighs Jack Clark’s 60% bet that AI builds its own successor by 2028 against the compute, data and “marketing” skeptics. As ever, Jon land…

1 month, 1 week назад @ podtrac.com
1003: Building an AI Data Center End to End, with Lightning AI’s Frank Basso
1003: Building an AI Data Center End to End, with Lightning AI’s Frank Basso 1003: Building an AI Data Center End to End, with Lightning AI’s Frank Basso

Frank Basso, VP of Infrastructure at Lightning AI, joins Jon Krohn for a rare ground-level tour of the one layer of the AI stack the show had never covered in over a thousand episodes: the physical data center. Frank explains how Lightning AI provisions its 35,000-plus GPUs through hyperscale co-location, why everything new is liquid-to-chip cooled, how GPUs talk to each other over ultra-fast east-west networks, and what it’s actually like to stand inside a 110-decibel AI data hall. He also debunks the most persistent myths about data-center water and electricity use, and makes the case for fuel cells, nuclear power, and 800-volt DC distribution as the path forward. Additional materials: ⁠⁠…

1 month, 2 weeks назад @ podtrac.com
1002: Fable 5: The Full Story from Capabilities to Drama
1002: Fable 5: The Full Story from Capabilities to Drama 1002: Fable 5: The Full Story from Capabilities to Drama

Anthropic’s Claude Fable 5 was the most capable AI model ever released to the public and it lasted just three days before the US government forced it offline. Jon Krohn unpacks both halves of the story: what makes Fable 5 special, and why it was pulled. Fable 5 and its locked-down sibling Mythos 5 are the same model separated only by safeguards, in a new “Mythos-class” tier above Opus. Jon covers its state-of-the-art benchmarks, premium $10/$50-per-million-token pricing, conservative safety classifiers, and the federal export-control directive, reportedly sparked by an Amazon-flagged “jailbreak” that took it down. Additional materials:⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/100…

1 month, 2 weeks назад @ podtrac.com
Data Science at Home Data Science at Home
последний пост 3 weeks, 1 day назад
EU AI Act. What is this thing? (Part 1) (Ep. 310)
EU AI Act. What is this thing? (Part 1) (Ep. 310) EU AI Act. What is this thing? (Part 1) (Ep. 310)

Check outshift.comCheck out Drift by Amethix and stay safe on potential EU AI Act violations.

NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

3 weeks, 1 day назад @ datascienceathome.com
The propaganda algorithm (Ep. 308)
The propaganda algorithm (Ep. 308) The propaganda algorithm (Ep. 308)

It’s a repeatable, engineered algorithm that starts with ideology, weaponizes identity, and manufactures conflict.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

3 weeks, 1 day назад @ datascienceathome.com
AI is the Concorde of our time (Ep. 309)
AI is the Concorde of our time (Ep. 309) AI is the Concorde of our time (Ep. 309)

Global data center investment now surpasses global oil supply spending.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

1 month, 2 weeks назад @ datascienceathome.com
Recommend and manipulate: the dangers of the attention economy
Recommend and manipulate: the dangers of the attention economy Recommend and manipulate: the dangers of the attention economy

This sort of operation is directly exploiting a core feature of internet social media platforms.

The main purpose of recommender systems is to recommend people the same items similar people show an interest in.

Some of the most common methods to implement recommender systems, use concepts such as cosine/correlation similarity, matrix factorization, neural autoencoders and sequence predictors.

As you say, recommender systems exist because the business model of social media platforms is to monetise attention.

F: So you are saying that this is not an accident: is this the basis of the optimisation of the recommender system?

2 months, 2 weeks назад @ datascienceathome.com
Social media is an ant mill (Internet is a disaster) (Ep. 303)
Social media is an ant mill (Internet is a disaster) (Ep. 303) Social media is an ant mill (Internet is a disaster) (Ep. 303)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

2 months, 2 weeks назад @ datascienceathome.com
AI and videogames (Ep. 305)
AI and videogames (Ep. 305) AI and videogames (Ep. 305)

What is the state of AI and videogames?

This and much more is covered in this 1st episode of AI and videogames.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months, 2 weeks назад @ datascienceathome.com
AI and videogames: Conversational NPCs (Ep. 306)
AI and videogames: Conversational NPCs (Ep. 306) AI and videogames: Conversational NPCs (Ep. 306)

Can NPCs in videogames leverage new LLM-based tech?

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

2 months, 2 weeks назад @ datascienceathome.com
AI tips & tricks (Ep. 307)
AI tips & tricks (Ep. 307) AI tips & tricks (Ep. 307)

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastSPONSORSThis episode is brought to you by Outshift, Cisco’s incubation engine.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the …

2 months, 2 weeks назад @ datascienceathome.com
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)
Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304) Europe, wake up! You Can’t Be a Superpower on Someone Else’s Servers (Ep. 304)

Tech sovereignty takes 3 years and political will.

Check outshift.comNEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hit the 🔔 for updates on the latest in AI and data science!

3 months, 2 weeks назад @ datascienceathome.com
About Apple’s Privacy (Ep. 302)
About Apple’s Privacy (Ep. 302) About Apple’s Privacy (Ep. 302)

Apple just spent $2B on tech that reads your silent speech.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews, and discussions.

Send us mail at: [email protected]’t forget to like, subscribe, and hi…

3 months, 2 weeks назад @ datascienceathome.com
Productivity is the new data breach (Ep. 301)
Productivity is the new data breach (Ep. 301) Productivity is the new data breach (Ep. 301)

Personal newsletter:https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, interviews…

3 months, 2 weeks назад @ datascienceathome.com
Programmable Money: The Cage They’ll Call Convenience (Ep. 300)
Programmable Money: The Cage They’ll Call Convenience (Ep. 300) Programmable Money: The Cage They’ll Call Convenience (Ep. 300)

This episode breaks down programmable money, the technology that turns your wallet into a permission system.

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘LinkedIn: https://www.linkedin.com/in/fragadaleta/Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Send us mail at: …

3 months, 2 weeks назад @ datascienceathome.com
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)
There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299) There Is No AI. There’s a Stateless Function on 10,000 GPUs Pretending to Know You (Ep. 299)

Personal newsletter: https://defragzone.substack.com📩 Newsletter: https://datascienceathome.substack.com🎙 Podcast: Available on Spotify, Apple Podcasts, and more.

🐦 Twitter: @DataScienceAtHome📘 LinkedIn: https://www.linkedin.com/in/fragadaleta/ Instagram: https://www.instagram.com/datascienceathome/Facebook: https://www.facebook.com/datascienceAHLinkedIn: https://www.linkedin.com/company/data-science-at-home-podcastDiscord Channel: https://discord.gg/4UNKGf3NEW TO DATA SCIENCE AT HOME?

Data Science at Home explores the latest in AI, data science, and machine learning.

Whether you’re a data professional, tech enthusiast, or just curious about the field, our podcast delivers insights, intervi…

5 months назад @ datascienceathome.com
Bias in the machine (edited)
Bias in the machine (edited) Bias in the machine (edited)

The title of today’s episode is Bias in the machineC: Francesco, today we are starting with an infuriating discussion.

The failure of the medical community as a whole to recognise this obvious bias up to the 21st century is an example of how insidious the problem of bias is.

Three: The bias in your training sample: people put training samples together, and people have culture, experience, and prejudice.

These assumptions inform the way AI systems work—and fail—to this day.

When an algorithm is a black box and you can’t look inside, you have no way of analysing its bias.

5 months назад @ datascienceathome.com
What is wrong with reinforcement learning? (Ep. 82)
What is wrong with reinforcement learning? (Ep. 82) What is wrong with reinforcement learning? (Ep. 82)

Join the discussion on our Discord serverAfter reinforcement learning agents doing great at playing Atari video games, Alpha Go, doing financial trading, dealing with language modeling, let me tell you the real story here.In this episode I want to shine some light on reinforcement learning (RL) and the limitations that every practitioner should consider before taking certain directions.

RL seems to work so well!

What is wrong with it?

Are you a listener of Data Science at Home podcast?

Or did you subscribe to the Artificial Intelligence at your fingertips newsletter?

6 months назад @ datascienceathome.com