Oxen.ai Blog

Welcome to the Oxen.ai blog 🐂. The Oxen team is dedicated to helping AI creators and builders go from research to production. We love to learn and share our learnings with the community. Sign up for any of our future events on our luma page.

Oxen's Model Report - July 23rd, 2026
Oxen's Model Report - July 23rd, 2026

Every time I decide to write one of these I come up with an initial shortlist. By the time I actually sit down to write, a couple of days have usually passed and I have to complete...

Eloy Martinez
Eloy Martinez
7/23/2026
18 min read
The Ultimate Guide to Topaz Upscaling Video Models: Proteus, Starlight, Wonder, and Hyperion
The Ultimate Guide to Topaz Upscaling Video Models: Proteus, Starlight, Wonder, and Hyperion

The hardest part of AI video isn't pressing generate. It's finishing. How do you get high-quality 4K HDR content into your post-production pipeline? That's where models like the To...

Greg Schoeninger
Greg Schoeninger
7/13/2026
7 min read
Everything new in Oxen.ai - July 7th, 2026
Everything new in Oxen.ai - July 7th, 2026

Hey Herd, It's been awhile since we've done a proper feature update, so wanted to come up for air to show you the latest and greatest. The Oxen team has been heads down plowing th...

Greg Schoeninger
Greg Schoeninger
7/7/2026
6 min read
Oxen.ai vs Higgsfield: pay-as-you-go AI video and image models
Oxen.ai vs Higgsfield: pay-as-you-go AI video and image models

Higgsfield is great for solo creators who want a polished, creative-first studio. Its curated cinematic presets and effects get you to a good-looking result quickly, all wrapped i...

Eloy Martinez
Eloy Martinez
6/30/2026
4 min read
Oxen.ai vs Magnific (FreePik): which AI creative platform should you use?
Oxen.ai vs Magnific (FreePik): which AI creative platform should you use?

Magnific and Oxen.ai both put the top AI image and video models in one place, alongside upscaling and editing tools. If you are choosing between them, the differences come down to ...

Eloy Martinez
Eloy Martinez
6/30/2026
3 min read
Oxen's Model Report - May 8th, 2026
Oxen's Model Report - May 8th, 2026

Welcome back to another iteration of everybody's favorite moooodel report. Every time I sit down to write one of these I'm shocked by how much there is to cover (as you might've se...

Eloy Martinez
Eloy Martinez
5/8/2026
13 min read
Writing a fine-tuning and deployment pipeline isn't as easy as it looks (Gemma 4 Version)
Writing a fine-tuning and deployment pipeline isn't as easy as it looks (Gemma 4 Version)

🚹You're about to embark on a journey of ups and downs and many aha moments. We're pulling back the curtain on every painful hour so you never have to spend them yourself. Fine-tun...

Eloy Martinez
Eloy Martinez
4/16/2026
- engineering
19 min read
Oxen's Model Report - April 9th, 2026
Oxen's Model Report - April 9th, 2026

Welcome back to another iteration of everybody’s favorite moooodel report. In AI, any given day feels like a decade, and the past couple of weeks have felt like a couple of centuri...

Eloy Martinez
Eloy Martinez
4/9/2026
9 min read
When to Fine-Tune an Image Model
When to Fine-Tune an Image Model

💡Curious about fine-tuning multi-modal models? This Friday, we're diving into the new Qwen3.5 series, what makes it great and how to train it on images and video. Join us live for...

Eloy Martinez
Eloy Martinez
4/1/2026
7 min read
Oxen's Model Report - March 11th, 2026
Oxen's Model Report - March 11th, 2026

Welcome back to another iteration of our favorite moooodel report. This week we've got an absolutely packed lineup, from motion-controlled video generation to open-weight language ...

Eloy Martinez
Eloy Martinez
3/12/2026
8 min read
Frank's Red Hot
Frank's Red Hot

0:00 /0:45 1× An AI generated goat rapping alongside Ludacris in Frank's RedHot's "Eat The GOAT" Super Bowl ad. The goat was fully genera...

Eloy Martinez
Eloy Martinez
2/20/2026
- Showcase
0
Isometric.nyc
Isometric.nyc

A giant isometric pixel-art map of New York City, inspired by SimCity 2000 and Rollercoaster Tycoon. Andy Coenen fine-tuned an image model on Oxen with just 40 training examples to...

Eloy Martinez
Eloy Martinez
2/20/2026
- Showcase
1 min read
Bell Canada
Bell Canada

0:00 /0:30 1× A fully AI generated commercial for the Canadian telecom giant Bell. Made in collaboration with KnuckleheadTV and amazing ...

Eloy Martinez
Eloy Martinez
2/20/2026
- Showcase
1 min read
Oxen's Model Report
Oxen's Model Report

Welcome to this week's Oxen moooodel report. We know the AI space moves like crazy. There's a new model, paper, podcast, or hot take every single day. To help y'all keep up, we're ...

Eloy Martinez
Eloy Martinez
2/17/2026
8 min read
How a $1 Qwen3-VL Fine-Tune Beat Gemini 3
How a $1 Qwen3-VL Fine-Tune Beat Gemini 3

Can a $1 fine-tune beat a state-of-the-art closed-source model? ModelAccuracyTime (98 samples)Cost/RunBase Qwen3-VL-8B54.1%~10 sec$0.003Gemini 3 Flash82.7%2 min 46 sec$0.016FT Q...

Eloy Martinez
Eloy Martinez
2/11/2026
15 min read
How to Train a LTX-2 Character LoRA with Oxen.ai
How to Train a LTX-2 Character LoRA with Oxen.ai

LTX-2 is a video generation model, that not only can generation video frames, but audio as well. This model is fully open source, meaning the weights and the code are available for...

Greg Schoeninger
Greg Schoeninger
1/28/2026
11 min read
How to Use WAN 2.1-VACE to Generate Hollywood-Level Video Edits
How to Use WAN 2.1-VACE to Generate Hollywood-Level Video Edits

Imagine you are shooting a film and you realize that you have the actor wearing the wrong jacket in a scene. Do you bring the whole cast back in to re-shoot? Depending on the actor...

Greg Schoeninger
Greg Schoeninger
11/6/2025
9 min read
How We Cut Inference Costs from $46K to $6.5K Fine-Tuning Qwen-Image-Edit
How We Cut Inference Costs from $46K to $6.5K Fine-Tuning Qwen-Image-Edit

At Oxen.ai, we think a lot about what it takes to run high-quality inference at scale. It’s one thing to produce a handful of impressive results with a cutting-edge image editing m...

Eloy Martinez
Eloy Martinez
10/26/2025
13 min read
How to Set Noise Timesteps When Fine-Tuning Diffusion Models for Image Generation
How to Set Noise Timesteps When Fine-Tuning Diffusion Models for Image Generation

Fine-tuning Diffusion Models such as Stable Diffusion, FLUX.1-dev, or Qwen-Image can give you a lot of bang for your buck. Base models may not be trained on a certain concept or st...

Greg Schoeninger
Greg Schoeninger
10/10/2025
- Practical ML
6 min read
Fine-Tuned Qwen-Image-Edit vs Nano-Banana and FLUX Kontext Dev
Fine-Tuned Qwen-Image-Edit vs Nano-Banana and FLUX Kontext Dev

Welcome back to Fine-Tuning Friday, where each week we try to put some models to the test and see if fine-tuning an open-source model can outperform whatever state of the art (SOTA...

Greg Schoeninger
Greg Schoeninger
9/27/2025
12 min read
We Fine-Tuned GPT OSS 20B to Rap Like Eminem
We Fine-Tuned GPT OSS 20B to Rap Like Eminem

OpenAI came out with GPT-OSS 120B and 20B in August 2025. The first “Open” LLMs from OpenAI since GPT-2, over six years ago. The idea of fine-tuning a frontier OpenAI model was exc...

Greg Schoeninger
Greg Schoeninger
9/3/2025
14 min read
How We're Building a “Tab Tab” Code Completion Model
How We're Building a “Tab Tab” Code Completion Model

Welcome to Fine-Tuning Fridays, where we share our learnings from fine-tuning open source models for real world tasks. We’ll walk you through what models work, what models don’t an...

Greg Schoeninger
Greg Schoeninger
7/29/2025
8 min read
How to Fine-Tune a FLUX.1-dev LoRA with Code, Step by Step
How to Fine-Tune a FLUX.1-dev LoRA with Code, Step by Step

FLUX.1-dev is one of the most popular open-weight models available today. Developed by Black Forest Labs, it has 12 billion parameters. The goal of this post is to provide a barebo...

Greg Schoeninger
Greg Schoeninger
6/28/2025
- Fine-Tune Fridays
20 min read
How to Fine-Tune PixArt to Generate a Consistent Character
How to Fine-Tune PixArt to Generate a Consistent Character

Can we fine-tune a small diffusion transformer (DiT) to generate OpenAI-level images by distilling off of OpenAI images? The end goal is to have a small, fast, cheap model that we ...

Greg Schoeninger
Greg Schoeninger
6/19/2025
- Fine-Tune Fridays
21 min read
How to Fine-Tune Qwen3 on Text2SQL to GPT-4o level performance
How to Fine-Tune Qwen3 on Text2SQL to GPT-4o level performance

Welcome to a new series from the Oxen.ai Herd called Fine-Tuning Fridays! Each week we will take an open source model and put it head to head against a closed source foundation mod...

Greg Schoeninger
Greg Schoeninger
5/28/2025
- Fine-Tune Fridays
15 min read
Fine-Tuning Fridays
Fine-Tuning Fridays

Fine-tuning Fridays is a series from the Oxen.ai Herd where each week we do a deep dive into a model and put it head to head against the competition. We will be giving you practica...

Greg Schoeninger
Greg Schoeninger
5/16/2025
- Fine-Tune Fridays
2 min read
How RWKV-7 Goose Works đŸȘż + Notes from the Author
How RWKV-7 Goose Works đŸȘż + Notes from the Author

In this special Arxiv Dive, we're joined by Eugene Cheah - author, lead in RWKV org, CEO of Featherless AI, to discuss the development process and key decisions behind these models...

Greg Schoeninger
Greg Schoeninger
4/15/2025
- Arxiv Dives
17 min read
How Phi-4 Cracked Small Multimodality
How Phi-4 Cracked Small Multimodality

Phi-4 extends the existing Phi model’s capabilities by adding vision and audio all in the same model. This means you can do everything from understand images, generate code, recogn...

Greg Schoeninger
Greg Schoeninger
3/26/2025
- Arxiv Dives
8 min read
Training a Rust 1.5B Coder LM with Reinforcement Learning (GRPO)
Training a Rust 1.5B Coder LM with Reinforcement Learning (GRPO)

Group Relative Policy Optimization (GRPO) has proven to be a useful algorithm for training LLMs to reason and improve on benchmarks. DeepSeek-R1 showed that you can bootstrap a mod...

Greg Schoeninger
Greg Schoeninger
3/6/2025
- Practical ML
16 min read
Why GRPO is Important and How it Works
Why GRPO is Important and How it Works

Last week on Arxiv Dives we dug into research behind DeepSeek-R1, and uncovered that one of the techniques they use in the their training pipeline is called Group Relative Policy O...

Greg Schoeninger
Greg Schoeninger
2/12/2025
- Arxiv Dives
12 min read
🧠 GRPO VRAM Requirements For the GPU Poor
🧠 GRPO VRAM Requirements For the GPU Poor

Since the release of DeepSeek-R1, Group Relative Policy Optimization (GRPO) has become the talk of the town for Reinforcement Learning in Large Language Models due to its effective...

Greg Schoeninger
Greg Schoeninger
2/6/2025
- Practical ML
9 min read
How DeepSeek R1, GRPO, and Previous DeepSeek Models Work
How DeepSeek R1, GRPO, and Previous DeepSeek Models Work

In January 2025, DeepSeek took a shot directly at OpenAI by releasing a suite of models that “Rival OpenAI’s o1.” From their website: In the spirit of Arxiv Dives we are going to...

Greg Schoeninger
Greg Schoeninger
2/4/2025
- Arxiv Dives
15 min read
No Hype DeepSeek-R1 Reading List
No Hype DeepSeek-R1 Reading List

DeepSeek-R1 is a big step forward in the open model ecosystem for AI with their latest model competing with OpenAI's o1 on a variety of metrics. There is a lot of hype, and a lot o...

Greg Schoeninger
Greg Schoeninger
1/30/2025
- Arxiv Dives
27 min read
Oxen v0.25.0 Migration
Oxen v0.25.0 Migration

Today we released oxen v0.25.0 🎉 which comes with a few performance optimizations, including how we traverse the Merkle Tree to find files and folders. The main improvement is how...

Greg Schoeninger
Greg Schoeninger
1/28/2025
3 min read
đŸŒČ Merkle Tree VNodes
đŸŒČ Merkle Tree VNodes

In this post we peel back some of the layers of Oxen.ai’s Merkle Tree and show how we make it suitable for projects with large directories. If you are unfamiliar with Merkle Trees ...

Greg Schoeninger
Greg Schoeninger
1/27/2025
8 min read
đŸŒČ Merkle Tree 101
đŸŒČ Merkle Tree 101

Intro Merkle Trees are important data structures for ensuring integrity, deduplication, and verification of data at scale. They are used heavily in tools such as Git, Bitcoin, IPF...

Greg Schoeninger
Greg Schoeninger
1/27/2025
9 min read
arXiv Dive: RAGAS - Retrieval Augmented Generation Assessment
arXiv Dive: RAGAS - Retrieval Augmented Generation Assessment

RAGAS is an evaluation framework for Retrieval Augmented Generation (RAG). A paper released by Exploding Gradients, AMPLYFI, and CardiffNLP. RAGAS gives us a suite of metrics that ...

Greg Schoeninger
Greg Schoeninger
1/21/2025
- Arxiv Dives
13 min read
The Best AI Data Version Control Tools [2025]
The Best AI Data Version Control Tools [2025]

Data is often seen as static. It's common to just dump your data into S3 buckets in tarballs or upload to Hugging Face and leave it at that. Yet nowadays, data needs to evolve and ...

Greg Schoeninger
Greg Schoeninger
12/27/2024
6 min read
OpenCoder: The OPEN Cookbook For Top-Tier Code LLMs
OpenCoder: The OPEN Cookbook For Top-Tier Code LLMs

Welcome to the last arXiv Dive of 2024! Every other week we have been diving into interesting research papers in AI/ML. In this blog we’ll be diving into Open Coder, a paper and co...

Greg Schoeninger
Greg Schoeninger
12/24/2024
- Arxiv Dives
14 min read
LLaVA-CoT: Let Vision Language Models Reason Step-By-Step
LLaVA-CoT: Let Vision Language Models Reason Step-By-Step

When it comes to large language models, it is still the early innings. Many of them still hallucinate, fail to follow instructions, or generally don’t work. The only way to combat ...

Greg Schoeninger
Greg Schoeninger
12/10/2024
- Arxiv Dives
12 min read
How Upcycling MoEs Beat Dense LLMs
How Upcycling MoEs Beat Dense LLMs

In this Arxiv Dive, Nvidia researcher, Ethan He, presents his co-authored work Upcycling LLMs in Mixture of Experts (MoE). He goes into what a MoE is, the challenges behind upcycli...

Greg Schoeninger
Greg Schoeninger
11/19/2024
- Arxiv Dives
1 min read
Thinking LLMs: General Instruction Following with Thought Generation
Thinking LLMs: General Instruction Following with Thought Generation

The release of OpenAI-O1 has motivated a lot of people to think deeply about
thoughts 💭. Thinking before you speak is a skill that some people have better than others 😉, but a sk...

Greg Schoeninger
Greg Schoeninger
11/11/2024
- Arxiv Dives
14 min read
The Prompt Report Part 2: Plan and Solve, Tree of Thought, and Decomposition Prompting
The Prompt Report Part 2: Plan and Solve, Tree of Thought, and Decomposition Prompting

In the last blog, we went over prompting techniques 1-3 of The Prompt Report. This arXiv Dive, we were lucky to have the authors of the paper join us to go through some of the more...

Greg Schoeninger
Greg Schoeninger
10/31/2024
- Arxiv Dives
17 min read
The Prompt Report Part 1: A Systematic Survey of Prompting Techniques
The Prompt Report Part 1: A Systematic Survey of Prompting Techniques

For this blog we are switching it up a bit. In past Arxiv Dives, we have gone deep into the underlying model architectures and techniques that make large language models and other ...

Greg Schoeninger
Greg Schoeninger
10/9/2024
- Arxiv Dives
12 min read
arXiv Dive: How Flux and Rectified Flow Transformers Work
arXiv Dive: How Flux and Rectified Flow Transformers Work

Flux made quite a splash with its release on August 1st, 2024 as the new state of the art generative image model outperforming SDXL, SDXL-Turbo, Pixart, and DALL-E. While the model...

Greg Schoeninger
Greg Schoeninger
9/18/2024
- Arxiv Dives
9 min read
How Well Can Llama 3.1 8B Detect Political Spam? [4/4]
How Well Can Llama 3.1 8B Detect Political Spam? [4/4]

It only took about 11 minutes to fine-tuned Llama 3.1 8B on our political spam synthetic dataset using ReFT. While this is extremely fast, beating out our previous record of 14 min...

Eric Laurence
Eric Laurence
9/14/2024
3 min read
Fine-Tuning Llama 3.1 8B in Under 12 Minutes [3/4]
Fine-Tuning Llama 3.1 8B in Under 12 Minutes [3/4]

Meta has recently released Llama 3.1, including their 405 billion parameter model which is the most capable open model to date and the first open model on the same level as GPT 4. ...

Eric Laurence
Eric Laurence
9/5/2024
3 min read
arXiv Dive: How Meta Trained Llama 3.1
arXiv Dive: How Meta Trained Llama 3.1

Llama 3.1 is a set of Open Weights Foundation models released by Meta, which marks the first time an open model has caught up to GPT-4, Anthropic, or other closed models in the eco...

Greg Schoeninger
Greg Schoeninger
8/27/2024
- Arxiv Dives
12 min read
How to De-duplicate and Clean Synthetic Data [2/4]
How to De-duplicate and Clean Synthetic Data [2/4]

Synthetic data has shown promising results for training and fine tuning large models, such as Llama 3.1 and the models behind Apple Intelligence, and to produce datasets from minim...

Eric Laurence
Eric Laurence
8/23/2024
6 min read
Create Your Own Synthetic Data With Only 5 Political Spam Texts [1/4]
Create Your Own Synthetic Data With Only 5 Political Spam Texts [1/4]

With the 2024 elections coming up, spam and political texts are more prevalent than ever as political campaigns increasingly turn towards texting potential voters. Over 15 billion ...

Eric Laurence
Eric Laurence
8/1/2024
5 min read
Fine-tuning Llama 3 in 14 minutes using ReFT
Fine-tuning Llama 3 in 14 minutes using ReFT

If you have been fine-tuning models recently, you have most likely used LoRA. While LoRA has been the dominant PEFT technique for a long time thanks to its efficiency and effective...

Eric Laurence
Eric Laurence
7/25/2024
8 min read
ArXiv Dives: How ReFT works
ArXiv Dives: How ReFT works

ArXiv Dives is a series of live meetups that take place on Fridays with the Oxen.ai community. We believe that it is not only important to read the papers, but dive into the code t...

Greg Schoeninger
Greg Schoeninger
7/21/2024
- Arxiv Dives
10 min read
ArXiv Dives:💃 Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
ArXiv Dives:💃 Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling

Modeling sequences with infinite context length is one of the dreams of Large Language models. Some LLMs such as Transformers suffer from quadratic computational complexity, making...

Greg Schoeninger
Greg Schoeninger
6/26/2024
- Arxiv Dives
5 min read
ArXiv Dives: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
ArXiv Dives: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

The ability to interpret and steer large language models is an important topic as they become more and more a part of our daily lives. As the leader in AI safety, Anthropic takes o...

Greg Schoeninger
Greg Schoeninger
6/4/2024
- Arxiv Dives
9 min read
ArXiv Dives: Efficient DiT Fine-Tuning with PixART for Text to Image Generation
ArXiv Dives: Efficient DiT Fine-Tuning with PixART for Text to Image Generation

Diffusion Transformers have been gaining a lot of steam since OpenAI's demo of Sora back in March. The problem, when we think of training text-to-image models, we usually think mil...

Greg Schoeninger
Greg Schoeninger
5/29/2024
- Arxiv Dives
9 min read
ArXiv Dives: Evaluating LLMs for Code Completion with HumanEval
ArXiv Dives: Evaluating LLMs for Code Completion with HumanEval

Large Language Models have shown very good ability to generalize within a distribution, and frontier models have shown incredible flexibility under prompting. Now that there is so...

Greg Schoeninger
Greg Schoeninger
5/17/2024
- Arxiv Dives
15 min read
How to Train Diffusion for Text from  Scratch
How to Train Diffusion for Text from Scratch

This is part two of a series on Diffusion for Text with Score Entropy Discrete Diffusion (SEDD) models. Today we will be diving into the code for diffusion models for text, and see...

Greg Schoeninger
Greg Schoeninger
4/30/2024
- Arxiv Dives
16 min read
ArXiv Dives: Text Diffusion with SEDD
ArXiv Dives: Text Diffusion with SEDD

Diffusion models have been popular for computer vision tasks. Recently models such as Sora show how you can apply Diffusion + Transformers to generate state of the art videos with ...

Greg Schoeninger
Greg Schoeninger
4/16/2024
- Arxiv Dives
11 min read
ArXiv Dives: The Era of 1-bit LLMs, All Large Language Models are in 1.58 Bits
ArXiv Dives: The Era of 1-bit LLMs, All Large Language Models are in 1.58 Bits

This paper presents BitNet b1.58 where every weight in a Transformer can be represented as a {-1, 0, 1} instead of a floating point number. The model matches full precision transfo...

Greg Schoeninger
Greg Schoeninger
4/8/2024
- Arxiv Dives
9 min read
ArXiv Dives: Evolutionary Optimization of Model Merging Recipes
ArXiv Dives: Evolutionary Optimization of Model Merging Recipes

Today, we’re diving into a fun paper by the team at Sakana.ai called “Evolutionary Optimization of Model Merging Recipes”. The high level idea is that we have so many open weights ...

Greg Schoeninger
Greg Schoeninger
4/1/2024
- Arxiv Dives
10 min read
ArXiv Dives: I-JEPA
ArXiv Dives: I-JEPA

Today, we’re diving into the I-JEPA paper. JEPA stands for Joint-Embedding Predictive Architecture and if you have been following Yann LeCunn, is a technique he has been hyping up ...

Greg Schoeninger
Greg Schoeninger
3/26/2024
- Arxiv Dives
13 min read
How to train Mistral 7B as a "Self-Rewarding Language Model"
How to train Mistral 7B as a "Self-Rewarding Language Model"

About a month ago we went over the "Self-Rewarding Language Models" paper by the team at Meta AI with the Oxen.ai Community. The paper felt very approachable and reproducible, so w...

Greg Schoeninger
Greg Schoeninger
3/20/2024
- Practical ML
17 min read
Downloading Datasets with Oxen.ai
Downloading Datasets with Oxen.ai

Oxen.ai makes it quick and easy to download any version of your data wherever and whenever you need it. When we say quick, we mean raw speed. Oxen chunks and transfers data faster...

Greg Schoeninger
Greg Schoeninger
3/18/2024
- Getting Started
4 min read
Uploading Datasets to Oxen.ai
Uploading Datasets to Oxen.ai

Oxen.ai makes it quick and easy to upload your datasets, keep track of every version and share them with your team or the world. Oxen datasets can be as small as a single csv or as...

Greg Schoeninger
Greg Schoeninger
3/18/2024
- Getting Started
4 min read
ArXiv Dives - Diffusion Transformers
ArXiv Dives - Diffusion Transformers

Diffusion transformers achieve state-of-the-art quality generating images by replacing the commonly used U-Net backbone with a transformer that operates on latent patches. They rec...

Greg Schoeninger
Greg Schoeninger
3/12/2024
- Arxiv Dives
14 min read
"Road to Sora" Paper Reading List
"Road to Sora" Paper Reading List

This post is an effort to put together a reading list for our Friday paper club called ArXiv Dives. Since there has not been an official paper released yet for Sora, the goal is fo...

Greg Schoeninger
Greg Schoeninger
3/5/2024
- Arxiv Dives
21 min read
ArXiv Dives - Medusa
ArXiv Dives - Medusa

Abstract In this paper, they present MEDUSA, an efficient method that augments LLM inference by adding extra decoding heads to predict multiple subsequent tokens in parallel. The ...

Greg Schoeninger
Greg Schoeninger
3/4/2024
- Arxiv Dives
5 min read
ArXiv Dives - Lumiere
ArXiv Dives - Lumiere

This paper introduces Lumiere – a text-to-video diffusion model designed for synthesizing videos that portray realistic, diverse and coherent motion – a pivotal challenge in video ...

Greg Schoeninger
Greg Schoeninger
2/27/2024
- Arxiv Dives
11 min read
ArXiv Dives - Depth Anything
ArXiv Dives - Depth Anything

This paper presents Depth Anything, a highly practical solution for robust monocular depth estimation. Depth estimation traditionally requires extra hardware and algorithms such as...

Greg Schoeninger
Greg Schoeninger
2/19/2024
- Arxiv Dives
16 min read
Arxiv Dives - Toolformer: Language models can teach themselves to use tools
Arxiv Dives - Toolformer: Language models can teach themselves to use tools

Large Language Models (LLMs) show remarkable capabilities to solve new tasks from a few textual instructions, but they also paradoxically struggle with basic functionality such as ...

Greg Schoeninger
Greg Schoeninger
2/12/2024
- Arxiv Dives
10 min read
Arxiv Dives - Self-Rewarding Language Models
Arxiv Dives - Self-Rewarding Language Models

The goal of this paper is to see if we can create a self-improving feedback loop to achieve “superhuman agents”. Current language models are bottlenecked by labeled data from human...

Greg Schoeninger
Greg Schoeninger
2/6/2024
- Arxiv Dives
13 min read
Arxiv Dives - Direct Preference Optimization (DPO)
Arxiv Dives - Direct Preference Optimization (DPO)

This paper provides a simple and stable alternative to RLHF for aligning Large Language Models with human preferences called "Direct Preference Optimization" (DPO). They reformulat...

Greg Schoeninger
Greg Schoeninger
1/30/2024
- Arxiv Dives
12 min read
Arxiv Dives - Efficient Streaming Language Models with Attention Sinks
Arxiv Dives - Efficient Streaming Language Models with Attention Sinks

This paper introduces the concept of an Attention Sink which helps Large Language Models (LLMs) maintain the coherence of text into the millions of tokens while also maintaining a ...

Greg Schoeninger
Greg Schoeninger
1/20/2024
- Arxiv Dives
12 min read
Arxiv Dives - How Mixture of Experts works with Mixtral 8x7B
Arxiv Dives - How Mixture of Experts works with Mixtral 8x7B

Mixtral 8x7B is an open source mixture of experts large language model released by the team at Mistral.ai that outperforms Llama-2 70B and GPT-3.5 on a variety natural language und...

Greg Schoeninger
Greg Schoeninger
1/13/2024
- Arxiv Dives
12 min read
Arxiv Dives - LLaVA 🌋 an open source Large Multimodal Model (LMM)
Arxiv Dives - LLaVA 🌋 an open source Large Multimodal Model (LMM)

What is LLaVA? LLaVA is a Multi-Modal model that connects a Vision Encoder and an LLM for general purpose visual and language understanding. Paper: https://arxiv.org/abs/2304.084...

Greg Schoeninger
Greg Schoeninger
1/7/2024
- Arxiv Dives
12 min read
Practical ML Dive - Building RAG from Open Source Pt 1
Practical ML Dive - Building RAG from Open Source Pt 1

RAG was introduced by the Facebook AI Research (FAIR) team in May of 2020 as an end-to-end way to include document search into a sequence-to-sequence neural network architecture. ...

Greg Schoeninger
Greg Schoeninger
1/6/2024
- Practical ML
14 min read
Arxiv Dives - How Mistral 7B works
Arxiv Dives - How Mistral 7B works

What is Mistral 7B? Mistral 7B is an open weights large language model by Mistral.ai that was build for performance and efficiency. It outshines models that are twice it's size, i...

Greg Schoeninger
Greg Schoeninger
12/23/2023
- Arxiv Dives
10 min read
Practical ML Dive - How to train Mamba for Question Answering
Practical ML Dive - How to train Mamba for Question Answering

What is Mamba 🐍? There is a lot of hype about Mamba being a fast alternative to the Transformer architecture. The paper released in December of 2023 claims 5x faster throughput w...

Greg Schoeninger
Greg Schoeninger
12/21/2023
- Practical ML
22 min read
Mamba: Linear-Time Sequence Modeling with Selective State Spaces - Arxiv Dives
Mamba: Linear-Time Sequence Modeling with Selective State Spaces - Arxiv Dives

What is Mamba 🐍? Mamba at it's core is a recurrent neural network architecture, that outperforms Transformers with faster inference and improved handling of long sequences of len...

Greg Schoeninger
Greg Schoeninger
12/15/2023
- Arxiv Dives
15 min read
Practical ML Dive - How to customize a Vision Transformer on your own data
Practical ML Dive - How to customize a Vision Transformer on your own data

Welcome to Practical ML Dives, a series spin off of Arxiv Dives. In Arxiv Dives, we cover state of the art research papers, and dive into the gnitty gritty details of how AI model...

Greg Schoeninger
Greg Schoeninger
12/14/2023
- Arxiv Dives
20 min read
Arxiv Dives - Zero-shot Image Classification with CLIP
Arxiv Dives - Zero-shot Image Classification with CLIP

CLIP explores the efficacy of learning image representations from scratch with 400 million image-text pairs, showcasing zero-shot transfer capabilities across diverse computer visi...

Greg Schoeninger
Greg Schoeninger
12/8/2023
- Arxiv Dives
14 min read
How NOT to store unstructured machine learning datasets
How NOT to store unstructured machine learning datasets

Training data is typically the most valuable part of any machine learning project. As we converge on model architectures like the transformer that perform well on many tasks, it is...

Greg Schoeninger
Greg Schoeninger
12/8/2023
6 min read
đŸ§Œ SUDS - A Guide to Structuring Unstructured Data
đŸ§Œ SUDS - A Guide to Structuring Unstructured Data

At Oxen.ai we value high quality datasets. We have many years of experience training and evaluating models, and have seen many interesting data formats. Interesting is something we...

Greg Schoeninger
Greg Schoeninger
12/8/2023
12 min read
Arxiv Dives - Vision Transformers (ViT)
Arxiv Dives - Vision Transformers (ViT)

With all of the hype around Transformers for natural language processing and text, the authors of this paper beg the question - can we apply self-attention and Transformers to imag...

Greg Schoeninger
Greg Schoeninger
12/1/2023
- Arxiv Dives
13 min read
Reading List For Andrej Karpathy’s “Intro to Large Language Models” Video
Reading List For Andrej Karpathy’s “Intro to Large Language Models” Video

Andrej Karpathy recently released an hour long talk on “The busy person’s intro to large language models” that had some great tidbits whether you are an expert in machine learning ...

Greg Schoeninger
Greg Schoeninger
11/27/2023
11 min read
Arxiv Dives - A Mathematical Framework for Transformer Circuits - Part 2
Arxiv Dives - A Mathematical Framework for Transformer Circuits - Part 2

Every Friday at Oxen.ai we host a paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. We believe diving into the details of research papers is the best way to build fund...

Greg Schoeninger
Greg Schoeninger
11/21/2023
- Arxiv Dives
16 min read
Arxiv Dives - A Mathematical Framework for Transformer Circuits - Part 1
Arxiv Dives - A Mathematical Framework for Transformer Circuits - Part 1

Every Friday at Oxen.ai we host a paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. We believe diving into the details of research papers is the best way to build fund...

Greg Schoeninger
Greg Schoeninger
11/11/2023
- Arxiv Dives
13 min read
Data Version Control 101 with Oxen
Data Version Control 101 with Oxen

This intro tutorial from Oxen.ai shows how Oxen can make versioning your data as easy as versioning your code. Oxen is built to track and store changes for everything from a singl...

Greg Schoeninger
Greg Schoeninger
11/9/2023
12 min read
Arxiv Dive Manifesto
Arxiv Dive Manifesto

Every Friday the team at Oxen.ai gets together and goes over research papers, blog posts, or books that help us stay up to date with the latest in Machine Learning and AI. We call ...

Greg Schoeninger
Greg Schoeninger
11/5/2023
- Arxiv Dives
4 min read
Arxiv Dives - Attention Is All You Need
Arxiv Dives - Attention Is All You Need

Every Friday at Oxen.ai we host a paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. We believe diving into the details of research papers is the best way to build fund...

Greg Schoeninger
Greg Schoeninger
11/4/2023
- Arxiv Dives
17 min read
Arxiv Dives - How LoRA fine-tuning works
Arxiv Dives - How LoRA fine-tuning works

Every Friday at Oxen.ai we host a paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. We believe diving into the details of research papers is the best way to build fund...

Greg Schoeninger
Greg Schoeninger
10/27/2023
- Arxiv Dives
10 min read
How to run Llama-2 on CPU after fine-tuning with LoRA
How to run Llama-2 on CPU after fine-tuning with LoRA

Running Large Language Models (LLMs) on the edge is a fascinating area of research, and opens up many use cases that require data privacy or lower cost profiles. With libraries lik...

Greg Schoeninger
Greg Schoeninger
10/23/2023
10 min read
Arxiv Dives - Generating Speech from Text with Fast Speech-2
Arxiv Dives - Generating Speech from Text with Fast Speech-2

Every Friday at Oxen.ai we host a public paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. We believe diving into the details of research papers is the best way to bui...

Greg Schoeninger
Greg Schoeninger
10/21/2023
- Arxiv Dives
11 min read
Arxiv Dives - Llama-2 Explained
Arxiv Dives - Llama-2 Explained

Every Friday at Oxen.ai we host a public paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. These are the notes from the group session for reference. If you would like ...

Greg Schoeninger
Greg Schoeninger
10/13/2023
- Arxiv Dives
10 min read
Arxiv Dives - Stable Diffusion
Arxiv Dives - Stable Diffusion

Every Friday at Oxen.ai we host a public paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. These are the notes from the group session for reference. If you would like ...

Greg Schoeninger
Greg Schoeninger
10/9/2023
- Arxiv Dives
8 min read
Arxiv Dives - How Segment Anything Works
Arxiv Dives - How Segment Anything Works

Every Friday at Oxen.ai we host a public paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. These are the notes from the group session for reference. If you would like ...

Greg Schoeninger
Greg Schoeninger
9/29/2023
- Arxiv Dives
10 min read
Arxiv Dives - Retrieval Augmented Generation (RAG)
Arxiv Dives - Retrieval Augmented Generation (RAG)

Every Friday at Oxen.ai we host a public paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. These are the notes from the group session for reference. If you would like ...

Greg Schoeninger
Greg Schoeninger
9/25/2023
- Arxiv Dives
9 min read
Arxiv Dives - Training Language Models to Follow Instructions (InstructGPT)
Arxiv Dives - Training Language Models to Follow Instructions (InstructGPT)

Join the "Nerd Herd" Every Friday at Oxen.ai we host a public paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. These are the notes from the group session for referen...

Greg Schoeninger
Greg Schoeninger
9/15/2023
- Arxiv Dives
8 min read
Arxiv Dives - Language Models are Unsupervised Multitask Learners (GPT-2)
Arxiv Dives - Language Models are Unsupervised Multitask Learners (GPT-2)

Every Friday at Oxen.ai we host a public paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. These are the notes from the group session for reference. If you would like ...

Greg Schoeninger
Greg Schoeninger
9/8/2023
- Arxiv Dives
8 min read
Creating a Cute Custom Character with Stable Diffusion and Dreambooth
Creating a Cute Custom Character with Stable Diffusion and Dreambooth

Introduction Stable Diffusion is an incredible open-source tool for fast, effective generation of novel images across a wide variety of domains. Despite its power and convenience...

Greg Schoeninger
Greg Schoeninger
8/1/2023
8 min read