# Oxen.ai > Oxen.ai is the platform for building AI on your own data. Use 200+ models through one API, version your datasets at scale, and fine-tune custom models. Oxen.ai gives AI teams one platform to run 200+ models through a single API, version datasets of any size, fine-tune custom models on their own data, and deploy them — with an open-source CLI and Python library for fast data version control. ## Docs Every docs page is also available as Markdown by appending .md to its URL. - [Docs index for LLMs](https://docs.oxen.ai/llms.txt): Machine-readable index of every documentation page - [Full docs as one file](https://docs.oxen.ai/llms-full.txt): The entire documentation corpus in a single Markdown file - [Getting started](https://docs.oxen.ai/getting-started/intro.md): What Oxen.ai is and how the pieces fit together - [Install the CLI](https://docs.oxen.ai/getting-started/install.md): Install the oxen CLI and push your first dataset - [Inference](https://docs.oxen.ai/getting-started/inference.md): Use 200+ models through one API - [Fine-tuning](https://docs.oxen.ai/getting-started/fine-tuning.md): Fine-tune custom models on your own data - [Data version control](https://docs.oxen.ai/examples/data/versioning.md): Version, branch, and share terabyte-scale datasets - [Python API](https://docs.oxen.ai/python-api/index.md): The oxenai Python library reference - [HTTP API](https://docs.oxen.ai/http-api/index.md): REST API reference - [Self-hosted oxen-server](https://docs.oxen.ai/getting-started/oxen-server.md): Run the open-source server yourself ## Product - [Models](https://www.oxen.ai/ai/models): Browse 200+ hosted models with pay-as-you-go pricing - [Explore datasets](https://www.oxen.ai/explore): Public dataset repositories hosted on Oxen.ai - [Pricing](https://www.oxen.ai/pricing): Plans and compute pricing - [Community](https://www.oxen.ai/community): ArXiv Dives paper club and community events ## Blog - [Run DeepSeek, Kimi, and GLM in Your Terminal with oh-my-pi and Oxen.ai](https://www.oxen.ai/blog/run-deepseek-kimi-and-glm-in-your-terminal-with-oh-my-pi-and-oxen-ai): It seems like every day on X there is a post about a new model overtaking Claude or GPT. The belle of the ball this week was DeepSeek V4.1 Flash which is ~97% … - [WAN 3.0 Prompting Guide](https://www.oxen.ai/blog/wan-30-prompting-guide): The next series of video generation model from Alibaba WAN is powerful, but can take a minute to master. It's not the best model for complex scenes or motions,… - [Speeding Up the Oxen.ai File System](https://www.oxen.ai/blog/speeding-up-our-merkle-tree-with-lmdb): We started Oxen.ai not just to make the latest AI models easy to use, but to help people customize and train their own. A core piece of that is data storage. O… - [MiniMax H3: Open Weights & Prompting Techniques Complete Guide.](https://www.oxen.ai/blog/minimax-h3): Open weights video generation has finally caught up. Last week, MiniMax released their H3 general purpose video generation model - and it is good. It is not on… - [Seedance 2.5, MiniMax H3, FLUX 3, and WAN 3 are now live on Oxen!](https://www.oxen.ai/blog/seedance-2-5-minimax-h3-flux-3-and-wan-3-are-now-live-on-oxen): Four brand new video releases landed over the past week. Everybody chose to show off and we're very happy they did. Also, amazing news to report, open source v… - [Oxen's Model Report - July 23rd, 2026](https://www.oxen.ai/blog/oxens-model-report-july-23rd-2026): Every time I decide to write one of these I come up with an initial shortlist. By the time I actually sit down to write, a couple of days have usually passed a… - [The Ultimate Guide to Topaz Upscaling Video Models: Proteus, Starlight, Wonder, and Hyperion](https://www.oxen.ai/blog/the-ultimate-guide-to-topaz-upscaling-video-models-proteus-starlight-wonder-and-hyperion): The hardest part of AI video isn't pressing generate. It's finishing. How do you get high-quality 4K HDR content into your post-production pipeline? That's whe… - [Everything new in Oxen.ai - July 7th, 2026](https://www.oxen.ai/blog/everything-new-in-oxen-ai-july-7th-2026): Hey Herd, It's been awhile since we've done a proper feature update, so wanted to come up for air to show you the latest and greatest. The Oxen team has been h… - [Oxen.ai vs Higgsfield: pay-as-you-go AI video and image models](https://www.oxen.ai/blog/oxen-ai-vs-higgsfield-pay-as-you-go-ai-video-and-image-models): Higgsfield is great for solo creators who want a polished, creative-first studio. Its curated cinematic presets and effects get you to a good-looking result qu… - [Oxen.ai vs Magnific (FreePik): which AI creative platform should you use?](https://www.oxen.ai/blog/oxen-ai-vs-magnific-freepik-which-ai-creative-platform-should-you-use): Magnific and Oxen.ai both put the top AI image and video models in one place, alongside upscaling and editing tools. If you are choosing between them, the diff… - [Oxen's Model Report - May 8th, 2026](https://www.oxen.ai/blog/oxens-model-report-may-8th-2026): Welcome back to another iteration of everybody's favorite moooodel report. Every time I sit down to write one of these I'm shocked by how much there is to cove… - [Writing a fine-tuning and deployment pipeline isn't as easy as it looks (Gemma 4 Version)](https://www.oxen.ai/blog/writing-a-fine-tuning-and-deployment-pipeline-isnt-as-easy-as-it-looks-gemma-4-version): 🚨You're about to embark on a journey of ups and downs and many aha moments. We're pulling back the curtain on every painful hour so you never have to spend th… - [Oxen's Model Report - April 9th, 2026](https://www.oxen.ai/blog/oxens-model-report-3): Welcome back to another iteration of everybody’s favorite moooodel report. In AI, any given day feels like a decade, and the past couple of weeks have felt lik… - [When to Fine-Tune an Image Model](https://www.oxen.ai/blog/when-to-fine-tune-an-image-model): 💡Curious about fine-tuning multi-modal models? This Friday, we're diving into the new Qwen3.5 series, what makes it great and how to train it on images and vi… - [Oxen's Model Report - March 11th, 2026](https://www.oxen.ai/blog/oxens-model-report-2): Welcome back to another iteration of our favorite moooodel report. This week we've got an absolutely packed lineup, from motion-controlled video generation to … - [Frank's Red Hot](https://www.oxen.ai/blog/franks-red-hot-super-bowl): 0:00 /0:45 1× An AI generated goat rapping alongside Ludacris in Frank's RedHot's "Eat The GOAT" Super Bowl ad. The goat was fully generated using Oxen AI, bri… - [Isometric.nyc](https://www.oxen.ai/blog/isometric-nyc): A giant isometric pixel-art map of New York City, inspired by SimCity 2000 and Rollercoaster Tycoon. Andy Coenen fine-tuned an image model on Oxen with just 40… - [Bell Canada](https://www.oxen.ai/blog/bell-canada-x-oxen-ai): 0:00 /0:30 1× A fully AI generated commercial for the Canadian telecom giant Bell. Made in collaboration with KnuckleheadTV and amazing artists such as Bruce A… - [Oxen's Model Report](https://www.oxen.ai/blog/oxens-model-report): Welcome to this week's Oxen moooodel report. We know the AI space moves like crazy. There's a new model, paper, podcast, or hot take every single day. To help … - [How a $1 Qwen3-VL Fine-Tune Beat Gemini 3](https://www.oxen.ai/blog/how-a-1-qwen3-vl-fine-tune-beat-gemini-3): Can a $1 fine-tune beat a state-of-the-art closed-source model? ModelAccuracyTime (98 samples)Cost/RunBase Qwen3-VL-8B54.1%~10 sec$0.003Gemini 3 Flash82.7%2 mi… - [How to Train a LTX-2 Character LoRA with Oxen.ai](https://www.oxen.ai/blog/how-to-train-a-ltx-2-character-lora-with-oxen-ai): LTX-2 is a video generation model, that not only can generation video frames, but audio as well. This model is fully open source, meaning the weights and the c… - [How to Use WAN 2.1-VACE to Generate Hollywood-Level Video Edits](https://www.oxen.ai/blog/how-to-use-wan-2-1-vace-to-generate-hollywood-level-video-edits): Imagine you are shooting a film and you realize that you have the actor wearing the wrong jacket in a scene. Do you bring the whole cast back in to re-shoot? D… - [How We Cut Inference Costs from $46K to $6.5K Fine-Tuning Qwen-Image-Edit](https://www.oxen.ai/blog/how-we-cut-inference-costs-from-46k-to-7-5k-fine-tuning-qwen-image-edit): At Oxen.ai, we think a lot about what it takes to run high-quality inference at scale. It’s one thing to produce a handful of impressive results with a cutting… - [How to Set Noise Timesteps When Fine-Tuning Diffusion Models for Image Generation](https://www.oxen.ai/blog/how-to-set-noise-timesteps-when-fine-tuning-diffusion-models-for-image-generation): Fine-tuning Diffusion Models such as Stable Diffusion, FLUX.1-dev, or Qwen-Image can give you a lot of bang for your buck. Base models may not be trained on a … - [Fine-Tuned Qwen-Image-Edit vs Nano-Banana and FLUX Kontext Dev](https://www.oxen.ai/blog/fine-tuned-qwen-image-edit-vs-nano-banana-and-flux-kontext-dev): Welcome back to Fine-Tuning Friday, where each week we try to put some models to the test and see if fine-tuning an open-source model can outperform whatever s… - [We Fine-Tuned GPT OSS 20B to Rap Like Eminem](https://www.oxen.ai/blog/we-fine-tuned-gpt-oss-20b-to-rap-like-eminem): OpenAI came out with GPT-OSS 120B and 20B in August 2025. The first “Open” LLMs from OpenAI since GPT-2, over six years ago. The idea of fine-tuning a frontier… - [How We're Building a “Tab Tab” Code Completion Model](https://www.oxen.ai/blog/building-a-tab-tab-code-completion-model): Welcome to Fine-Tuning Fridays, where we share our learnings from fine-tuning open source models for real world tasks. We’ll walk you through what models work,… - [How to Fine-Tune a FLUX.1-dev LoRA with Code, Step by Step](https://www.oxen.ai/blog/how-to-fine-tune-a-flux-1-dev-lora-with-code-step-by-step): FLUX.1-dev is one of the most popular open-weight models available today. Developed by Black Forest Labs, it has 12 billion parameters. The goal of this post i… - [How to Fine-Tune PixArt to Generate a Consistent Character](https://www.oxen.ai/blog/fine-tuning-a-diffusion-transformer-to-generate-a-consistent-character): Can we fine-tune a small diffusion transformer (DiT) to generate OpenAI-level images by distilling off of OpenAI images? The end goal is to have a small, fast,… - [How to Fine-Tune Qwen3 on Text2SQL to GPT-4o level performance](https://www.oxen.ai/blog/how-to-fine-tune-qwen3-to-gpt-4o-level-performance): Welcome to a new series from the Oxen.ai Herd called Fine-Tuning Fridays! Each week we will take an open source model and put it head to head against a closed … - [Fine-Tuning Fridays](https://www.oxen.ai/blog/fine-tuning-fridays): Fine-tuning Fridays is a series from the Oxen.ai Herd where each week we do a deep dive into a model and put it head to head against the competition. We will b… - [How RWKV-7 Goose Works 🪿 + Notes from the Author](https://www.oxen.ai/blog/how-rwkv-7-goose-works-notes-from-the-author): In this special Arxiv Dive, we're joined by Eugene Cheah - author, lead in RWKV org, CEO of Featherless AI, to discuss the development process and key decision… - [How Phi-4 Cracked Small Multimodality](https://www.oxen.ai/blog/how-phi-4-cracked-small-multimodality): Phi-4 extends the existing Phi model’s capabilities by adding vision and audio all in the same model. This means you can do everything from understand images, … - [Training a Rust 1.5B Coder LM with Reinforcement Learning (GRPO)](https://www.oxen.ai/blog/training-a-rust-1-5b-coder-lm-with-reinforcement-learning-grpo): Group Relative Policy Optimization (GRPO) has proven to be a useful algorithm for training LLMs to reason and improve on benchmarks. DeepSeek-R1 showed that yo… - [Why GRPO is Important and How it Works](https://www.oxen.ai/blog/why-grpo-is-important-and-how-it-works): Last week on Arxiv Dives we dug into research behind DeepSeek-R1, and uncovered that one of the techniques they use in the their training pipeline is called Gr… - [🧠 GRPO VRAM Requirements For the GPU Poor](https://www.oxen.ai/blog/grpo-vram-requirements-for-the-gpu-poor): Since the release of DeepSeek-R1, Group Relative Policy Optimization (GRPO) has become the talk of the town for Reinforcement Learning in Large Language Models… - [How DeepSeek R1, GRPO, and Previous DeepSeek Models Work](https://www.oxen.ai/blog/how-deepseek-r1-grpo-and-previous-deepseek-models-work): In January 2025, DeepSeek took a shot directly at OpenAI by releasing a suite of models that “Rival OpenAI’s o1.” From their website: In the spirit of Arxiv Di… - [No Hype DeepSeek-R1 Reading List](https://www.oxen.ai/blog/no-hype-deepseek-r1-reading-list): DeepSeek-R1 is a big step forward in the open model ecosystem for AI with their latest model competing with OpenAI's o1 on a variety of metrics. There is a lot… - [Oxen v0.25.0 Migration](https://www.oxen.ai/blog/oxen-v0-25-0-migration): Today we released oxen v0.25.0 🎉 which comes with a few performance optimizations, including how we traverse the Merkle Tree to find files and folders. The ma… - [🌲 Merkle Tree VNodes](https://www.oxen.ai/blog/merkle-tree-vnodes): In this post we peel back some of the layers of Oxen.ai’s Merkle Tree and show how we make it suitable for projects with large directories. If you are unfamili… - [🌲 Merkle Tree 101](https://www.oxen.ai/blog/merkle-tree-101): Intro Merkle Trees are important data structures for ensuring integrity, deduplication, and verification of data at scale. They are used heavily in tools such … - [arXiv Dive: RAGAS - Retrieval Augmented Generation Assessment](https://www.oxen.ai/blog/arxiv-dive-ragas): RAGAS is an evaluation framework for Retrieval Augmented Generation (RAG). A paper released by Exploding Gradients, AMPLYFI, and CardiffNLP. RAGAS gives us a s… - [The Best AI Data Version Control Tools [2025]](https://www.oxen.ai/blog/the-best-ai-data-version-control-tools): Data is often seen as static. It's common to just dump your data into S3 buckets in tarballs or upload to Hugging Face and leave it at that. Yet nowadays, data… - [OpenCoder: The OPEN Cookbook For Top-Tier Code LLMs](https://www.oxen.ai/blog/opencoder-the-open-cookbook-for-top-tier-code-llms): Welcome to the last arXiv Dive of 2024! Every other week we have been diving into interesting research papers in AI/ML. In this blog we’ll be diving into Open … - [LLaVA-CoT: Let Vision Language Models Reason Step-By-Step](https://www.oxen.ai/blog/llava-cot-let-vision-language-models-reason-step-by-step-2): When it comes to large language models, it is still the early innings. Many of them still hallucinate, fail to follow instructions, or generally don’t work. Th… - [How Upcycling MoEs Beat Dense LLMs](https://www.oxen.ai/blog/how-upcycling-moes-beat-dense-llms): In this Arxiv Dive, Nvidia researcher, Ethan He, presents his co-authored work Upcycling LLMs in Mixture of Experts (MoE). He goes into what a MoE is, the chal… - [Thinking LLMs: General Instruction Following with Thought Generation](https://www.oxen.ai/blog/thinking-llms-general-instruction-following-with-thought-generation): The release of OpenAI-O1 has motivated a lot of people to think deeply about…thoughts 💭. Thinking before you speak is a skill that some people have better tha… - [The Prompt Report Part 2: Plan and Solve, Tree of Thought, and Decomposition Prompting](https://www.oxen.ai/blog/the-prompt-report-part-2-thought-generation-tree-of-thought-and-decomposition-prompting): In the last blog, we went over prompting techniques 1-3 of The Prompt Report. This arXiv Dive, we were lucky to have the authors of the paper join us to go thr… - [The Prompt Report Part 1: A Systematic Survey of Prompting Techniques](https://www.oxen.ai/blog/the-prompt-report-part-1-a-systematic-survey-of-prompting-techniques): For this blog we are switching it up a bit. In past Arxiv Dives, we have gone deep into the underlying model architectures and techniques that make large langu… - [arXiv Dive: How Flux and Rectified Flow Transformers Work](https://www.oxen.ai/blog/arxiv-dive-how-flux-and-rectified-flow-transformers-work): Flux made quite a splash with its release on August 1st, 2024 as the new state of the art generative image model outperforming SDXL, SDXL-Turbo, Pixart, and DA… - [How Well Can Llama 3.1 8B Detect Political Spam? [4/4]](https://www.oxen.ai/blog/how-well-can-llama-3-1-8b-detect-political-spam-4-4): It only took about 11 minutes to fine-tuned Llama 3.1 8B on our political spam synthetic dataset using ReFT. While this is extremely fast, beating out our prev… - [Fine-Tuning Llama 3.1 8B in Under 12 Minutes [3/4]](https://www.oxen.ai/blog/fine-tuning-llama-3-1-8b-in-under-12-minutes): Meta has recently released Llama 3.1, including their 405 billion parameter model which is the most capable open model to date and the first open model on the … - [arXiv Dive: How Meta Trained Llama 3.1](https://www.oxen.ai/blog/llama-3-1-herd-of-models): Llama 3.1 is a set of Open Weights Foundation models released by Meta, which marks the first time an open model has caught up to GPT-4, Anthropic, or other clo… - [How to De-duplicate and Clean Synthetic Data [2/4]](https://www.oxen.ai/blog/filtering-synthetic-data-for-quality): Synthetic data has shown promising results for training and fine tuning large models, such as Llama 3.1 and the models behind Apple Intelligence, and to produc… - [Create Your Own Synthetic Data With Only 5 Political Spam Texts [1/4]](https://www.oxen.ai/blog/create-your-own-synthetic-data-with-only-5-political-spam-texts): With the 2024 elections coming up, spam and political texts are more prevalent than ever as political campaigns increasingly turn towards texting potential vot… - [Fine-tuning Llama 3 in 14 minutes using ReFT](https://www.oxen.ai/blog/fine-tuning-llama-3-in-14-minutes-using-reft): If you have been fine-tuning models recently, you have most likely used LoRA. While LoRA has been the dominant PEFT technique for a long time thanks to its eff… - [ArXiv Dives: How ReFT works](https://www.oxen.ai/blog/arxiv-dives-how-reft-works): ArXiv Dives is a series of live meetups that take place on Fridays with the Oxen.ai community. We believe that it is not only important to read the papers, but… - [ArXiv Dives:💃 Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling](https://www.oxen.ai/blog/samba): Modeling sequences with infinite context length is one of the dreams of Large Language models. Some LLMs such as Transformers suffer from quadratic computation… - [ArXiv Dives: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet](https://www.oxen.ai/blog/scaling-monosemanticity-claude-3): The ability to interpret and steer large language models is an important topic as they become more and more a part of our daily lives. As the leader in AI safe… - [ArXiv Dives: Efficient DiT Fine-Tuning with PixART for Text to Image Generation](https://www.oxen.ai/blog/arxiv-dives-evaluating-llms-for-code-completion-with-humaneval): Diffusion Transformers have been gaining a lot of steam since OpenAI's demo of Sora back in March. The problem, when we think of training text-to-image models,… - [ArXiv Dives: Evaluating LLMs for Code Completion with HumanEval](https://www.oxen.ai/blog/arxiv-dives-human-eval): Large Language Models have shown very good ability to generalize within a distribution, and frontier models have shown incredible flexibility under prompting. … - [How to Train Diffusion for Text from Scratch](https://www.oxen.ai/blog/how-to-train-diffusion-for-text-from-scratch): This is part two of a series on Diffusion for Text with Score Entropy Discrete Diffusion (SEDD) models. Today we will be diving into the code for diffusion mod… - [ArXiv Dives: Text Diffusion with SEDD](https://www.oxen.ai/blog/arxiv-dives-text-diffusion-with-sedd): Diffusion models have been popular for computer vision tasks. Recently models such as Sora show how you can apply Diffusion + Transformers to generate state of… - [ArXiv Dives: The Era of 1-bit LLMs, All Large Language Models are in 1.58 Bits](https://www.oxen.ai/blog/arxiv-dives-bitnet-1-58): This paper presents BitNet b1.58 where every weight in a Transformer can be represented as a {-1, 0, 1} instead of a floating point number. The model matches f… - [ArXiv Dives: Evolutionary Optimization of Model Merging Recipes](https://www.oxen.ai/blog/arxiv-dives-evolutionary-model-merging): Today, we’re diving into a fun paper by the team at Sakana.ai called “Evolutionary Optimization of Model Merging Recipes”. The high level idea is that we have … - [ArXiv Dives: I-JEPA](https://www.oxen.ai/blog/arxiv-dives-i-jepa): Today, we’re diving into the I-JEPA paper. JEPA stands for Joint-Embedding Predictive Architecture and if you have been following Yann LeCunn, is a technique h… - [How to train Mistral 7B as a "Self-Rewarding Language Model"](https://www.oxen.ai/blog/how-to-train-mistral-7b-to-be-a-self-rewarding-language-model): About a month ago we went over the "Self-Rewarding Language Models" paper by the team at Meta AI with the Oxen.ai Community. The paper felt very approachable a… - [Downloading Datasets with Oxen.ai](https://www.oxen.ai/blog/download-your-first-dataset-with-oxen-ai): Oxen.ai makes it quick and easy to download any version of your data wherever and whenever you need it. When we say quick, we mean raw speed. Oxen chunks and t… - [Uploading Datasets to Oxen.ai](https://www.oxen.ai/blog/how-to-version-data-and-see-changes-oxenai): Oxen.ai makes it quick and easy to upload your datasets, keep track of every version and share them with your team or the world. Oxen datasets can be as small … - [ArXiv Dives - Diffusion Transformers](https://www.oxen.ai/blog/arxiv-dives-diffusion-transformers): Diffusion transformers achieve state-of-the-art quality generating images by replacing the commonly used U-Net backbone with a transformer that operates on lat… - ["Road to Sora" Paper Reading List](https://www.oxen.ai/blog/road-to-sora-reading-list): This post is an effort to put together a reading list for our Friday paper club called ArXiv Dives. Since there has not been an official paper released yet for… - [ArXiv Dives - Medusa](https://www.oxen.ai/blog/arxiv-dives-medusa): Abstract In this paper, they present MEDUSA, an efficient method that augments LLM inference by adding extra decoding heads to predict multiple subsequent toke… - [ArXiv Dives - Lumiere](https://www.oxen.ai/blog/arxiv-dives-lumiere): This paper introduces Lumiere – a text-to-video diffusion model designed for synthesizing videos that portray realistic, diverse and coherent motion – a pivota… - [ArXiv Dives - Depth Anything](https://www.oxen.ai/blog/arxiv-dives-depth-anything): This paper presents Depth Anything, a highly practical solution for robust monocular depth estimation. Depth estimation traditionally requires extra hardware a… - [Arxiv Dives - Toolformer: Language models can teach themselves to use tools](https://www.oxen.ai/blog/toolformer-language-models-can-teach-themselves-to-use-tools): Large Language Models (LLMs) show remarkable capabilities to solve new tasks from a few textual instructions, but they also paradoxically struggle with basic f… - [Arxiv Dives - Self-Rewarding Language Models](https://www.oxen.ai/blog/arxiv-dives-self-rewarding-language-models): The goal of this paper is to see if we can create a self-improving feedback loop to achieve “superhuman agents”. Current language models are bottlenecked by la… - [Arxiv Dives - Direct Preference Optimization (DPO)](https://www.oxen.ai/blog/arxiv-dives-direct-preference-optimization-dpo): This paper provides a simple and stable alternative to RLHF for aligning Large Language Models with human preferences called "Direct Preference Optimization" (… - [Arxiv Dives - Efficient Streaming Language Models with Attention Sinks](https://www.oxen.ai/blog/arxiv-dives-efficient-streaming-language-models-with-attention-sinks): This paper introduces the concept of an Attention Sink which helps Large Language Models (LLMs) maintain the coherence of text into the millions of tokens whil… - [Arxiv Dives - How Mixture of Experts works with Mixtral 8x7B](https://www.oxen.ai/blog/arxiv-dives-mixture-of-experts-moe-with-mixtral-8x7b): Mixtral 8x7B is an open source mixture of experts large language model released by the team at Mistral.ai that outperforms Llama-2 70B and GPT-3.5 on a variety… - [Arxiv Dives - LLaVA 🌋 an open source Large Multimodal Model (LMM)](https://www.oxen.ai/blog/arxiv-dive-how-to-llava-works): What is LLaVA? LLaVA is a Multi-Modal model that connects a Vision Encoder and an LLM for general purpose visual and language understanding. Paper: https://arx… - [Practical ML Dive - Building RAG from Open Source Pt 1](https://www.oxen.ai/blog/practical-ml-dive-how-to-build-rag-from-open-source): RAG was introduced by the Facebook AI Research (FAIR) team in May of 2020 as an end-to-end way to include document search into a sequence-to-sequence neural ne… - [Arxiv Dives - How Mistral 7B works](https://www.oxen.ai/blog/arxiv-dive-how-to-mistral-7b-works): What is Mistral 7B? Mistral 7B is an open weights large language model by Mistral.ai that was build for performance and efficiency. It outshines models that ar… - [Practical ML Dive - How to train Mamba for Question Answering](https://www.oxen.ai/blog/practical-ml-dive-how-to-train-mamba-for-question-answering): What is Mamba 🐍? There is a lot of hype about Mamba being a fast alternative to the Transformer architecture. The paper released in December of 2023 claims 5x… - [Mamba: Linear-Time Sequence Modeling with Selective State Spaces - Arxiv Dives](https://www.oxen.ai/blog/mamba-linear-time-sequence-modeling-with-selective-state-spaces-arxiv-dives): What is Mamba 🐍? Mamba at it's core is a recurrent neural network architecture, that outperforms Transformers with faster inference and improved handling of l… - [Practical ML Dive - How to customize a Vision Transformer on your own data](https://www.oxen.ai/blog/practical-ml-dive-how-to-customize-a-vision-transformer-on-your-own-data): Welcome to Practical ML Dives, a series spin off of Arxiv Dives. In Arxiv Dives, we cover state of the art research papers, and dive into the gnitty gritty det… - [Arxiv Dives - Zero-shot Image Classification with CLIP](https://www.oxen.ai/blog/arxiv-dives-zero-shot-image-classification-with-clip): CLIP explores the efficacy of learning image representations from scratch with 400 million image-text pairs, showcasing zero-shot transfer capabilities across … - [How NOT to store unstructured machine learning datasets](https://www.oxen.ai/blog/how-not-to-store-unstructured-machine-learning-datasets): Training data is typically the most valuable part of any machine learning project. As we converge on model architectures like the transformer that perform well… - [🧼 SUDS - A Guide to Structuring Unstructured Data](https://www.oxen.ai/blog/suds-a-guide-to-structuring-unstructured-data): At Oxen.ai we value high quality datasets. We have many years of experience training and evaluating models, and have seen many interesting data formats. Intere… - [Arxiv Dives - Vision Transformers (ViT)](https://www.oxen.ai/blog/arxiv-dives-vision-transformers-vit): With all of the hype around Transformers for natural language processing and text, the authors of this paper beg the question - can we apply self-attention and… - [Reading List For Andrej Karpathy’s “Intro to Large Language Models” Video](https://www.oxen.ai/blog/reading-list-for-andrej-karpathys-intro-to-large-language-models-video): Andrej Karpathy recently released an hour long talk on “The busy person’s intro to large language models” that had some great tidbits whether you are an expert… - [Arxiv Dives - A Mathematical Framework for Transformer Circuits - Part 2](https://www.oxen.ai/blog/arxiv-dives-a-mathematical-framework-for-transformer-circuits-part-two): Every Friday at Oxen.ai we host a paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. We believe diving into the details of research papers is the b… - [Arxiv Dives - A Mathematical Framework for Transformer Circuits - Part 1](https://www.oxen.ai/blog/arxiv-dives-a-mathematical-framework-for-transformer-circuits): Every Friday at Oxen.ai we host a paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. We believe diving into the details of research papers is the b… - [Data Version Control 101 with Oxen](https://www.oxen.ai/blog/oxen-version-control-101): This intro tutorial from Oxen.ai shows how Oxen can make versioning your data as easy as versioning your code. Oxen is built to track and store changes for eve… - [Arxiv Dive Manifesto](https://www.oxen.ai/blog/arxiv-dive-manifesto): Every Friday the team at Oxen.ai gets together and goes over research papers, blog posts, or books that help us stay up to date with the latest in Machine Lear… - [Arxiv Dives - Attention Is All You Need](https://www.oxen.ai/blog/arxiv-dives-attention-is-all-you-need): Every Friday at Oxen.ai we host a paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. We believe diving into the details of research papers is the b… - [Arxiv Dives - How LoRA fine-tuning works](https://www.oxen.ai/blog/arxiv-dives-how-lora-fine-tuning-works): Every Friday at Oxen.ai we host a paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. We believe diving into the details of research papers is the b… - [How to run Llama-2 on CPU after fine-tuning with LoRA](https://www.oxen.ai/blog/how-to-run-llama-2-on-cpu-after-fine-tuning-with-lora): Running Large Language Models (LLMs) on the edge is a fascinating area of research, and opens up many use cases that require data privacy or lower cost profile… - [Arxiv Dives - Generating Speech from Text with Fast Speech-2](https://www.oxen.ai/blog/arxiv-dives-fast-speech-2): Every Friday at Oxen.ai we host a public paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. We believe diving into the details of research papers i… - [Arxiv Dives - Llama-2 Explained](https://www.oxen.ai/blog/arxiv-dives-how-llama-2-works): Every Friday at Oxen.ai we host a public paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. These are the notes from the group session for referenc… - [Arxiv Dives - Stable Diffusion](https://www.oxen.ai/blog/arxiv-dives-stable-diffusion): Every Friday at Oxen.ai we host a public paper club called "Arxiv Dives" to make us smarter Oxen 🐂 🧠. These are the notes from the group session for referenc… ## Optional - [Oxen CLI on GitHub](https://github.com/Oxen-AI/Oxen): Open-source data version control, written in Rust - [Discord](https://discord.gg/s3tBEn7Ptg): Community chat - [X / Twitter](https://twitter.com/oxen_ai): Product updates