U

Unsloth

U
Unsloth AI

Windows ARM64 Binaries

Studio: keep the GPU order the user asked for instead of re-emitting … …it ascending (#11034) * Studio: keep the GPU order the user asked for instead of re-emitting it ascending A multi-GPU CUDA launch built the child's CUDA_VISIBLE_DEVICES from gpu_indices, which every producer sorts, so a parent CUDA_VISIBLE_DEVICES=1,0 reached llama-server as 0,1. A numeric mask carries enumeration ORDER as wel…

U
Unsloth AI

Large Performance Gains + Fixes

This is a large performance and reliability + bug fix release for Unsloth Highlights 1.2-1.7x faster diffusion. AMD 20% perf boost vs ROCM via Vulkan 2x faster updating, remove SAC + AV false positives for Windows Blender MCP, detect Hermes, AMD gibberish fixed (reported to AMD) Over 250+ bug fixes, 60% smaller binaries and performance improvements Strix iGPU BIOS popup - 3x faster inference if mo…

U
Unsloth AI

Large Perf Improvements + Fixes

This is a large performance and reliability + bug fix release for Unsloth Highlights AMD uses Vulkan by default - 20% perf boost for prefill, decoding vs ROCM Windows llama-server.exe is now signed, reducing false positives for SAC AMD gibberish issues in Strix, iGPUs fixed in (upstream - reported to AMD) Over 200+ bug fixes, 50% smaller binaries and performance improvements Updated PyTorch to 2.1…

U
Unsloth AI v5.3-Flash

2x Faster Qwen3.8-Flash + GLM-5.3-Flash MTP

Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it. Also our new release includes 170+ training, chat, hardware, and performance improvements. Highlights Smoother model loading (less errors) across local servers and connected providers. Faster and less laggy UI with follow-up turns much faster for all chats. Safer chat edits that…

U
Unsloth AI v5.3-Flash

2x Faster Qwen3.8-Flash + GLM-5.3-Flash MTP

Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it. Also our new release includes 170+ training, chat, hardware, and performance improvements. Highlights Smoother model loading (less errors) across local servers and connected providers. Safer chat edits that preserve tool cards, reply details, and conversation branches. New local…

U
Unsloth AI v5.3-Flash

Qwen3.8-Flash-Next + GLM-5.3-Flash

Qwen3.8-Flash-Next and GLM-5.3-Flash can now run locally in Unsloth! Run Qwen3.8-Flash on 75GB RAM, GLM-5.3-Flash on 102GB RAM+VRAM 5x Faster inference for RAM offloading "Infinite" repeated compaction now works 100+ chat, reliability and performance improvements Qwen Guide: https://unsloth.ai/docs/models/qwen3.8-next Qwen GGUFs: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF GLM Guide: ht…

U
Unsloth AI

Bug Fixes + Auto compaction + LAN Remote Access

Thanks for the support for Qwen3.8-27B and Unsloth Desktop! This is a bug fix release with 170+ PRs. MLX fixed - Some MLX and Mac runtimes did not run correctly LAN API keyless / password-less + Keyboard shortcuts XET / HTTP download toggle - clearer download progress AMD bug fixes + 170 bug, reliability & performance fixes Features Auto Compaction (Experimental) for longer chats beyond context li…

U
Unsloth AI

Bug Fixes + Auto compaction + LAN Remote Access

Thanks for the support for Qwen3.8-27B and Unsloth Desktop! This is a bug fix release with 170+ PRs. MLX fixed - Some MLX and Mac runtimes did not run correctly LAN API keyless / password-less is now supported XET / HTTP download toggle - clearer download progress AMD bug fixes for Strix Halo, all RDNA GPUs + 170 bug fixes Features Auto Compaction (Experimental) for longer chats beyond context lim…

U
Unsloth AI

Auto compaction (preview) + LAN Remote Access

Thanks for the support for Qwen3.8-27B and Unsloth Desktop last week! For this release, we merged 200+ PRs to introduce many new features, fixes including: Auto Compaction (Experimental) for longer chats beyond context limits Remote & LAN Access (Preview) for easy network access without Cloudflare links Faster Chat - Improved streaming performance, reduced UI lag, and smoother long conversations.…

U
Unsloth AI

pre-split-full: Merge main into studio-mmproj-fit

Three conflicts, all small. The llama_cpp.py one is two independent additions at the same spot, the projector pin state and the Metal context refusal; both are kept. The mmproj-fallback test file is likewise both sides' tests, main's loadFallbackNotice coverage alongside the wording assertion here. image-input-support.ts takes main's line. The missing .ts on that runtime import is what made the fr…

U
Unsloth AI

Qwen3.8-27B

Qwen3.8-27B and Qwen3.8-2.4T can now be run locally in Unsloth! Run on 17GB RAM via Unsloth Dynamic GGUFs. You can also fine-tune Qwen3.8-27B in Unsloth. Qwen3.8-27B is by far the strongest model for its size. We also uploaded NVFP4 quants. Guide: https://unsloth.ai/docs/models/qwen3.8 GGUF: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF See 1-bit Qwen3.8-2.4T GGUF running in Unsloth: Highlights…

U
Unsloth AI v0.1.71-beta

v0.1.71-beta

Offer the media pickers only what the host can run, and name the H3 s…

U
Unsloth AI v0.1.702-beta

v0.1.702-beta

Unsloth Desktop is here! The first desktop app to run and train AI models locally. Research, export and deploy from the same open-source app on Windows, macOS and Linux. v0.1.702-beta Update (August 13th) Added tool calling / web search & more for all external providers Fixed bypass permissions not working for sandboxing UI and UX fixes - VRAM usage is now tunable 10% faster inference + reduced VR…

U
Unsloth AI

Introducing Unsloth Desktop 🦥

Unsloth Desktop is here! The first desktop app to run and train AI models locally. Research, export and deploy from the same open-source app on Windows, macOS and Linux. 🦥 Download Unsloth Desktop for Linux, Windows, MacOS Here's what you can do with Unsloth Desktop: Get up to 50% more accurate tool calling with self-healing calls and sandboxed code execution. Run Muse Glimmer 30B, Kimi K3, Qwen3.…

U
Unsloth AI

Introducing Unsloth Desktop 🦥

Unsloth Desktop is here! The first desktop app to run and train AI models locally. Research, export and deploy from the same open-source app on Windows, macOS and Linux. 🦥 Download Unsloth Desktop for Linux, Windows, MacOS Here's what you can do with Unsloth Desktop: Get up to 50% more accurate tool calling with self-healing calls and sandboxed code execution. Run Muse Glimmer 30B, Kimi K3, Qwen3.…

U
Unsloth AI

Meta Muse Glimmer

Meta has released Muse Glimmer, the first open model from Meta Superintelligence Labs. Muse Glimmer is a dense, 30B parameter model by Meta, built for local agentic and coding workflows. Released under the Apache 2.0 license you can run Muse Glimmer 30B via Unsloth and Unsloth Dynamic quants for the best performance. Run Muse Glimmer Fine-tune Muse Glimmer Muse Glimmer 30B can run locally on 20GB…

U
Unsloth AI

Meta Muse Glimmer

Meta has released Muse Glimmer, the first open model from Meta Superintelligence Labs. Muse Glimmer is a dense, 30B parameter model by Meta, built for local agentic and coding workflows. Released under the Apache 2.0 license you can run Muse Glimmer 30B via Unsloth and Unsloth Dynamic quants for the best performance. Run Muse Glimmer Fine-tune Muse Glimmer Muse Glimmer 30B can run locally on 20GB…

U
Unsloth AI

DSpark + DeepSeek-V4 Flash 0731

Hey everyone! For folks who missed the news - Kimi K3 & DeepSeek v4 Flash can now run locally with Unsloth Dynamic GGUFs! We now added more efficient and faster downloading for Colab, low memory systems and also high memory and CPU systems - we auto fallback to HTTP as well if XET is stuck August 7th Update Many bug fixes + Mac fixes and smoother installations. DeepSeek V4 Flash 0731 + DSpark You…

U
Unsloth AI v0.1.527-beta

Unsloth v0.1.527-beta

What's Changed Studio: judge the cached-pipeline signals on the snapshot the row actually loads by @danielhanchen in #7851 Bump install.sh / install.ps1 pin to unsloth>=2026.8.3 by @danielhanchen in #7860 Studio: make the sidebar menu separators visible in dark mode by @shimmyshimmer in #7865 Studio: run sandbox matplotlib headless by @NilayYadav in #7789 Don't send desktop startup requests to an…

U
Unsloth AI v0.1.525-beta

v0.1.525-beta: Improve desktop installer progress UX (#8102)

Improve desktop installer progress UX Keep timed installer copy phase-neutral Align macOS titlebar controls Refine setup loading copy Show live installer progress [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci Rotate installer phase copy Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>

U
Unsloth AI v0.1.52-beta

v0.1.52-beta: Align desktop titlebar controls by platform (#7837)

Align desktop titlebar controls by platform [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci Align sidebar navigation icons Hide collapsed sidebar in Tauri [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci Refine collapsed desktop navigation [pre-commit.ci] auto fixes from pre-commit.com hooks fo…

U
Unsloth AI

Kimi K3 + DeepSeek-V4 Flash 0731 + Deep Research + Parallel Chat

Hey everyone! Kimi K3 & DeepSeek v4 Flash can now run locally with Unsloth Dynamic GGUFs, Unsloth can keep multiple chats generating in parallel, and the new Deep Research mode plans, reads and cites sources using your local model. This release also brings better AMD and Intel GPU support, DoRA training, and many installer, MLX, export and inference fixes. DeepSeek V4 Flash 0731 (August 2 Update)…

U
Unsloth AI v0.1.511-beta

v0.1.511-beta

Compare the workflow path as posix so the guard passes on Windows (#7…