M

MLX LM (Apple ml-explore)

M
MLX LM (Apple ml-explore) AI v0.31.3

v0.31.3

Highlights Lots of bugfixes Thread local generation stream to accompany MLX v0.31.2 What's Changed Bump the patch version by @angeloskath in #1124 Fix batch dimension mismatch in BatchKVCache and BatchRotatingKVCache extend() by @razorback16 in #1141 Fix parallel tool call handling in server by @kernelpool in #1170 Fix MiniMax M2 parallel tool calling by @kernelpool in #1171 Fix missing tree_reduc…

M
MLX LM (Apple ml-explore) AI v0.31.2

v0.31.2

Highlights Caching system prompt and user messages for non-trimmable caches Batch generator refactoring What's Changed Bump the patch version by @angeloskath in #959 Presence and frequency penalties by @angeloskath in #971 Eval self.left_padding whenever it is updated in BatchRotatingKVCache by @rltakashige in #960 Late binding caused incorrect cache checkpoint by @angeloskath in #976 Move to meta…

M
MLX LM (Apple ml-explore) AI v0.31.0

v0.31.0

What's Changed Fix save/load of CacheList by @angeloskath in #886 Share model by @angeloskath in #871 Fix mixed quant predicates for MLA models by @spicyneuron in #892 Add JoyAI LLM Flash by @kernelpool in #894 perplexity: add --trust-remote-code option by @ivanfioravanti in #896 server: add usage.prompt_tokens_details.cached_tokens to json response by @percontation in #849 Fix qwen3.5 casting to…

M
MLX LM (Apple ml-explore) AI v0.30.7

v0.30.7

What's Changed Fix Kimi Linear by @kernelpool in #853 Bump version for next release by @awni in #865 Pythonic tool calling for LFM2 models by @viktike in #864 Fix DeepSeek V3.2 indexer and weight loading by @kernelpool in #866 Make validation set optional in training process by @Goekdeniz-Guelmez in #857 Mistral tool parser by @awni in #874 LongCat MLA by @kernelpool in #868 [MODEL] support qwen3.…

M
MLX LM (Apple ml-explore) AI v0.30.6

v0.30.6

What's Changed Transformers v5 by @awni in #811 Add LongCat Flash tool parser by @kernelpool in #810 Add Kimi-K2.5 by @kernelpool in #813 Bump mlx version and version by @awni in #816 Fix NemotronH config compatibility with HuggingFace format by @LuqDaMan in #820 Fix for Exception - MultiLinear.to_quantized() missing 'mode' by @inferencers in #809 Fix Kimi K2.5 tool call handling by @kernelpool in…

M
MLX LM (Apple ml-explore) AI v0.30.5

v0.30.5

What's Changed import logging as it throws no logging error in place of actual error by @Maanas-Verma in #778 server: use OpenAI compatible finish_reason by @percontation in #782 move Xielu Activation in Apertus to activations.py by @Goekdeniz-Guelmez in #772 bump transformers by @awni in #746 Update glm4_moe_lite to store KV latent in cache by @N8python in #780 Adding TeleChat3 by @Goekdeniz-Guel…

M
MLX LM (Apple ml-explore) AI v0.30.4

v0.30.4

What's Changed Add AWQ/GPTQ weight transformation utilities by @ericcurtin in #730 Add IQuest Coder V1 Loop variant by @kernelpool in #716 Fix sliding window batching by @awni in #738 Fix Batch Generation: Add extract method to ArraysCache for item retrieval by @Goekdeniz-Guelmez in #740 Make MambaCache compatible with batch generation for nemotron-h by @nikhilmitrax in #690 Add a server benchmark…

M
MLX LM (Apple ml-explore) AI v0.30.1

v0.30.1

What's Changed custom dsv32 chat template by @awni in #693 shard glm by @awni in #698 support minimax m2 by @awni in #700 Enhance load_config function to check for config file existence and i… by @cubist38 in #701 batch_generate fails with Phi3 (LongRoPE) when prompts have different lengths by @vyaivanove in #707 Fix GIL starvation in _generate thread when batch is idle by @sjug in #706 Ignore gen…