ARahim3

mlx-dspark

ARahim3

Up to 3× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3, Ornith-1.0, ternary Bonsai-27B.

AI 简介

mlx-dspark 是一个在 Apple Silicon 上原生运行的轻量级 speculative decoding 推理加速工具,支持 DeepSeek 的 DSpark 和 z-lab 的 DFlash 两种无损草稿模型,专为 Qwen3 和 Gemma-4 系列大语言模型优化。其核心特点是:基于 MLX 框架实现零精度损失的加速(输出与标准解码完全一致),支持命令行、Python API 及 OpenAI 兼容服务接口,无需外部服务器框架。适用于本地 Mac 环境下的低延迟 LLM 推理、开发调试及桌面端 AI 应用集成,尤其适合资源受限但需高响应一致性的场景。

Python
MIT License
210
Stars
18
Forks
1
Watchers
4
Issues

Star 增长

今日0
近 7 天0
近 30 天+62
综合评分50.04
默认分支main