A decoder-only large language model built entirely from scratch in Rust using the Candle framework -- no Python, no PyTorch. Now at v3.0.0 (Phases 29-40 complete), it implements Gated DeltaNet hybrid attention, DeepSeek Sparse Attention, fine-grained MoE with a shared expert, multi-token prediction, on-policy distillation, native quantization-aware training, native video/document understanding, long-horizon tool-use chains, and multiple thinking modes. Released as source-only, with zero pretrained model artifacts distributed.

Fund this project

Unverified URL

The funding manifest has not provided proof via wellKnown that this link is associated with it. Learn more.

Continue