Needle 3: The 8MB On-Device Model That Beats 10× Larger Models
Needle 3 is Cactus Compute's third-generation on-device foundation model — a single 8-29MB binary that handles tool calling, structured extraction, and text embeddings entirely offline, with no cloud and no network. Its defining trick is the 'intelligence ladder': one set of weights where every depth from 2 to 20 layers is a complete, deployable model, so a developer picks a 2-layer version for a smartwatch or the full 20-layer for a flagship phone — all from one training run. This deep dive covers the architecture (Monarch Hadamard MLP, grouped-query attention, an engram n-gram memory that lets the 121M model do the arithmetic of a 50M one), the benchmark claims and their fine print — 'passes DeepSeek V4 Flash' means fine-tuned on DroidCall, not a general win — the Pebble smartwatch partnership, and the open-core business model. The honest verdict: a serious specialist for offline, act-on-voice use cases, not a general-purpose chat model.























