AI product Open source

Needle 2

Needle 2 is an open-source, 45-million-parameter foundation model for tool calling, device use, and structured extraction on small devices. Its weights are compressed into a single 14 MB binary and integrated with its inference engine; the repository reports that a full session uses about 28 MB of RAM and performs inference without network access after setup. The model converts text and tool descriptions into schema-constrained JSON calls using a byte-level grammar, attaches a learned confidence score for threshold-based escalation, and can retrieve the top five tools from a larger catalogue. It uses a 256-token sliding window with tools pinned as key-value sinks to bound memory, and its Simple Attention Network combines a Hadamard MLP, grouped-query attention, engram key-value memory, and multi-lane hyper-connections. The Python package supports inference, LoRA fine-tuning, and checkpoint export; the inference engine is fetched from Hugging Face and cached, while the training stack is an optional installation.

View repository Visit site Mentioned in 3 videos ↓

What Needle 2 is used for

3 uses taken from transcripts — each links to the moment in the video.

  • A 45-million-parameter, 14 MB open-source model for tool calling, device use, and structured extraction on small devices. It runs in low memory, returns schema-constrained JSON tool calls, and provides confidence scores for escalation.

  • A 14 MB open-source tool-calling model designed for small devices such as phones and wearables. It runs offline with constrained token generation, supports LoRA fine-tuning, and provides confidence scores for escalation.

  • An open-source 45-million-parameter model for tool calling, device use, and structured extraction. Its small binary runs with low memory, returns schema-constrained JSON tool calls, and provides confidence scores for escalation.

Videos mentioning Needle 2

3 in the library.