AI product Open source
Needle 2 is an open-source, 45-million-parameter foundation model for tool calling, device use, and structured extraction on small devices. Its weights are compressed into a single 14 MB binary and integrated with its inference engine; the repository reports that a full session uses about 28 MB of RAM and performs inference without network access after setup. The model converts text and tool descriptions into schema-constrained JSON calls using a byte-level grammar, attaches a learned confidence score for threshold-based escalation, and can retrieve the top five tools from a larger catalogue. It uses a 256-token sliding window with tools pinned as key-value sinks to bound memory, and its Simple Attention Network combines a Hadamard MLP, grouped-query attention, engram key-value memory, and multi-lane hyper-connections. The Python package supports inference, LoRA fine-tuning, and checkpoint export; the inference engine is fetched from Hugging Face and cached, while the training stack is an optional installation.
3 uses taken from transcripts — each links to the moment in the video.
A 45-million-parameter, 14 MB open-source model for tool calling, device use, and structured extraction on small devices. It runs in low memory, returns schema-constrained JSON tool calls, and provides confidence scores for escalation.
A 14 MB open-source tool-calling model designed for small devices such as phones and wearables. It runs offline with constrained token generation, supports LoRA fine-tuning, and provides confidence scores for escalation.
An open-source 45-million-parameter model for tool calling, device use, and structured extraction. Its small binary runs with low memory, returns schema-constrained JSON tool calls, and provides confidence scores for escalation.
3 in the library.