AI product Open source

Bespoke Nimble

Bespoke Nimble is an open-source model, training recipe, and local scoring library from Bespoke Labs for making typed decisions about text. Given a context and a flat schema of boolean or fixed-choice fields, it selects an allowed answer and returns candidate logits and probabilities without generating a reasoning explanation or free-form answer.

View repository Mentioned in 1 video ↓

Overview

Nimble fine-tunes Qwen3.5-9B with LoRA using contrastive training pairs: two nearly identical examples differ in one relevant fact, which changes the correct label. At inference, the scorer reads the logits for one-token candidate codes and converts them to probabilities with softmax; the Python library constructs the typed result rather than parsing generated JSON. The MLX ParallelScorer processes shared context once and scores fields in parallel, while the CUDA scorer scores each field separately.

The repository includes the model recipe, curation and evaluation code, datasets, examples, tests, and deployment documentation. Bespoke-Nimble-9B can run locally on Apple Silicon or an NVIDIA BF16-capable GPU. It accepts text only, does not support nested schemas or answers outside the supplied choices, and its probabilities are not guarantees of correctness.

What Bespoke Nimble is used for

1 use taken from transcripts — each links to the moment in the video.

  • A Qwen 3.5-9B fine-tune that answers typed multiple-choice and yes-or-no questions without producing reasoning text. It returns candidate answers as JSON with probabilities.

Videos mentioning Bespoke Nimble

1 in the library.