AI product Open source
Naive-N0.5-Flash is an open-weight mixture-of-experts language model from NaiveAI for coding and AI research. It has 309 billion total parameters, 15.5 billion active parameters, and a native one-million-token context window implemented with a hybrid of 128-token Sliding-Window Attention and DeepSeek Sparse Attention; the latter's indexer scans the history while its backbone attends to the top 2,048 selected tokens, with no full-attention layers. The model is released under the MIT License with inference code, and supports FP8 mixed-precision inference through Transformers on FP8-capable NVIDIA GPUs; the repository reports approximately 315 GB of model weights. NaiveAI also describes NaiveRT, an inference system using mega-kernel fusion, Programmatic Dependent Launch, and speculative decoding.
1 use taken from transcripts — each links to the moment in the video.
An open-weight coding and AI-research model with 309 billion parameters, sparse attention, and a native million-token context window.
1 in the library.