AI product Open source
An open-source project that runs a 28.9-million-parameter language model locally on an ESP32-S3 microcontroller, without network connectivity. It uses a tiered memory layout based on Per-Layer Embeddings: frequently accessed activations and normalization weights stay in SRAM, the model core and output head reside in PSRAM, and a roughly 25-million-parameter embedding table is stored in flash; only the few rows needed for each token are read into faster memory. The included TinyStories model generates short stories, while a Barista model provides espresso question answering. The repository notes that the TinyStories model is not designed for general question answering, instruction following, code generation, or factual knowledge, and provides separate scripts for downloading and verifying model assets and deploying a selected model to the board.
1 use taken from transcripts — each links to the moment in the video.
Fits a 28.9-million-parameter language model onto an ESP32-S3, generating every token locally without a network connection. It uses flash for the embedding table and faster memory for the rows needed per token.
1 in the library.