AI product Open source
APEX is an open, verification-first RTL design for an LLM-inference tile developed with Sigmantic AI. It implements one transformer decoder layer—including attention, softmax, RMSNorm, RoPE, SwiGLU, residual processing, and KV-cache compression—and runs Qwen2.5-0.5B through the verified pipeline on FPGA hardware.
The KV-cache codec sits in the datapath: cached keys and values are compressed as they are produced and decompressed when consumed, while an importance unit allocates precision according to which parts of the context matter. Each RTL block is checked against an executable NumPy golden model with bit-exact comparisons and mutation-tested testbenches. The repository covers the inference tile rather than a complete system; DRAM control, PCIe, and NoC are out of scope.
2 uses taken from transcripts — each links to the moment in the video.
An RTL implementation of a transformer decoder layer running Qwen 2.5-b on an FPGA. It places KV-cache compression in the data path and validates blocks against a NumPy reference model.
An open hardware design for large-language-model inference that compresses the key-value cache directly in the data path. It quantizes and decompresses cached context, tracks importance to allocate precision, and targets memory-efficient, verifiable inference on FPGA hardware.
2 in the library.