FPGA transformer-inference layer
Apex implements one transformer decoder layer in RTL on an FPGA and places the KV-cache codec inside the data path.
From Github Awesome — GitHub Trending Weekly #45: comet, blobatar, cumora, herdr, barehands, ip-as-logo-skill, Aura, md2hd at 06:48
Problem: KV-cache compression performed later in software can add overhead to inference hardware.
For: Teams researching specialized hardware for transformer inference.
Examples
- Qwen 2.5-B: the transformer model targeted by the implementation.
Soon you can unlock the full business plan.
Behind this: 6 build steps · 2 tools and how each is used · how to validate demand · 1 thing the video never answers.