Apex Compute demonstrates real-time, efficient billion-parameter LLM inference on FPGA hardware, significantly outperforming NVIDIA Jetson in both speed and energy efficiency. Running the Gemma 3 1B model via llama.cpp, our architecture achieves 15.6 tokens/sec at just 4.5 W, more than 2× faster and 3× more efficient than CUDA on Jetson Orin Nano.
Despite operating on slower memory, lower frequency, older process technology, and an FPGA—not an ASIC—Apex Compute consistently accelerates core transformer operations by at least an order of magnitude compared to CUDA. With a straightforward transition to a 12nm ASIC and LPDDR5 bandwidth, we expect more than 20× system speed-up and the possibility of a sub-1-watt chip capable of running billion-parameter models.
This capability lets robotics, drones, wearables, and smart-home devices run advanced AI models offline with minimal overhead, bringing AI power anywhere in the physical world.
Explore the implementation and sample Gemma 3 1B model on GitHub, or purchase a Kintex-7 board with two Apex Compute Unified Engines.
* Gemma 3 1B was tested with llama.cpp on NVIDIA Jetson with CUDA enabled. Because llama.cpp still performs some computation on the CPU, the CPU was limited to one thread. GPU-only performance is approximately 8 tokens/second, lower than the real-time demo above.
Apex Compute Architecture vs CUDA
| Operation | Apex μs | CUDA μs | Speedup | 12nm ASIC projection |
|---|---|---|---|---|
| RMSNorm of input (1,1152) | 0.785 | 21.448 | 27.33× | 65× |
| Q K V projection (1,1024)(1,256)(1,256) | 112.054 | 251.709 | 2.25× | 22.5× |
| RMSNorm for Q and K (1,1024)(1,256) | 2.313 | 46.363 | 20.04× | 48× |
| RoPE for Q and K (4,256)(1,256) | 2.837 | 19.161 | 6.75× | 16× |
| kqT + softmax + transpose v | 7.895 | 37.408 | 4.74× | 11.4× |
| Matmul for attention 4× (1,9)×(9,256) | 3.426 | 71.1 | 20.75× | 8.2× |
| MVM for attention output projection | 75.054 | 152.567 | 2.03× | 20× |
| RMSNorm output projection (1,1152) | 0.785 | 21.451 | 27.33× | 65× |
| Residual add for attention output | 2.313 | 12.388 | 5.36× | 12.9× |
| RMSNorm for pre-MLP | 0.836 | 21.06 | 25.19× | 60× |
| MVM for MLP gate + GELU | 505.622 | 836.437 | 1.65× | 16.5× |
| MVM for MLP up | 504.809 | 864.549 | 1.71× | 17.1× |
| MVM for MLP down + element-wise multiply | 505.148 | 325.476 | 0.64× | 6.4× |
| RMSNorm for post-MLP | 0.785 | 20.988 | 26.74× | 64× |
| Output residual add | 2.647 | 12.262 | 4.63× | 11× |
| Large MVM (1,1152)×(1152,262144) | 19041.4 | 59500 | 3.12× | 31.2× |
| Model speed | 15.6 tokens/s | 7.8 tokens/s | 2.00× | ≥ 20× |
| Gemma 3 — 1B | Apex (4.5 W) | CUDA Jetson (7 W) | Apex / CUDA | 12nm ASIC projection |
|---|---|---|---|---|
| Tokens / sec | 15.63 | 7.59 | 2.06× faster | ≥ 20× faster |
| Joules / token | 0.287 | 0.922 | 3.20× more efficient | 10–20× (simulation pending) |
Prototype hardware comparison
| Specification | Apex Kintex UltraScale+ KU5P | NVIDIA Jetson Orin Nano 8GB | Comparison |
|---|---|---|---|
| Memory | 1333 MHz, 32-bit DDR4 | 2133 MHz, 128-bit LPDDR5 | 6.4× slower than NVIDIA |
| Engine speed | 333 MHz | 408 MHz GPU | 1.23× slower than NVIDIA |
| Precision | INT4 / FP4 / INT8 matrix, BF16 vector | BF16 / FP32 / INT4 / INT8 | — |
| Power | 4.5 W | 7 W compute module only | 1.56× less power |
| Process | 16nm FinFET | 8nm Samsung | NVIDIA uses a newer process |
Potential 12nm ASIC speed-up
Our design is primarily memory-bandwidth-bound for weight matrix multiplication and compute-bound for SRAM-based operations. By transitioning to LPDDR5, we can realize a proportional performance increase directly in our architecture because there are no structural bottlenecks or external overheads in the hardware.
A standard LPDDR5 controller on a 12nm process node can achieve 6400 Mbps on a 128-bit interface, delivering approximately 104 GB/s of bandwidth—about 10× higher than our current FPGA-based system. With this improvement, the design enables roughly a 20× speed-up compared to NVIDIA Jetson platforms.
At this bandwidth, our compute engine only needs to operate at 800 MHz to fully utilize the 256-pair element-processing in the Apex Compute Unified Engine Unit. We already reach 333 MHz timing closure on a 16nm FPGA, so scaling to approximately 800 MHz on a 12nm ASIC is straightforward and expected. There is a high chance we can achieve a sub-1 W compute chip capable of running billion-parameter AI models efficiently.
More benchmark results
PyTorch is less efficient than optimized CUDA kernels for LLMs. We’ve included reference benchmarks below for comparison.
Apex Compute Architecture vs PyTorch CUDA
| Operation | Apex μs | PyTorch CUDA μs | Speedup |
|---|---|---|---|
| RMSNorm of input | 1.678 | 102.309 | 60.97× |
| Q K V projection | 112.054 | 215.66 | 1.92× |
| RMSNorm for Q and K | 2.313 | 189.582 | 81.96× |
| RoPE for Q and K | 2.837 | 114.915 | 40.51× |
| kqT + softmax + transpose v | 7.895 | 79.749 | 10.10× |
| Matmul for attention | 3.426 | 20.736 | 6.05× |
| MVM for attention output projection | 75.054 | 134.853 | 1.80× |
| RMSNorm output projection | 0.785 | 114.009 | 145.23× |
| Residual add for attention output | 2.313 | 9.44 | 4.08× |
| RMSNorm for pre-MLP | 0.836 | 101.668 | 121.61× |
| MVM for MLP gate + GELU | 505.622 | 655.281 | 1.30× |
| MVM for MLP up | 504.809 | 641.656 | 1.27× |
| MVM for MLP down + element-wise multiply | 505.148 | 637.111 | 1.26× |
| RMSNorm for post-MLP | 0.785 | 117.07 | 149.13× |
| Output residual add | 2.647 | 9.473 | 3.58× |
| Large MVM | 19041.4 | 67140.2 | 3.53× |
| Model speed | 15.6 tokens/s | 2.9 tokens/s | 5.2× |
Run it yourself
Explore the open implementation or contact us about evaluating Apex Compute for your edge-AI workload.
Contact us