Already delivers 2× performance vs. NVIDIA Jetson Orin on Gemma 3 1B.
Watch · Jetson Orin vs. Apex FPGA, side-by-sideMatrix multiplication, softmax and element-wise ops all run near peak efficiency.
RTL on Kintex UltraScale+ in real time. M=K=N=1024, 42 GFLOPS BF16 peak.
Performance by sequence length, in GFLOPS. Peak: 36.87 GFLOPS — 87.8% of theoretical — at sequence length 3072 and head_dim 256.
A 256×2048 @ 2048×1024 matmul split 192/64 across two engines, recorded by the 8,192-timestamp trace buffer.
Engine 1 finishes its smaller partition first and halts on the hardware synchronization flag until Engine 0 completes.
Our engine runs at 333 MHz on 16nm FPGA fabric with a 32-bit DDR4 bus. The Orin Nano is an 8nm ASIC with 128-bit LPDDR5.
Tested with llama.cpp on Jetson with CUDA enabled and the CPU limited to one thread. GPU-only performance is approximately 8 tokens/s.
Already delivers 2× performance vs. NVIDIA Jetson Orin on Gemma 3 1B.
Watch · Jetson Orin vs. Apex FPGA, side-by-sideGlobalFoundries partnership unlocks an additional 10× gain on production silicon.
One Apex engine on an off-the-shelf FPGA — at a third of the clock speed — already hits 90% of Axelera's shipping ASIC throughput.
Axelera and Apex FPGA figures measured end-to-end on Llama 3.2 1B. Jetson and Apex FPGA figures measured end-to-end on Gemma 3 1B. The 12nm simulation assumes 104 GB/s LPDDR5 bandwidth and 32 Apex Compute engines running in parallel. The 2-engine and 6 GB/s DDR bandwidth scenario is reproducible on FPGA via our public repo.
GlobalFoundries will manufacture our first edge-AI SoC, built around the Unified Engine. MIPS is our SoC design partner for the surrounding system architecture.
Memory-bandwidth-bound on weight matmul, compute-bound on SRAM operations. No structural bottleneck in between.
LPDDR5 provides roughly 104 GB/s, about 10× the bandwidth of our current system. The ASIC engine is designed to run at 1.6 GHz.
One compiled prefill program handles any prompt up to 192 tokens — and the instruction binary is 10× smaller.
INT4 or NVFP4 is selected for each 64-element K-block by minimizing MSE. Norms and embeddings stay BF16.
One command reflashes, warm-boots the fabric and re-scans PCIe. No power cycle, no host reboot.
Every model runs in the automated test suite.
An FPGA PCIe card with two Unified Engines, the XDMA driver, the Python runtime and a working Gemma 3 1B example. The package includes hardware design updates.
Run LLMs, vision transformers and VLA architectures on-device.
Vision and navigation on-board, within a power budget that keeps the aircraft aloft.
Perception, fusion and planning with bounded latency, on the vehicle.
Multimodal and localization models on-device. The whole policy stays on the robot.
For deployments where data cannot leave the building.
We're building GPU IP alongside the Unified Engine, targeting Vulkan and DirectX — so AI acceleration and graphics run on the same platform.