Already delivers 2× performance vs. NVIDIA Jetson Orin on Gemma 3 1B.
Watch · Jetson Orin vs. Apex FPGA, side-by-sideMatrix multiplication, softmax, element-wise operations and many other operators execute at high efficiency.
RTL on Kintex UltraScale+ in real time. M=K=N=1024, 42 GFLOPS BF16 peak.
Performance by sequence length, in GFLOPS. Peak: 36.87 GFLOPS — 87.8% of theoretical — at sequence length 3072 and head_dim 256.
A 256×2048 @ 2048×1024 matmul split 192/64 across two engines, recorded by the 8,192-timestamp trace buffer.
Engine 1 finishes its smaller partition first and halts on the hardware synchronization flag until Engine 0 completes.
Our engine runs at 333 MHz on 16nm FPGA fabric with a 32-bit DDR4 bus. The Orin Nano is an 8nm ASIC with 128-bit LPDDR5.
Tested with llama.cpp on Jetson with CUDA enabled and the CPU limited to one thread. GPU-only performance is approximately 8 tokens/s.
Already delivers 2× performance vs. NVIDIA Jetson Orin on Gemma 3 1B.
Watch · Jetson Orin vs. Apex FPGA, side-by-sideGlobalFoundries partnership unlocks an additional 10× gain on production silicon.
A single Apex engine on an off-the-shelf FPGA (with one-third of the frequency) already gets 90% of Axelera's shipping ASIC throughput.
Axelera and Apex FPGA figures measured end-to-end on Llama 3.2 1B. Jetson and Apex FPGA figures measured end-to-end on Gemma 3 1B. The 12nm simulation assumes 104 GB/s LPDDR5 bandwidth and 32 Apex Compute engines running in parallel. The 2-engine and 6 GB/s DDR bandwidth scenario is reproducible on FPGA via our public repo.
GlobalFoundries will manufacture our first complete edge-AI SoC, built around the Unified Engine, with MIPS as the SoC design partner for the surrounding system architecture.
Memory-bandwidth-bound on weight matmul, compute-bound on SRAM operations. No structural bottleneck in between.
LPDDR5 provides roughly 104 GB/s, about 10× the bandwidth of our current system. The ASIC engine is designed to run at 1.6 GHz.
One compiled prefill program handles any prompt up to 192 tokens. The instruction binary shrank roughly 10×.
INT4 or NVFP4 is selected for each 64-element K-block by minimizing MSE. Norms and embeddings stay BF16.
One command reflashes, warm-boots the fabric and re-scans PCIe. No power cycle, no host reboot.
Every model runs in the automated test suite.
An FPGA PCIe card with two Unified Engines, the XDMA driver, the Python runtime and a working Gemma 3 1B example. The package includes hardware design updates.
Run LLMs, vision transformers and VLA architectures on-device.
Vision and navigation on-board, within a power budget that keeps the aircraft aloft.
Perception, fusion and planning with bounded latency, on the vehicle.
Multimodal and localization models on-device. The whole policy stays on the robot.
For deployments where data cannot leave the building.
We are developing GPU IP to complement the Unified Engine with modern graphics and compute capabilities. The design targets Vulkan and DirectX, bringing AI acceleration and graphics together on the same platform.