Engine Benchmarks SoC GPU Careers

20× performance vs. NVIDIA Jetson.Under 10 W. At one-fifth the cost.

20× Projected performance compared to NVIDIA Jetson
8 TFLOPS BF16 compute
32 Unified Engines
104 GB/s LPDDR5 bandwidth
<10 W Target compute power
The Unified Engine

FPGA prototype is ready for licensing and deployment.

FPGA Logic footprint
78,348 CLB LUTs
50,045CLB registers
197DSP slices
30 / 16URAM / BRAM tiles
1.05 MBon-chip SRAM
366 MHzTiming closure on Kintex UltraScale+
47 GFLOPSBF16 peak; matrix multiplication reaches 95% utilization
256 or 512-bitAXI master width
BF16 / FP8 / INT8 / INT4 / FP4 / TQ4Matrix, vector and block-quantized formats

Logic circuits optimized for transformer architecture

Matrix multiplication, softmax, element-wise operations and many other operators execute at high efficiency.

Unified Engine on the fabric Binary processing layer Compute chip layer
Benchmarks

Up to 95.6% FLOPS utilization on matrix operations.

RTL on Kintex UltraScale+ in real time. M=K=N=1024, 42 GFLOPS BF16 peak.

OperationGFLOPSof peak
BF16 matrix multiplication — A·Bᵀ 40.17 95.6%
Matmul + bias + activation (GELU / SiLU) 40.03 95.3%
Quantized matrix multiplication — BF16 × INT4/FP4 40.03 95.3%
Softmax fused into matrix multiplication 37.76 89.9%
Memory-efficient attention kernel ~90%
Streaming quantized matvec — decode mode 31.33 74.6%
Attention kernel, causal, bias enabled

Q·Kᵀ, scaling, bias, softmax and the value matrix multiply in one kernel.

head_dim 64 head_dim 128 head_dim 256
40 30 20 10 0
Theoretical peak 42 GFLOPS
64 128 256 512 1024 2048 3072 4096 6144 8192

Performance by sequence length, in GFLOPS. Peak: 36.87 GFLOPS — 87.8% of theoretical — at sequence length 3072 and head_dim 256.

Multi-engine, cycle-accurate

Tensor parallelism across multiple engines.

A 256×2048 @ 2048×1024 matmul split 192/64 across two engines, recorded by the 8,192-timestamp trace buffer.

Engine 0192×2048 @ 2048×1024
Queue
DMA from DRAM
Compute
DMA to DRAM
Halt
Engine 164×2048 @ 2048×1024
Queue
DMA from DRAM
Compute
DMA to DRAM
Halt
05 ms10 ms15 ms20 ms

Engine 1 finishes its smaller partition first and halts on the hardware synchronization flag until Engine 0 completes.

Head to head · Gemma 3 1B

Faster and more efficient than the NVIDIA Jetson Orin Nano.

Our engine runs at 333 MHz on 16nm FPGA fabric with a 32-bit DDR4 bus. The Orin Nano is an 8nm ASIC with 128-bit LPDDR5.

2.06× faster decoding 15.63 vs 7.59 tokens/s, llama.cpp with CUDA
3.20× more energy efficient 0.287 vs 0.922 joules per token
5.2× faster than PyTorch CUDA 2.9 tokens/s on PyTorch CUDA
SpecificationApex · Kintex UltraScale+ KU5PJetson Orin Nano 8GB
Memory1333 MHz, 32-bit DDR42133 MHz, 128-bit LPDDR5
Engine clock333 MHz408 MHz GPU
ProcessFPGA fabric8nm Samsung, ASIC
PrecisionBF16 / FP8 / INT8 / INT4 / FP4, plus block-quantized TQ4BF16 / FP32 / INT8 / INT4
Power4.5 W total7 W compute module only
Gemma 3 1B decode15.63 tokens/s7.59 tokens/s

Apex Compute architecture vs CUDA

Microseconds per call, Gemma 3 1B decode
OperationApex µsCUDA µsSpeedup
RMSNorm of input (1,1152)0.78521.44827.33×
Q·K·V projection112.054251.7092.25×
RMSNorm for Q and K2.31346.36320.04×
RoPE for Q and K2.83719.1616.75×
k·qᵀ + softmax + transpose v7.89537.4084.74×
Attention matmul, 4×3.42671.10020.75×
Attention output projection75.054152.5672.03×
MLP gate + GELU505.622836.4371.65×
MLP up projection504.809864.5491.71×
MLP down + element-wise multiply505.148325.4760.64×
Large MVM (1,1152)×(1152,262144)19,04159,5003.12×
Whole model15.63 t/s7.59 t/s2.06×

Tested with llama.cpp on Jetson with CUDA enabled and the CPU limited to one thread. GPU-only performance is approximately 8 tokens/s.

Performance roadmap

A 20× leap, in two stages.

02
12nm silicon · tape-out in progress

GlobalFoundries partnership unlocks an additional 10× gain on production silicon.

Kernel utilization Reproduce on GitHub ↗
Attention kernelup to 90%
Softmax utilization> 90%
Matmul utilization> 90%
Normalized raw performance · Gemma 3 1B / Llama 3.2 1B 20×
Jetson Orin NanoSamsung 8nm ASIC · Shipping
Axelera Metis AIPU4 cores @ 800 MHz · Shipping2.2×
Apex · FPGA1 engine @ 333 MHz · FPGA prototype
Apex · 12nm ASIC32 engines @ 1.6 GHz · Tape-out Q3 '2620×
A single Apex engine on an off-the-shelf FPGA (with one-third of the frequency) already gets 90% of Axelera's shipping ASIC throughput.

Axelera and Apex FPGA figures measured end-to-end on Llama 3.2 1B. Jetson and Apex FPGA figures measured end-to-end on Gemma 3 1B. The 12nm simulation assumes 104 GB/s LPDDR5 bandwidth and 32 Apex Compute engines running in parallel. The 2-engine and 6 GB/s DDR bandwidth scenario is reproducible on FPGA via our public repo.

A complete SoC for edge AI.

GlobalFoundries will manufacture our first complete edge-AI SoC, built around the Unified Engine, with MIPS as the SoC design partner for the surrounding system architecture.

Projected · 12 nm ASIC

Projected 12 nm ASIC performance.

Memory-bandwidth-bound on weight matmul, compute-bound on SRAM operations. No structural bottleneck in between.

LPDDR5 provides roughly 104 GB/s, about 10× the bandwidth of our current system. The ASIC engine is designed to run at 1.6 GHz.

≥20×system speedup over Jetson
<10 Wtarget compute power
104 GB/sLPDDR5 bandwidth, 128-bit at 6400 Mbps
1.6 GHzASIC engine clock
Compiler and runtime

Software written for exactly one machine.

One program, any sequence length

One compiled prefill program handles any prompt up to 192 tokens. The instruction binary shrank roughly 10×.

Instruction scheduling visual

Quantization chosen per block

INT4 or NVFP4 is selected for each 64-element K-block by minimizing MSE. Norms and embeddings stay BF16.

Quantization pipeline visual

Field updates without a reboot

One command reflashes, warm-boots the fabric and re-scans PCIe. No power cycle, no host reboot.

Modular update blocks visual

Text, vision, speech and localization models run on the engine today.

Every model runs in the automated test suite.

Prototype hardware

Try it on your own bench.

An FPGA PCIe card with two Unified Engines, the XDMA driver, the Python runtime and a working Gemma 3 1B example. The package includes hardware design updates.

Apex Compute FPGA prototype card
Where it goes

AI at the edge.

Run LLMs, vision transformers and VLA architectures on-device.

Autonomous drone

Drones

Vision and navigation on-board, within a power budget that keeps the aircraft aloft.

Autonomous vehicle perception

Autonomous vehicles

Perception, fusion and planning with bounded latency, on the vehicle.

Industrial robot arm

Robotics

Multimodal and localization models on-device. The whole policy stays on the robot.

Off-cloud compute

Off-cloud systems

For deployments where data cannot leave the building.

Also at Apex In development

AI acceleration, with graphics alongside.

We are developing GPU IP to complement the Unified Engine with modern graphics and compute capabilities. The design targets Vulkan and DirectX, bringing AI acceleration and graphics together on the same platform.

Vulkan 1.4API target
DirectX 12 UltimateFeature set target
Modern workloadsGraphics and compute together
Coherent memoryUnified device hierarchy
Backed by
Angels from
Zero Matter Tekion Argmax Insider One

Evaluate the engine on your workload.

Send us the model and the power envelope.

Request a demo
Apex Compute

© 2026 Apex Compute. All rights reserved.