GPUs / Apple

Apple M4

Memory
LPDDR5X
Bus
128-bit
Published peak
120 GB/s
Capacity
16 GB

Measured

Read
60.3GB/s
Write
78.9GB/s
Copy
83.7GB/s
Random
9.88GB/s
Latency
285ns
Sustain
83.4GB/s

Medians of 1 verified run from 1 operator. Read reaches 50 % of the published peak.

Local AI speed

ModelParameters4-bit, tokens/s8-bit, tokens/s
Llama 3.1 8B8.03B≤ 15≤ 7.5
Qwen 2.5 14B14.7B≤ 8.2≤ 4.1
gpt-oss-20b21B, 3.6B active≤ 34Does not fit
Mistral Small 24B24B≤ 5.0Does not fit
Qwen 2.5 32B32.5BDoes not fitDoes not fit
Llama 3.3 70B70.6BDoes not fitDoes not fit
gpt-oss-120b117B, 5.1B activeDoes not fitDoes not fit
Upper bounds: median read ÷ bytes read per token. Mixture-of-experts models read only their active experts. Fit checks weights against the 16 GB configuration; the KV cache needs room on top.

Top operators

#OperatorScoreReadDate
1OP-FEA44A43.760.3 GB/s2026-09-27

Full Apple M4 leaderboard