GPUs / Apple
Apple M4
- Memory
- LPDDR5X
- Bus
- 128-bit
- Published peak
- 120 GB/s
- Capacity
- 16 GB
Measured
- Read
- 60.3GB/s
- Write
- 78.9GB/s
- Copy
- 83.7GB/s
- Random
- 9.88GB/s
- Latency
- 285ns
- Sustain
- 83.4GB/s
Medians of 1 verified run from 1 operator. Read reaches 50 % of the published peak.
Local AI speed
| Model | Parameters | 4-bit, tokens/s | 8-bit, tokens/s |
|---|---|---|---|
| Llama 3.1 8B | 8.03B | ≤ 15 | ≤ 7.5 |
| Qwen 2.5 14B | 14.7B | ≤ 8.2 | ≤ 4.1 |
| gpt-oss-20b | 21B, 3.6B active | ≤ 34 | Does not fit |
| Mistral Small 24B | 24B | ≤ 5.0 | Does not fit |
| Qwen 2.5 32B | 32.5B | Does not fit | Does not fit |
| Llama 3.3 70B | 70.6B | Does not fit | Does not fit |
| gpt-oss-120b | 117B, 5.1B active | Does not fit | Does not fit |