GPUs / NVIDIA
RTX 4060
- Memory
- GDDR6
- Bus
- 128-bit
- Published peak
- 272 GB/s
- Capacity
- 8 GB
Measured
- Read
- 250.5GB/s
- Write
- 244.6GB/s
- Copy
- 244.1GB/s
- Random
- 39.1GB/s
- Latency
- 235ns
- Sustain
- 246.6GB/s
Medians of 5 verified runs from 1 operator. Read reaches 92 % of the published peak.
Local AI speed
| Model | Parameters | 4-bit, tokens/s | 8-bit, tokens/s |
|---|---|---|---|
| Llama 3.1 8B | 8.03B | ≤ 62 | Does not fit |
| Qwen 2.5 14B | 14.7B | ≤ 34 | Does not fit |
| gpt-oss-20b | 21B, 3.6B active | Does not fit | Does not fit |
| Mistral Small 24B | 24B | Does not fit | Does not fit |
| Qwen 2.5 32B | 32.5B | Does not fit | Does not fit |
| Llama 3.3 70B | 70.6B | Does not fit | Does not fit |
| gpt-oss-120b | 117B, 5.1B active | Does not fit | Does not fit |