GPUs / NVIDIA
RTX 3080
- Memory
- GDDR6X
- Bus
- 320-bit
- Published peak
- 760 to 912 GB/s
- Capacity
- from 10 GB
Measured
- Read
- 610.0GB/s
- Write
- 548.2GB/s
- Copy
- 576.1GB/s
- Random
- 48.1GB/s
- Latency
- 274ns
- Sustain
- 572.0GB/s
Medians of 2 verified runs from 1 operator.
Local AI speed
| Model | Parameters | 4-bit, tokens/s | 8-bit, tokens/s |
|---|---|---|---|
| Llama 3.1 8B | 8.03B | ≤ 152 | ≤ 76 |
| Qwen 2.5 14B | 14.7B | ≤ 83 | Does not fit |
| gpt-oss-20b | 21B, 3.6B active | Does not fit | Does not fit |
| Mistral Small 24B | 24B | Does not fit | Does not fit |
| Qwen 2.5 32B | 32.5B | Does not fit | Does not fit |
| Llama 3.3 70B | 70.6B | Does not fit | Does not fit |
| gpt-oss-120b | 117B, 5.1B active | Does not fit | Does not fit |