int8bit-quantization-v100-32gb

IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-sA-appendix.md

Created 2026-08-25T02:58:12+00:00

OPT-13B and GPT-Neo-20B were run on a single V100 Volta GPU (32 GB VRAM) using int8bit quantization with no notable performance drop

Summary

Running two of the larger open-source language models (13B and 20B parameters) on a single older-generation GPU is practical when you compress the model weights to 8-bit precision, and the output quality stays essentially intact. This means a modest, single-GPU setup is enough to deploy models in that size range, removing the need for multi-card clusters or cutting-edge hardware to get usable results.