DumpsterCluster: Repurposing Retired GPUs for LLM Inference
A research article available on arXiv (2608.14614) examines the potential of creating a 'DumpsterCluster' using decommissioned GPUs for contemporary large language model (LLM) inference. The researchers built a cluster comprising 128 GPUs from used parts and operated it for a year. Currently, this DumpsterCluster is valued at around $22,000, in stark contrast to the $600,000 price tag of an 8-GPU B200 system, showcasing significant cost benefits. By employing pipeline-parallel optimizations, the V100-based cluster delivers competitive throughput for LLaMA-70B, proving its feasibility for production. Nonetheless, the deployment uncovers critical context dependencies: older GPUs require much more energy per token, making the total cost of ownership advantageous only in areas with low electricity costs. The findings emphasize the promising future for retired accelerators, contributing to both economic and environmental sustainability in AI infrastructure.
Key facts
- Paper arXiv:2608.14614 investigates repurposing retired GPUs for LLM inference.
- A 128-GPU DumpsterCluster was built from second-hand components and run for one year.
- DumpsterCluster costs $22K versus $600K for an 8-GPU B200 system.
- Pipeline-parallel optimizations enable competitive LLaMA-70B throughput on V100 GPUs.
- Older GPUs have higher energy consumption per token.
- Total cost of ownership is favorable only in regions with inexpensive electricity.
- The study addresses economic and environmental sustainability of AI hardware.
- The cluster was built from scratch using only second-hand components.
Entities
Institutions
- arXiv