Back to AI intel
趋势
搞钱
$2800 Rig with 8x Radeon Pro V620 Achieves High-Speed Qwen3.8-Flash-Next Inference.
AI intel briefing
Core summary
One sentence to understand this update
A $2800 rig featuring 8x Radeon Pro V620 GPUs (256GB VRAM) and a custom vLLM fork can run Qwen3.8-Flash-Next at 60-100 t/s decode and over 3000 t/s prefill.
Impact & opportunity
What this could mean
Builders and researchers can learn from this cost-effective hardware configuration to achieve high-performance local large model inference deployments on a limited budget.
Source
View original