Back to AI intel
趋势
搞钱
NVIDIA Puzzle-75B-A9B NVFP4 Performance on 3x3090 GPUs Noted
AI intel briefing
Core summary
One sentence to understand this update
A discussion highlights NVIDIA Puzzle-75B-A9B NVFP4 achieving 132 tokens/second on 3x3090 GPUs, questioning the lack of other models optimized for this 75B MoE configuration.
Impact & opportunity
What this could mean
Developers with multi-24GB GPU setups can consider models like Puzzle-75B-A9B for efficient inference, and this discussion may spur exploration of under-served MoE model sizes.
Source
View original