Back to AI intel
趋势
搞钱

NVIDIA Puzzle-75B-A9B NVFP4 Performance on 3x3090 GPUs Noted

AI intel briefing

Core summary

One sentence to understand this update

A discussion highlights NVIDIA Puzzle-75B-A9B NVFP4 achieving 132 tokens/second on 3x3090 GPUs, questioning the lack of other models optimized for this 75B MoE configuration.

Impact & opportunity

What this could mean

Developers with multi-24GB GPU setups can consider models like Puzzle-75B-A9B for efficient inference, and this discussion may spur exploration of under-served MoE model sizes.