Moondream 3.1-9B-A2B, a new vision-language model with a 9B Mixture-of-Experts architecture (2B active), offers state-of-the-art visual reasoning and…
OriginalAI Intel
Last 90 days · 2635 total
GLM-5.2 is now directly selectable within Claude Code through Hugging Face Inference Providers and hf-claude, making open models more accessible in d…
OriginalGLM-5.2 is temporarily available for free when used with Hugging Face Inference Providers for a duration of five hours.
OriginalAnthropic has collaborated with AE Studio on new research exploring AI capabilities that can be both helpful and dangerous.
OriginalFetchSandbox is a new tool for API integration testing that intelligently remembers and tracks past failures.
OriginalThe latest llama.cpp release b9969 improves Vulkan performance on Adreno GPUs and fixes `llama-cli` crashing with long prompts for q4_0 quantized net…
OriginalOllama v0.31.2-rc2 introduces support for offloading multimodal projector (mmproj) operations to integrated GPUs (iGPUs) on non-Metal systems by usin…
OriginalvLLM v0.25.0 is out, featuring 558 commits and making Model Runner V2 the default for all dense models, building upon previous quantized model suppor…
OriginalThe Hugging Face Gemma Challenge saw over 100 AI agents and humans collaborate to make Gemma 4 inference 5x faster on a single NVIDIA A10G GPU.
OriginalClaude Code v2.1.207 makes Auto mode generally available on Bedrock, Vertex AI, and Foundry without requiring an opt-in, and fixes terminal issues.
Original