A new arXiv paper introduces AgentLens, a production-assessed benchmark designed to provide detailed trajectory reviews for evaluating interactive co…
OriginalAI Intel
Last 90 days · 2579 total
New research on arXiv investigates the integration of Large Language Model (LLM)-powered reasoning into agent-based modeling (ABM), enhancing ABM's c…
OriginalA new arXiv paper introduces TriRoute, a unified learned routing mechanism for jointly adapting attention, Mixture-of-Experts (MoE), and KV-Cache all…
OriginalA new medical finetuned model, Reasoning-Medical0.1-27B, based on Qwen3.5-27B, is claiming to surpass MedGemma in performance for medical reasoning t…
OriginalClaude Code's latest release, v2.1.202, introduces a "Dynamic workflow size" setting to configure agent counts for dynamic workflows.
OriginalA Reddit user argues that the standard free ChatGPT model performs poorly, suggesting it might be a sub-20B model, especially when compared to local…
OriginalOllama v0.31.2-rc2 now allows integrated GPU (iGPU) offloading for multimodal projectors with "fit padding," improving efficiency by better handling…
OriginalMistral AI has introduced Robostral Navigate, a new cutting-edge model designed for robotics navigation.
OriginalxAI released Grok 4.5, with GLM-5.2 appearing in its own benchmarks, showing 2.6 percentage points behind on SWE Bench Pro.
OriginalWillow Frontier Pro has launched, claiming to be the fastest and most accurate dictation model globally.
Original