A new paper introduces Kara, an efficient reasoning LLM serving method that utilizes sliding-window KV cache compression to reduce decoding latency a…
OriginalAI Intel
Last 90 days · 2526 total
Research introduces PACE, a neuro-symbolic framework for generating plausible and actionable counterfactual explanations to clarify machine learning…
OriginalA research paper presents the Wiola architecture, a novel and original design for efficient Small Language Models (SLMs), built from first principles.
OriginalResearch introduces TokenScope, a method for token-level explainability and interpretability in Large Language Models (LLMs) specifically for code-or…
OriginalA paper presents Auto-FL-Research, an agentic search system designed to automatically discover and optimize federated learning (FL) algorithms.
OriginalResearch focuses on fixed-set robustness in Programming by Example (PBE) systems, investigating example corruption and semantic partition recovery.
OriginalA research paper explores safeguarding LLM agents from misalignment by utilizing provenance analysis to ensure their actions align with user intent.
OriginalA GitHub project demonstrates an auto-charging system for Steam Controllers, utilizing computer vision to guide a magnetic charging puck.
OriginalThe new llama.cpp b9860 release introduces a public C API to expose model file type (quantization) names like "Q8_0" or "Q4_K - Medium."
OriginalClaude Code updated to v2.1.197, making Claude Sonnet 5 the new default model with a native 1M-token context window and promotional pricing until Aug…
Original