Skip to content
arXiv cs.AI · Papers

Akashic: A Low-Overhead LLM Inference Service with MemAttention

arXiv:2607.05708v1 Announce Type: new Abstract: Recent LLM-based agent systems continuously accumulate context across multi-turn interactions, tool invocations, and cross-session workflows. Replaying the full history for every request quickly becomes impractical: long contexts increase prefill cost, may exceed context