Aivora

AivoraGlobal AI knowledge, made clear

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast Diffusion LLMs
arXivLLM

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast Diffusion LLMs

Flash-dLLM:為擴散大語言模型打造的 I/O 感知 KV 快取與平行解碼加速框架

Flash-dLLM is a training-free inference acceleration framework that dramatically speeds up Diffusion LLMs (dLLMs) using an IO-aware fused KV cache and a self-contained draft-and-verify decoding strategy.

2 min read

Latest Research

CliffCompaction: Cost-Efficient Context Compaction for Long-Horizon Coding Agents
arXivAI Coding

CliffCompaction: Cost-Efficient Context Compaction for Long-Horizon Coding Agents

CliffCompaction:長任務程式碼 Agent 的高效能 context 自動壓縮技術

CliffCompaction is an autocompaction technique for long-horizon coding agents that reduces token costs by up to 50% and prevents context drift by strictly truncating rather than rewriting original context.

2 min read
SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving
arXivAI Coding

SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving

SWE-Serve:首個針對「生產級推論服務」的 AI Agent 軟體工程基準測試

SWE-Serve is a new benchmark featuring 53 real-world SGLang tasks, designed to evaluate AI agents' capability to implement complex features and achieve production correctness across the inference serving stack.

2 min read
Hijacking MCP Agents: How A2M Exposes Semantic Supply-Chain Risks
arXivAI Safety

Hijacking MCP Agents: How A2M Exposes Semantic Supply-Chain Risks

破解 MCP 生態系:新型 A2M 攻擊框架揭示 AI Agent 的語意供應鏈安全威脅

This study introduces A2M, a two-stage black-box hijacking framework, exposing critical semantic supply-chain vulnerabilities in Model Context Protocol (MCP) agents and highlighting the need for stronger security.

2 min read
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast Diffusion LLMs
arXivLLM

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast Diffusion LLMs

Flash-dLLM:為擴散大語言模型打造的 I/O 感知 KV 快取與平行解碼加速框架

Flash-dLLM is a training-free inference acceleration framework that dramatically speeds up Diffusion LLMs (dLLMs) using an IO-aware fused KV cache and a self-contained draft-and-verify decoding strategy.

2 min read

AI Agent

View all

AI Coding

View all
CliffCompaction: Cost-Efficient Context Compaction for Long-Horizon Coding Agents
arXivAI Coding

CliffCompaction: Cost-Efficient Context Compaction for Long-Horizon Coding Agents

CliffCompaction:長任務程式碼 Agent 的高效能 context 自動壓縮技術

CliffCompaction is an autocompaction technique for long-horizon coding agents that reduces token costs by up to 50% and prevents context drift by strictly truncating rather than rewriting original context.

2 min read
SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving
arXivAI Coding

SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving

SWE-Serve:首個針對「生產級推論服務」的 AI Agent 軟體工程基準測試

SWE-Serve is a new benchmark featuring 53 real-world SGLang tasks, designed to evaluate AI agents' capability to implement complex features and achieve production correctness across the inference serving stack.

2 min read

AI Research

View all

AI Safety

View all
Hijacking MCP Agents: How A2M Exposes Semantic Supply-Chain Risks
arXivAI Safety

Hijacking MCP Agents: How A2M Exposes Semantic Supply-Chain Risks

破解 MCP 生態系:新型 A2M 攻擊框架揭示 AI Agent 的語意供應鏈安全威脅

This study introduces A2M, a two-stage black-box hijacking framework, exposing critical semantic supply-chain vulnerabilities in Model Context Protocol (MCP) agents and highlighting the need for stronger security.

2 min read