TokenCast: Accurate Token Consumption Forecasting for LLM Agents
TokenCast:精準預測 LLM Agent 執行過程中的 Token 消耗量
LLM agent token usage is highly variable, often shifting by an order of magnitude due to dynamic tool feedback and compounding context growth. TokenCast solves this by learning composable cost representations for execution segments, tracking both immediate token use and future context expansion. It dynamically refreshes forecasts as the agent runs. TokenCast achieves an average of 14.5% prediction error reduction and saves 21.3% of token overhead in offline budget-control simulations.
Key points
Dynamic Context Modeling
Solves the compounding cost problem where growing historical context inflates the input size of every subsequent model call.
Composable Cost Representation
Models individual execution segments to capture both immediate usage and subsequent context growth, compounding them for total estimates.
Real-time Low Latency
Requires zero extra LLM calls, achieving a mean cumulative prediction time of just 32.8 ms per run on SWE-bench Verified.
Effective Budget Control
Saves an average of 21.3% in tokens compared to a fixed-budget policy under matched trace completion rates.
How it works
Why it matters
As enterprises scale LLM agents for complex tasks like software engineering, unpredictable and runaway API costs present a major hurdle. TokenCast provides an extremely fast, real-time budgeting guardrail. By predicting final token usage mid-execution with negligible overhead, developers can proactively manage costs and prevent runaway agent bills, making agentic workflows significantly more viable for production.
Who it affects
- AI Developer
- Product Manager
- Enterprise Leader
- AI Researcher
How to use it
- 1Enterprise agent budget guardrails, triggering alerts or early terminations when projected costs exceed set thresholds.
- 2Dynamic routing and policy optimization, choosing cost-effective execution paths based on real-time projected token balances.
Limitations & caveats
- Requires offline training to learn the cost representations for specific agent frameworks, which may need recalibration for entirely new agent designs.
- Forecasting accuracy depends on early execution telemetry; highly erratic initial steps by the agent can degrade overall prediction quality.
Related
Scaling Long-Form Story Generation via Narrative State Tracking (NstAgent)
以敘事狀態追蹤突破長篇小說生成:免微調的 NstAgent 框架
This study introduces NstAgent, a training-free agentic framework that dynamically tracks characters, past events, and future requirements to scale coherent LLM story generation up to 100K words.
DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education
DeepEdu-v1:突破硬體與法規限制,專為越南教育量身打造的在地化 AI 代理模型
Addressing data-privacy laws and hardware limits in Vietnam, DeepEdu-v1 leverages the SCALE framework to optimize long-context inference and agentic learning, enabling low-latency, localized AI tutoring.

Google DeepMind Introduces Gemini 3.8 Live with Live Avatar: Real-Time Visual AI Agents for Enterprise
Google DeepMind 發表 Gemini 3.8 Live with Live Avatar:具備即時視覺化身與背景任務處理能力的企業級 AI 代理
Google DeepMind launches Gemini 3.8 Live with Live Avatar, integrating real-time video and speech with background tool execution to power highly responsive, multilingual virtual assistants.