arXivAI Agent
TokenCast: Accurate Token Consumption Forecasting for LLM Agents
TokenCast:精準預測 LLM Agent 執行過程中的 Token 消耗量
This paper introduces TokenCast, a real-time method that accurately forecasts dynamic LLM Agent token consumption within 32.8 ms without extra LLM calls, using composable cost representations.
2 min read