Scaling Long-Form Story Generation via Narrative State Tracking (NstAgent)
以敘事狀態追蹤突破長篇小說生成:免微調的 NstAgent 框架
While LLMs excel in creative writing, maintaining narrative coherence over long texts remains a major challenge, with most methods limited to under 10K words. This paper presents Narrative State Tracking Agent (NstAgent), a training-free framework that structures and tracks characters, history, and future requirements. By testing on extended benchmarks from 10K to 100K words, the researchers demonstrated that NstAgent prevents the usual quality and consistency degradation in long-form generation.
Key points
Scaling to Novel Length
Scales story generation from the typical 10K limit to 100K words, establishing a foundation for AI-generated novels.
Structured Narrative Tracking
Automatically tracks characters, past events, and future narrative requirements to maintain strict logical consistency.
Training-Free Agent Framework
Utilizes an out-of-the-box agentic workflow that works with existing LLMs without additional fine-tuning.
No Quality Degradation
Empirical tests show that writing quality and narrative coherence do not degrade even as the word count scales up.
How it works
Why it matters
Conventional LLMs struggle with long-form writing due to context limitations, often leading to logical inconsistencies or forgotten plots. NstAgent demonstrates that structured state management via an agent, without expensive model fine-tuning, enables current LLMs to write coherent 100K-word novels. This unlocks new possibilities for co-creative writing tools, deep worldbuilding, and automated content generation in entertainment.
Who it affects
- AI Developer
- AI Researcher
- Content Creator
How to use it
- 1Co-writing assistant: Helping human authors draft long-form novels while maintaining consistency in worldbuilding and character arcs.
- 2Game quest generation: Dynamically generating long, coherent quest storylines and lore for video games and virtual environments.
Limitations & caveats
- Dependent on the base LLM's reasoning capabilities; lower-tier models may fail to track states accurately.
- Inference costs and API calls may scale significantly as narrative states grow more complex over extremely long texts.
Related
TokenCast: Accurate Token Consumption Forecasting for LLM Agents
TokenCast:精準預測 LLM Agent 執行過程中的 Token 消耗量
This paper introduces TokenCast, a real-time method that accurately forecasts dynamic LLM Agent token consumption within 32.8 ms without extra LLM calls, using composable cost representations.
DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education
DeepEdu-v1:突破硬體與法規限制,專為越南教育量身打造的在地化 AI 代理模型
Addressing data-privacy laws and hardware limits in Vietnam, DeepEdu-v1 leverages the SCALE framework to optimize long-context inference and agentic learning, enabling low-latency, localized AI tutoring.

Google DeepMind Introduces Gemini 3.8 Live with Live Avatar: Real-Time Visual AI Agents for Enterprise
Google DeepMind 發表 Gemini 3.8 Live with Live Avatar:具備即時視覺化身與背景任務處理能力的企業級 AI 代理
Google DeepMind launches Gemini 3.8 Live with Live Avatar, integrating real-time video and speech with background tool execution to power highly responsive, multilingual virtual assistants.