arXivAI Research
LongHarness Bench: Stress-Testing Long-Context LLM Harnesses for Efficiency and Reasoning
LongHarness Bench:長文本 LLM 外掛機制的效能與效率壓力測試
LongHarness Bench is a new long-context LLM evaluation benchmark that tests both accuracy and computational efficiency, addressing the saturation of traditional metrics.
2 min read