arXivAI 程式開發
SWE-Serve:首個針對「生產級推論服務」的 AI Agent 軟體工程基準測試
SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving
SWE-Serve 是一個新型基準測試,包含來自 SGLang 的 53 個真實任務,旨在評估 AI Agent 在複雜推論伺服器架構中的程式碼實作與「生產環境正確性」。
2 分鐘閱讀
#swe-serve
1 篇文章
SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving
SWE-Serve 是一個新型基準測試,包含來自 SGLang 的 53 個真實任務,旨在評估 AI Agent 在複雜推論伺服器架構中的程式碼實作與「生產環境正確性」。