MLCommons introduces a new multi-turn benchmark for measuring LLM serving systems under growing context and closed-loop agent workflows. Featuring Kimi K2.6 and Qwen3.6-35B-A3B. **Category:** Benchmarks, LLM, Agentic AI, Inference