EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent

84d ago · Global · primary source: export.arxiv.org

Multi-source synthesis by The Embedding Report from 2 sources. Every numeric and quoted claim traces to a cited source body (see methodology).

Researchers have introduced EComAgentBench, a benchmark for evaluating shopping agents on long-horizon tasks, and GLM-5.2, a model designed for sustained engineering work.

EComAgentBench is a benchmark of 662 tasks grounded in real Amazon products and reviews, designed to test shopping agents' ability to uncover hidden intent and make decisions over long horizons[1]. Each task requires agents to uncover hidden intent, verify candidates against attributes and review evidence, and commit to a single product within 100 tool calls. The benchmark uses typed, source-tagged rubrics to grade every task, attributing each failure to a requirement and its source. Even the strongest models attain only 57.1% overall accuracy in the benchmark[1]. Meanwhile, GLM-5.2 is a long-horizon model built for sustained engineering work, with a 1M-token context that stably sustains long-horizon work. GLM-5.2 trails Opus 4.8 by only 1% on FrontierSWE benchmark, but still has room to grow on SWE-Marathon benchmark, trailing Opus 4.8 by 13%[2]. GLM-5.2 outperforms both Opus 4.7 and GPT-5.5 on PostTrainBench and is the strongest open-source model on standard coding benchmarks, improving on GLM-5.1 by a wide margin. GLM-5.2 also introduces effort level control, enabling users to balance model capability against task execution speed and computational cost, and its architecture reduces per-token FLOPs by 2.9× at a 1M context length[2].

research-paperapplicationbenchmark

Background sources we checked (7)
  • arxiv.org ↗ As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked. Benchmarks that expose full intent upfront and grade o…
  • huggingface.co ↗ Hugging Face Machine Learning Demos on arXiv Back to Articles ... # Hugging Face Machine Learning Demos on arXiv Published November 17, 2022 Update on GitHub Upvote 1 - - - - - Abubakar Abid abidlabs Follow …
  • info.arxiv.org ↗ ## Hugging Face Spaces ... Hugging Face code repositories, About Hugging Face ... Collaborators: Abubakar Abid, Omar Sanseviero, Ahsen Khaliq, and the Hugging Face team ... Hugging Face Spaces includes links to demos created by the community or the authors themselves. By going to…
  • huggingface.co ↗ Demos on Hugging Face Spaces allow a wide audience to try out state-of-the-art machine learning research without writing any code. Hugging Face and ArXiv have collaborated to embed these demos directly along side papers on ArXiv! ... Thanks to this integration, users can now find…
  • en.wikipedia.org ↗ Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co., Ltd., doing business as DeepSeek, is a Chinese artificial intelligence (AI) company that develops large language models (LLMs). Based in Hangzhou, Zhejiang, DeepSeek is owned and funded by High-Flyer, a Chin…
  • en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…
  • en.wikipedia.org ↗ Douwe Kiela is a Dutch-American research scientist and entrepreneur working in the field of artificial intelligence with a focus on machine learning and natural language processing. He is a research scientist director at Google DeepMind. He previously co-founded and served as CEO…

Sources cited (2)

  1. arxiv.org ↗ E
  2. huggingface.co ↗ D
Spot something wrong? Report an issue