ClawBench
Evaluates completion of everyday online tasks by AI agents.
Paper & authors
ClawBench: Can AI Agents Complete Everyday Online Tasks?
* Equal contribution; # Corresponding author.
Jiaheng LiuNANJING UNIVERSITY · LINK LAB
11 research contributions
Tool use, everyday online tasks, long-horizon workflows, and agents that build and improve their own skills.
Search by name or capability, or narrow by publication status and authorship.
Evaluates completion of everyday online tasks by AI agents.
ClawBench: Can AI Agents Complete Everyday Online Tasks?
* Equal contribution; # Corresponding author.
Tests whether agents can turn multimodal guides into reusable, evolving skills.
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
* Equal contribution; # Corresponding author.
Evaluates multimodal travel planning under tightly coupled constraints.
WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints
* Equal contribution; # Corresponding author.
Tests whether models can create and evolve their own agent harnesses.
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
* Equal contribution; # Corresponding author.
Evaluates whether models can improve themselves starting from vague goals.
Aspire: Can Models Self-Evolve from Vague Goals?
* Equal contribution; # Corresponding author.
Studies whether self-testing and self-judging translate into self-improvement.
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
* Equal contribution; # Corresponding author.
Evaluates general-purpose agents on market-validated, end-to-end workflows.
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
* Equal contribution; # Corresponding author.
Evaluates long-horizon computer-use tasks in real professional software environments.
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
* Equal contribution; # Corresponding author.
Evaluates long-horizon planning and execution in interactive economic environments.
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
* Equal contribution; # Corresponding author.
Tests tool use at multiple granularities, including multi-turn and multi-tool interactions.
MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
* Equal contribution; # Corresponding author.
Tests spatial planning and interaction in vision-based games.
ING-VP: MLLMs cannot Play Easy Vision-based Games Yet
* Equal contribution; # Corresponding author.
Try a broader keyword or clear the filters.
This collection includes co-authored benchmarks, companion evaluation datasets, and evaluation methods. Each work has one primary category; topic tags capture related capabilities. Corresponding-author labels refer to Jiaheng Liu.