MM-BrowseComp
Evaluates multimodal web browsing and information-seeking agents.
Paper & authors
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
* Equal contribution; # Corresponding author.
Jiaheng LiuNANJING UNIVERSITY · LINK LAB
11 research contributions
Deep research, evidence synthesis, report generation, table analysis, and professional knowledge work.
Search by name or capability, or narrow by publication status and authorship.
Evaluates multimodal web browsing and information-seeking agents.
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
* Equal contribution; # Corresponding author.
Evaluates research reports that interleave text and visual evidence.
TVIR: Building Deep Research Agents Towards Text-Visual Interleaved Report Generation
* Equal contribution; # Corresponding author.
Evaluates deep-research agents in realistic, reproducible settings with multiple files and modalities.
DR³-Eval: Towards Realistic and Reproducible Deep Research Evaluation
* Equal contribution; # Corresponding author.
Evaluates multi-turn agentic table question answering in realistic settings.
How Far Can LLM Agents Reason with Tables? Benchmarking Multi-Turn Agentic Table Question Answering in the Wild
* Equal contribution; # Corresponding author.
Evaluates multilingual, multitask table question answering.
M3TQA: Massively Multilingual Multitask Table Question Answering
* Equal contribution; # Corresponding author.
Tests table question answering across reasoning demands and visual layout complexity.
MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QA
* Equal contribution; # Corresponding author.
Evaluates professional consultation on real-life problems and implicit user needs.
xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems
* Equal contribution; # Corresponding author.
Evaluates academic paper-writing tasks associated with multi-turn generation research.
Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training
* Equal contribution; # Corresponding author.
Localizes errors within the trajectories of deep-research agents.
Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories
* Equal contribution; # Corresponding author.
Tests complex question answering and reasoning over tables.
TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
* Equal contribution; # Corresponding author.
Evaluates the usefulness and overall quality of deep-research reports.
How Far Are We from Genuinely Useful Deep Research Agents?
* Equal contribution; # Corresponding author.
Try a broader keyword or clear the filters.
This collection includes co-authored benchmarks, companion evaluation datasets, and evaluation methods. Each work has one primary category; topic tags capture related capabilities. Corresponding-author labels refer to Jiaheng Liu.