DeepResearch RewardBench
Evaluates reward models for assessing deep-research reports.
Paper & authors
DeepResearch RewardBench: Evaluating Reward Models for Deep Research Report
* Equal contribution; # Corresponding author.
Jiaheng LiuNANJING UNIVERSITY · LINK LAB
9 research contributions
Reward models, reasoning critique, learned judges, and efficient evaluation.
Search by name or capability, or narrow by publication status and authorship.
Evaluates reward models for assessing deep-research reports.
DeepResearch RewardBench: Evaluating Reward Models for Deep Research Report
* Equal contribution; # Corresponding author.
Evaluates chain-of-thought efficiency and redundancy using a graph-based framework.
CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs
* Equal contribution; # Corresponding author.
Tests critics' judgments of semantic correctness in mathematical formalization.
CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization
* Equal contribution; # Corresponding author.
Evaluates Chinese reward models alongside the COIG-P preference dataset.
COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values
* Equal contribution; # Corresponding author.
Evaluates reward models for long-form text generation.
Long-form RewardBench: Evaluating Reward Models for Long-form Generation
* Equal contribution; # Corresponding author.
Develops generative judges that reason before evaluating model outputs.
Think-J: Learning to Think for Generative LLM-as-a-Judge
* Equal contribution; # Corresponding author.
Tests error detection in long chains of reasoning.
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
* Equal contribution; # Corresponding author.
Studies compact multimodal evaluation across multiple tasks.
LIME: Less Is More for MLLM Evaluation
* Equal contribution; # Corresponding author.
Benchmarks post-training sparsity across algorithms and models.
PTSBench: A Comprehensive Post-Training Sparsity Benchmark Towards Algorithms and Models
* Equal contribution; # Corresponding author.
Try a broader keyword or clear the filters.
This collection includes co-authored benchmarks, companion evaluation datasets, and evaluation methods. Each work has one primary category; topic tags capture related capabilities. Corresponding-author labels refer to Jiaheng Liu.