← Evaluation overview

6 research contributions

Instructions & Long Context

Instruction following, multi-turn dialogue, role consistency, and long-context understanding and generation.

Benchmarks & related work

6 works

Search by name or capability, or narrow by publication status and authorship.

Instruction following & dialogue

Inverse IFEval

Tests whether models follow real instructions that conflict with learned training conventions.

Instruction followingRobustness
ICLR 2026Corresponding author
Paper & authors

Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?

Qinyan Zhang*, Xinping Lei*, Ruijie Miao*, Yu Fu, Haojie Fan, Le Chang, Jiafan Hou, Dingling Zhang, Zhongfei Hou, Ziqiang Yang, Changxin Pu, Fei Hu, Jingkai Liu, Xinjie Chen, Jianpeng Jiao, Jiaheng Liu#, Tong Yang#, Zaiyuan Wang#, Ge Zhang#, Wenhao Huang

* Equal contribution; # Corresponding author.

Instruction following & dialogue

RoleAgentBench

Evaluates role-playing agents built from scripts through interactive tasks.

Role playingMulti-turn
NeurIPS 2024 (Datasets and Benchmarks)
Paper & authors

RoleAgent: Building, Interacting, and Benchmarking High-quality Role-Playing Agents from Scripts

Jiaheng Liu*, Zehao Ni*, Haoran Que*, Tao Sun, Zekun Wang, Jian Yang, Jiakai Wang, Hongcheng Guo, Zhongyuan Peng, Ge Zhang, Jiayi Tian, Xingyuan Bu, Ke Xu, Wenge Rong, Junran Peng#, Zhaoxiang Zhang

* Equal contribution; # Corresponding author.

Instruction following & dialogue

MT-Bench-101

Provides fine-grained evaluation of multi-turn dialogue abilities.

DialogueMulti-turn
ACL 2024
Paper & authors

MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Ge Bai, Jie Liu, Xingyuan Bu#, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, Wanli Ouyang

* Equal contribution; # Corresponding author.

Instruction following & dialogue

RoleBench

Evaluates role-playing abilities and consistency in language models.

Role playingDialogue
Findings of ACL 2024Corresponding author
Paper & authors

RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

Noah Wang, Z.Y. Peng, Haoran Que*, Jiaheng Liu#, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, Man Zhang, Zhaoxiang Zhang#, Wanli Ouyang, Ke Xu, Wenhao Huang, Jie Fu, Junran Peng

* Equal contribution; # Corresponding author.

Long-context understanding & generation

HelloBench

Evaluates the ability to generate long, coherent text.

Long-form generationWriting
arXiv preprint, 2024
Paper & authors

HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models

Haoran Que, Feiyu Duan, Liqun He, Yutao Mou, Wangchunshu Zhou, Jiaheng Liu, Wenge Rong, Zekun Moore Wang, Jian Yang, Ge Zhang, Junran Peng, Zhaoxiang Zhang, Songyang Zhang, Kai Chen

* Equal contribution; # Corresponding author.

Long-context understanding & generation

LongIns

Tests instruction following and multitask reasoning over long contexts.

Long contextInstruction following
arXiv preprint, 2024
Paper & authors

LongIns: A Challenging Long-context Instruction-based Exam for LLMs

Shawn Gavin, Tuney Zheng, Jiaheng Liu, Quehry Que, Noah Wang, Jian Yang, Chenchen Zhang, Wenhao Huang, Ge Zhang#

* Equal contribution; # Corresponding author.

This collection includes co-authored benchmarks, companion evaluation datasets, and evaluation methods. Each work has one primary category; topic tags capture related capabilities. Corresponding-author labels refer to Jiaheng Liu.