← Evaluation overview 6 research contributions
Instructions & Long Context Instruction following, multi-turn dialogue, role consistency, and long-context understanding and generation.
Benchmarks & related work 6 works
Search by name or capability, or narrow by publication status and authorship.
Clear filters
Instruction following & dialogue
Inverse IFEval
Tests whether models follow real instructions that conflict with learned training conventions.
Instruction following Robustness
ICLR 2026 Corresponding author
Paper & authors Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
Qinyan Zhang* , Xinping Lei* , Ruijie Miao* , Yu Fu, Haojie Fan, Le Chang, Jiafan Hou, Dingling Zhang, Zhongfei Hou, Ziqiang Yang, Changxin Pu, Fei Hu, Jingkai Liu, Xinjie Chen, Jianpeng Jiao, Jiaheng Liu# , Tong Yang# , Zaiyuan Wang# , Ge Zhang# , Wenhao Huang
* Equal contribution; # Corresponding author.
Instruction following & dialogue
RoleAgentBench
Evaluates role-playing agents built from scripts through interactive tasks.
Role playing Multi-turn
NeurIPS 2024 (Datasets and Benchmarks)
Paper & authors RoleAgent: Building, Interacting, and Benchmarking High-quality Role-Playing Agents from Scripts
Jiaheng Liu* , Zehao Ni* , Haoran Que* , Tao Sun, Zekun Wang, Jian Yang, Jiakai Wang, Hongcheng Guo, Zhongyuan Peng, Ge Zhang, Jiayi Tian, Xingyuan Bu, Ke Xu, Wenge Rong, Junran Peng# , Zhaoxiang Zhang
* Equal contribution; # Corresponding author.
Instruction following & dialogue
MT-Bench-101
Provides fine-grained evaluation of multi-turn dialogue abilities.
Dialogue Multi-turn
ACL 2024
Paper & authors MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
Ge Bai, Jie Liu, Xingyuan Bu# , Yancheng He, Jiaheng Liu , Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, Wanli Ouyang
* Equal contribution; # Corresponding author.
Instruction following & dialogue
RoleBench
Evaluates role-playing abilities and consistency in language models.
Role playing Dialogue
Findings of ACL 2024 Corresponding author
Paper & authors RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
Noah Wang, Z.Y. Peng, Haoran Que* , Jiaheng Liu# , Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, Man Zhang, Zhaoxiang Zhang# , Wanli Ouyang, Ke Xu, Wenhao Huang, Jie Fu, Junran Peng
* Equal contribution; # Corresponding author.
Long-context understanding & generation
HelloBench
Evaluates the ability to generate long, coherent text.
Long-form generation Writing
arXiv preprint, 2024
Paper & authors HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
Haoran Que, Feiyu Duan, Liqun He, Yutao Mou, Wangchunshu Zhou, Jiaheng Liu , Wenge Rong, Zekun Moore Wang, Jian Yang, Ge Zhang, Junran Peng, Zhaoxiang Zhang, Songyang Zhang, Kai Chen
* Equal contribution; # Corresponding author.
Long-context understanding & generation
LongIns
Tests instruction following and multitask reasoning over long contexts.
Long context Instruction following
arXiv preprint, 2024
Paper & authors LongIns: A Challenging Long-context Instruction-based Exam for LLMs
Shawn Gavin, Tuney Zheng, Jiaheng Liu , Quehry Que, Noah Wang, Jian Yang, Chenchen Zhang, Wenhao Huang, Ge Zhang#
* Equal contribution; # Corresponding author.
No matching work Try a broader keyword or clear the filters.
Show more
This collection includes co-authored benchmarks, companion evaluation datasets, and evaluation methods. Each work has one primary category; topic tags capture related capabilities. Corresponding-author labels refer to Jiaheng Liu.
Jiaheng Liu · Nanjing University Updated September 2026