← Evaluation overview

9 research contributions

Safety & Factuality

Factual knowledge, multimodal hallucination, jailbreak resistance, over-refusal, and human-centered behavior.

Benchmarks & related work

9 works

Search by name or capability, or narrow by publication status and authorship.

Factuality & hallucination

OmniHalluc-L

Evaluates long-form omni-modal hallucination using counterfactual and modality-perturbation analysis.

HallucinationAudio-visual
EMNLP 2026 · AcceptedCorresponding author
Paper & authors

OmniHalluc-L: Counterfactual Benchmarking and Modality-Perturbation Calibration for Long-Form Omni Hallucination

Zixuan Dong, Jiafu Tang, Zhide Lei, Zhe Cao, Zijie Zhang, Yanghai Wang, Shihao Li, Xiaodong Wang, Baoyun Peng, Jiaheng Liu#

* Equal contribution; # Corresponding author.

Factuality & hallucination

MSQA

Evaluates multilingual and multicultural factual knowledge with natively sourced questions.

FactualityMultilingual
Findings of EMNLP 2026 · Accepted
Paper & authors

MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

Yukai Huang, Xianru Chen, Mingxiang Chen, Xinping Lei, Fangbing Deng, Jin Chen, Jiaheng Liu, Ge Zhang, Wenhao Huang

* Equal contribution; # Corresponding author.

Safety & human alignment

USB

Provides unified evaluation of multimodal safety and over-refusal.

Multimodal safetyOver-refusal
ACL 2026
Paper & authors

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models

Baolin Zheng#, Guanlin Chen, Qingyang Teng, Hongqiong Zhong, Yingshui Tan, Zhendong Liu, Weixun Wang, Jiaheng Liu, Jian Yang, Huiyun Jing, Jincheng Wei, Wenbo Su, Xiaoyong Zhu, Bo Zheng, Kaifu Zhang

* Equal contribution; # Corresponding author.

Safety & human alignment

SafeDialBench

Evaluates safety in multi-turn dialogues with diverse jailbreak attacks.

Jailbreak resistanceMulti-turn
ICLR 2026
Paper & authors

SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks

Hongye Cao, Sijia Jing, Yanming Wang, Ziyue Peng, Zhixin Bai, Zhe Cao, Meng Fang, Fan Feng, Jiaheng Liu, Boyan Wang, Tianpei Yang, Jing Huo#, Yang Gao, Fanyu Meng#, Xi Yang, Chao Deng, Junlan Feng

* Equal contribution; # Corresponding author.

Safety & human alignmentCompanion dataset

FigStep-benign

Evaluates over-refusal on benign multimodal requests constructed in the DREAM study.

Over-refusalMultimodal safety
NAACL 2025
Paper & authors

DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models

Jianyu Liu, Hangyu Guo, Ranjie Duan, Xingyuan Bu#, Yancheng He, Shilong Li, Hui Huang, Jiaheng Liu, Yucheng Wang, Chenchen Jing, Xingwei Qu, Xiao Zhang, Pei Wang, Yanan Wu, Jihao Gu, Yangguang Li, Jianke Zhu

* Equal contribution; # Corresponding author.

Safety & human alignment

Chinese SafetyQA

Tests the factual accuracy of safety-related knowledge.

Safety knowledgeFactuality
ACL 2025
Paper & authors

Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models

Yingshui Tan#, Boren Zheng, Baihui Zheng, Kerui Cao, Huiyun Jing, Jincheng Wei, Jiaheng Liu, Yancheng He, Wenbo Su, Xiaoyong Zhu, Bo Zheng, Kaifu Zhang

* Equal contribution; # Corresponding author.

Factuality & hallucination

Chinese SimpleQA

Evaluates short-form factual question answering in Chinese.

FactualityChinese
ACL 2025Corresponding author
Paper & authors

Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models

Yancheng He, Shilong Li, Jiaheng Liu#, Yingshui Tan, Weixun Wang, Hui Huang, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Zhuoran Lin, Dekai Sun, Zhicheng Zheng, Wenbo Su, Bo Zheng

* Equal contribution; # Corresponding author.

Safety & human alignment

BSA Bench(Beyond Safe Answers)

Tests whether reasoning models recognize risk beyond producing superficially safe answers.

Risk awarenessSafety reasoning
arXiv preprint, 2025
Paper & authors

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models

Baihui Zheng, Boren Zheng, Kerui Cao, Yingshui Tan#, Zhendong Liu, Weixun Wang, Jiaheng Liu, Jian Yang, Wenbo Su, Xiaoyong Zhu, Bo Zheng, Kaifu Zhang

* Equal contribution; # Corresponding author.

Safety & human alignment

GIEBench

Evaluates empathy through the lens of group identity.

EmpathyHuman alignment
arXiv preprint, 2024
Paper & authors

GIEBench: Towards Holistic Evaluation of Group Identity-based Empathy for Large Language Models

Leyan Wang, Yonggang Jin, Tianhao Shen, Tianyu Zheng, Xinrun Du, Chenchen Zhang, Wenhao Huang, Jiaheng Liu, Shi Wang, Ge Zhang#, Liuyu Xiang#, Zhaofeng He

* Equal contribution; # Corresponding author.

This collection includes co-authored benchmarks, companion evaluation datasets, and evaluation methods. Each work has one primary category; topic tags capture related capabilities. Corresponding-author labels refer to Jiaheng Liu.