Publications
My research covers foundation models, code intelligence, reinforcement learning, agents, evaluation, and computer vision. See also Google Scholar.
Papers are grouped by publication or conference year; preprints appear after peer-reviewed papers within each year. * Equal contribution; # Corresponding author. Authorship marks follow the paper or author-confirmed information.
209 publications
No matching publications. Try another keyword or clear the filters.
2026
- OmniHalluc-L: Counterfactual Benchmarking and Modality-Perturbation Calibration for Long-Form Omni Hallucination
- MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
- AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning
- TVIR: Building Deep Research Agents Towards Text-Visual Interleaved Report Generation
- WebCompass: Evaluating Multimodal Code Agents in Web Development from Vision, Instruction, and Interaction to Implementation
- DR³-Eval: Towards Realistic and Reproducible Deep Research Evaluation
- ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
- HGP: An on-device personalized agent memory via hybrid graph storage
- ClawBench: Can AI Agents Complete Everyday Online Tasks?
- MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
- MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark
- DeepResearch RewardBench: Evaluating Reward Models for Deep Research Report
- WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints
- OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
- T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
- SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
- VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
- From Diagrams to Code: Multilingual Programming with Visual Design
- Evolving Rollouts: Harnessing Historical Experience for Web Agent Evolution in Reinforcement Learning
- How Far Can LLM Agents Reason with Tables? Benchmarking Multi-Turn Agentic Table Question Answering in the Wild
- Enhancing Multilingual Reasoning via Steerable Model Merging
- When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors
- CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs
- Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning
- Multi-Docker-Eval: A ‘Shovel of the Gold Rush’ Benchmark on Automatic Environment Building for Software Engineering
- MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
- ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
- M3TQA: Massively Multilingual Multitask Table Question Answering
- CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization
- USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
- Table-R1: Region-based Reinforcement Learning for Table Understanding
- COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values
- MdEval: Massively Multilingual Code Debugging
- MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models
- MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QA
- Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
- IF-VidCap: Can Video Caption Models Follow Instructions?
- ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution
- Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
- DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning
- Tricks or Traps? A Deep Dive into RL for LLM Reasoning
- TaskCraft: Automated Generation of Agentic Tasks
- ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
- IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
- YuE: Scaling Open Foundation Models for Long-Form Music Generation
- SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
- Long-form RewardBench: Evaluating Reward Models for Long-form Generation
- Think-J: Learning to Think for Generative LLM-as-a-Judge
- HappyWorld-Bench
- GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions
- Rethinking Multi-Agent Collaboration: When More Is Less
- xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
- Aspire: Can Models Self-Evolve from Vague Goals?
- S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
- REER-PT: Reverse-Engineered Reasoning for Perplexity-Guided Pre-training Data Augmentation
- Procedura: Agentic 3D Modeling with Procedural Control
- Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training
- StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
- LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation
- AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
- ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
- Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
- Can Machines Really See Objects in Images? A Study Based on Syntactic Distance and Visual Self-Referential Instances
- CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents
- P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning
- Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
- OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning
- CoVEBench: Can Video Editing Models Handle Complex Instructions?
- Knowledge Index of Noah's Ark
- Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories
- OProver: A Unified Framework for Agentic Formal Theorem Proving
- Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution
- AlphaCrafter: Harnessing Multi-Agent Workflows for Cross-Sectional Quantitative Trading
- CodeTracer: Towards Traceable Agent States
- Understanding by Reconstruction: Reversing the Software Development Process for LLM Pretraining
- Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization
- EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
- Vibe AIGC: A New Paradigm for Content Generation via Agentic Orchestration
- The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
2025
- OAgents: An Empirical Study of Building Effective Agents
- AIR: Complex Instruction Generation via Automatic Iterative Refinement
- MIO: A Foundation Model on Multimodal Tokens
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
- Flow-GRPO: Training Flow Matching Models via Online RL
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- OmniBench: Towards The Future of Universal Omni-Language Models
- Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models
- DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models
- Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
- VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
- See the World, Discover Knowledge: A Chinese Factuality Evaluation for Large Vision Language Models
- MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training
- Quantification of Large Language Model Distillation
- ProgCo: Program Helps Self-Correction of Large Language Models
- Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models
- Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
- OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
- M2RC-EVAL: Massively Multilingual Repository-level Code Completion Evaluation
- 2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
- Can MLLMs Understand the Deep Implication Behind Chinese Images?
- PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment
- LIME: Less Is More for MLLM Evaluation
- MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
- KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks
- McEval: Massively Multilingual Code Evaluation
- MuPT: A Generative Symbolic Music Pretrained Transformer
- TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
- xCoT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought Reasoning
- Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
- SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models
- Molecular Graph Contrastive Learning with Line Graph
- PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
- Encyclo-K: Evaluating LLMs with Dynamically Composed Knowledge Statements
- AutoMV: An Automatic Multi-Agent System for Music Video Generation
- ViDiC: Video Difference Captioning
- How Far Are We from Genuinely Useful Deep Research Agents?
- AI Deception: Risks, Dynamics, and Controls
- From Code Foundation Models to Agents and Applications: A Comprehensive Survey and Practical Guide to Code Intelligence
- Scaling Latent Reasoning via Looped Language Models
- KAT-Coder Technical Report
- HiPO: Hybrid Policy Optimization for Dynamic Reasoning in LLMs
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
- Efficient Agents: Building Effective Agents While Reducing Cost
- IFEvalCode: Controlled Code Generation
- KAT-V1: Kwai-AutoThink Technical Report
- Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving
- A Survey on Latent Reasoning
- SPEAR: Structured Pruning for Spiking Neural Networks via Synaptic Operation Estimation and Reinforcement Learning
- Scaling Test-time Compute for LLM Agents
- Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
- Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
- A Comprehensive Survey on Long Context Language Modeling
- Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation
- CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models
- Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs
- Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models
- Aligning Instruction Tuning with Pre-training
2024
- GraphReader: Building Graph-based Agent to Enhance Long-Context Abilities of Large Language Models
- DDK: Distilling Domain Knowledge for Efficient Large Language Models
- II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
- D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
- RoleAgent: Building, Interacting, and Benchmarking High-quality Role-Playing Agents from Scripts
- Compressing Large Language Models by Joint Sparsification and Quantization
- UniCoder: Scaling Code Large Language Model via Universal Code
- Towards Real-world Scenario: Imbalanced New Intent Discovery
- MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
- ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models
- Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
- E2-LLM: Efficient and Extreme Length Extension of Large Language Models
- RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
- OWL: A Large Language Model for IT Operations
- LogFormer: A Pre-train and Tuning Pipeline for Log Anomaly Detection
- PTSBench: A Comprehensive Post-Training Sparsity Benchmark Towards Algorithms and Models
- NC-NCD: Novel Class Discovery for Node Classification
- Segment, Lift and Fit: Automatic 3D Shape Labeling from 2D Prompts
- MaterialSeg3D: Segmenting Dense Materials from 2D Priors for 3D Assets
- Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model
- m3P: Towards Multimodal Multilingual Translation with Multimodal Prompt
- MT4CrossOIE: Multi-stage Tuning for Cross-lingual Open Information Extraction
- LTA-PCS: Learnable Task-Agnostic Point Cloud Sampling
- VRDistill: Vote Refinement Distillation for Efficient Indoor 3D Object Detection
- FullStack Bench: Evaluating LLMs as Full Stack Coders
- PSA-VLM: Enhancing Vision-Language Model Safety through Progressive Concept-Bottleneck-Driven Alignment
- AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions
- Aligning CodeLLMs with Direct Preference Optimization
- ING-VP: MLLMs cannot Play Easy Vision-based Games Yet
- HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
- FuzzCoder: Byte-level Fuzzing Test via Large Language Model
- I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm
- LongIns: A Challenging Long-context Instruction-based Exam for LLMs
- GIEBench: Towards Holistic Evaluation of Group Identity-based Empathy for Large Language Models
- Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level
- R2C2-Coder: Enhancing and Benchmarking Real-world Repository-level Code Completion Abilities of Code Large Language Models
- MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series
- The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability Analysis
- MLAD: A Unified Model for Multi-system Log Anomaly Detection
- LD-T3D: A Large-scale and Diverse Benchmark for Text-based 3D Model Retrieval
2023
- M2C: Towards Automatic Multimodal Manga Complement
- Adaptive Contrastive Knowledge Distillation for BERT Compression
- GripRank: Bridging the Gap between Retrieval and Generation via the Generative Knowledge Improved Passage Ranking
- GD-MAE: Generative Decoder for MAE Pre-training on LiDAR Point Clouds
- LogLG: Weakly Supervised Log Anomaly Detection via Log-Event Graph Construction
- A Unified Efficient Deep Image Compression Framework and Its Application on Human-centric Task
- ICD-Face: Intra-class Compactness Distillation for Face Recognition
- 3D-QueryIS: A Query-based Framework for 3D Instance Segmentation
2022
- LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine Translation
- Cross-Lingual Cross-Modal Consolidation for Effective Multilingual Video Corpus Moment Retrieval
- AnchorFace: Boosting TAR@FAR for Practical Face Recognition
- Deep 3D Vessel Segmentation based on Cross Transformer Network
- Computer-aided Tuberculosis Diagnosis with Attribute Reasoning Assistance
- CoupleFace: Relation Matters for Face Recognition Distillation
- 3D-Pruning: A Model Compression Framework for Efficient 3D Action Recognition
- APSNet: Towards Adaptive Point Sampling for Efficient 3D Action Recognition
- GeometryMotion-Transformer: A Strong Backbone for 3D Action Recognition
- OneFace: One Threshold for All
2021
- DAM: Discrepancy Alignment Metric for Face Recognition
- GeometryMotion-Net: A Strong Two-stream Baseline for 3D Action Recognition
- JointPruning: Pruning Networks along Multiple Dimensions for Efficient Point Cloud Processing
2020
- Learning to Auto Weight: Entirely Data-driven and Highly Efficient Weighting Framework
- Block Proposal Neural Architecture Search