Introduction

Ning Lu 卢宁

Learning from the world,
guided by value.

Ph.D. in CSE · Hong Kong University of Science and Technology
I received my Ph.D. in CSE from the Hong Kong University of Science and Technology (HKUST), where I was supervised by Prof. Cunsheng Ding, Prof. Qi Wang, and Prof. Ke Tang and focused on trustworthy language models. Previously I was an intern researcher at Bytedance-Seed, where I focus on the Agentic Reinforcement Learning. Now I am focusing on Agentic Reinforcement Learning and World Modeling for LLM.

Research Interest:
- LLM RL for Agent and Reasoning: PaW, HIVE, AHDAgent, Is PRM Necessary?
- LLM Alignment & RedTeaming: SafeDelta, SICO

Background

Journey so far.

Education

  1. Hong Kong University of Science and Technology

    PhD in Computer Science and Engineering

  2. Southern University of Science and Technology

    Ranked 916th in Shandong Province Gaokao
    BSc in Computer Science and Engineering

Selected honors

Outstanding Entrance Scholarship
Outstanding Graduate of the Computer Science Department

Experience

  1. Tencent, Lightspeed Studio

    Research Intern (Project UP), focusing on LLM agents and world model.

  2. ByteDance, Seed

    • Contributed to the development of Doubao Voice Agent.
    • Contributed to the Seed-2.0-lite-0428 RLVR post-training.

  3. ByteDance, Data-Douyin

    Research Intern, focusing on LLM RL.

  4. Huawei, 2012 Lab

    Research Intern, focusing on fault-torlerant large language model pre-training

Updates

What’s new.

We introduce PaW, which co-trains policy learning and world modeling from the same RL rollouts. Read the paper
I successfully defended my Ph.D. thesis! Grateful to my advisors, committee, collaborators, and everyone who supported me along the way.
We introduce AHD Agent, the first LLM agent framework for automatic algorithm design. Read the paper
Excited to share that Seed-2.0-lite-0428, a project I contributed to, is now released, with Gemini-level multimodal understanding.
We propose HIVE, which accelerates LLM reinforcement learning through prompt-entropy-based data selection. Read the paper

Selected research

View all
Status Under review.

Policy and World Modeling Co-Training for Language Agents

Ning Lu*, Baijiong Lin*, Shengcai Liu, Jiahao Wu, Haoze Lv, Yanbin Wei, Lingting Zhu, Shengju Qian, Xin Wang, Ying-Cong Chen, Qi Wang, Ke Tang (* equal contribution)

Status Under review.

Train at the Moving Edge: Efficient RL for Large Reasoning Models via Rollout Selection

Jiahao Wu*, Ning Lu*, Shengcai Liu, Kun Wang, Yanting Yang, Li Qing, Ke Tang (* equal contribution)

Published at Conference on Neural Information Processing Systems (NeurIPS)

Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs

Zhangying Feng*, Qianglong Chen*, Ning Lu, Yongqian Li, Siqi Cheng, Shuangmu Peng, Duyu Tang, Shengcai Liu, Zhirui Zhang (* equal contribution)

Published at International Conference on Machine Learning (ICML)

Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets

Ning Lu, Shengcai Liu, Jiahao Wu, Weiyu Chen, Zhirui Zhang, Yew-Soon Ong, Qi Wang, Ke Tang

Published at Transactions on Machine Learning Research (TMLR)

Large Language Models can be Guided to Evade AI-generated Text Detection

Ning Lu, Shengcai Liu, Rui He, Qi Wang, Yew-Soon Ong, Ke Tang