Ning Lu · AI Researcher

I teach language agents
to think, act, and imagine.

My research explores agentic reinforcement learning and world models— building AI systems that learn through interaction, reason over long horizons, and remain aligned along the way.

A research philosophy

Intelligence is not only about producing an answer. It is about understanding what happens next.

01

Agents

Language models that use tools, learn from environments, and improve through experience.

02

Worlds

Models that anticipate consequences and turn imagination into better decisions.

03

Alignment

Methods that keep capable models safe, useful, and faithful as they adapt.

Selected research

Ideas, made concrete.

Recent work across agent learning, reasoning, world modeling, and language-model safety.

2026

Under review.

Policy and World Modeling Co-Training for Language Agents

The first policy and world-modeling co-training RL framework for LLM agents.

2026

Under review.

Train at the Moving Edge: Efficient RL for Large Reasoning Models via Rollout Selection

The first online policy-verified data selection framework for efficient RL training.

2025

Conference on Neural Information Processing Systems (NeurIPS)

Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs

Unifying problem solving and solution-process judgment.

2025

International Conference on Machine Learning (ICML)

Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets

The first safety-aware post-fine-tuning defense method for LLM alignment.

2024

Transactions on Machine Learning Research (TMLR)

Large Language Models can be Guided to Evade AI-generated Text Detection

Showing that LLMs themselves can evade AI detectors with fine-grained prompting.

About Ning

At the boundary of
reasoning and action.

I received my Ph.D. in CSE from the Hong Kong University of Science and Technology (HKUST), where I was supervised by Prof. Cunsheng Ding, Prof. Qi Wang, and Prof. Ke Tang and focused on trustworthy language models. Previously I was an intern researcher at Bytedance-Seed, where I focus on the Agentic Reinforcement Learning. Now I am focusing on Agentic Reinforcement Learning and World Modeling for LLM.

Research Interest:
- LLM RL for Agent and Reasoning: PaW, HIVE, AHDAgent, Is PRM Necessary?
- LLM Alignment & RedTeaming: SafeDelta, SICO

DegreePh.D. in CSE
Alma materHKUST
Current lensAgents × Worlds

The next question

What if an agent could
learn the world as it acts?

Let’s find out together