Academic Homepage

Studying how machines reason, explore, and generate.

I am Zhijian Zhou, a researcher working on reinforcement learning, large language model reasoning, autonomous agents, and diffusion-based generative models.

About

Bio

I am interested in building intelligent systems that can reason, explore, and optimize in complex environments. My recent work spans reinforcement learning with verifiable rewards, language-agent training, structured reasoning for large language models, and diffusion-based generation.

Experience

Industry Experience

  1. Sep 2025 — Present

    Tencent · Project UP (Youtu Lab)

    Research Intern

  2. Apr 2025 — Sep 2025

    Infinite Lightyear (Shanghai) Technology Co., Ltd.

    Research Intern

  3. Jun 2024 — Apr 2025

    Shanghai Academy of AI for Science

    Research Intern

Research

Research Interests

Reinforcement Learning

Diversity preservation, verifiable rewards, intrinsic motivation, and stable optimization for policy learning.

LLM Reasoning & Agents

Structured reasoning, agentic SQL, automated agent generation, and training methods that improve exploration and reliability.

Diffusion Models

Reinforcement-guided diffusion, stable molecule generation, and accelerated molecular conformation generation.

Selected Publications

Publications

# equal contribution  ·  * corresponding author

DisPPO: Quantile-Based Distributional Reinforcement Learning for Large Language Models

Z. Zhou#, L. Li#, X. Zhang#, Z. Liu, Y. Miao, Y. Liu, D. Chen, K. Li, X. Sun, R. Jiang*, X. Tan*, C. Qu*, Y. Qi*

International Conference on Machine Learning (ICML), 2026

SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models

X. Yang, M. Yang, J. Wang, Z. Zhou, Z. Tan, H. Li*

International Conference on Machine Learning (ICML), 2026

Nested Spatio-Temporal Time Series Forecasting

Y. Ai, Y. Zhou, R. Jiang, J. An, C. Qu, Z. Zhou, S. Wang, F. Cao, Z. Xu, F. Shen*, Y. Qi*

International Conference on Machine Learning (ICML), 2026

The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward

L. Li#, Z. Zhou#, J. Hao, J. K. Liu, Y. Miao, W. Pang, X. Tan, W. Chu, Z. Wang, S. Pan, C. Qu*, Y. Qi

International Conference on Learning Representations (ICLR), 2026

Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization

Y. Shi#, Y. Cai#*, S. Cai#, Z. Xu#, L. Chen*, Y. Qin, Z. Zhou, X. Fei, C. Qiu, X. Tan, G. Li, Z. Li, H. Lin, G. Cai, Y. Mao, Y. Wu, K. Li, X. Sun

arXiv preprint arXiv:2512.24615, 2025

Unleashing Flow Policies with Distributional Critics

D. Chen, Y. Liu, Z. Zhou, C. Qu*, Y. Qi*

arXiv preprint arXiv:2509.23087, 2025

Count Counts: Motivating Exploration in LLM Reasoning with Count-based Intrinsic Rewards

X. Zhang, R. Li, Z. Zhou, L. Li, Y. Qin, K. Li, X. Sun, X. Tan*, C. Qu*, Y. Qi*

International Conference on Learning Representations (ICLR), 2026

ChemHTS: Hierarchical Tool Stacking for Enhancing Chemical Agents

Z. Li#, J. Xiao#, B. Zhang, Z. Zhou, Q. He, F. Cao, J. Liang, Y. Qi*

arXiv preprint arXiv:2502.14327, 2025

ChemAmp: Amplified Chemistry Tools via Composable Agents

Z. Li#, P. Chang#, J. Xiao, Z. Zhou, Q. He, J. Liang, F. Cao, X. Yinghui, Y. Qi*

arXiv preprint arXiv:2505.21569, 2025

Guiding Diffusion Models with Reinforcement Learning for Stable Molecule Generation

Z. Zhou, J. An, Z. Liu, Y. Shi, X. Zhang, F. Cao, C. Qu*, Y. Qi*

arXiv preprint arXiv:2508.16521, 2025

DyJR: Preserving Diversity in Reinforcement Learning with Verifiable Rewards via Dynamic Jensen-Shannon Replay

L. Li, Z. Zhou, T. Wang, W. Xu, Z. Huang, W. Chu, Z. Wang, S. Pan, C. Qu, Y. Qi

arXiv preprint arXiv:2603.16157, 2026

SQL-ASTRA: Alleviating Sparse Feedback in Agentic SQL via Column-Set Matching and Trajectory Aggregation

L. Li#, Z. Zhou#, J. Long, P. Liu, W. Xu, Z. Wang, S. Pan*, C. Qu*

Annual Meeting of the Association for Computational Linguistics (ACL), 2026

Constraints-Guided Diffusion Reasoner for Neuro-Symbolic Learning

X. Zhang, Z. Zhou, W. Xu, Y. Miao, C. Qu*, Y. Qi*

AAAI Conference on Artificial Intelligence (AAAI), pp. 28446–28454, 2026

Equivariant Asynchronous Diffusion: An Adaptive Denoising Schedule for Accelerated Molecular Conformation Generation

J. An, N.-N. Zhang, Y.-F. Shi, Z. Zhou, C. Qu, F. Cao, Y. Qi

arXiv preprint arXiv:2603.10093, 2026

Determining Blockchain Transaction Timing and Fee with Observable Mempools

Q. Bai, Y. Xu, Z. Zhou, X. Wang

arXiv preprint arXiv:2512.21923, 2025

From Implicit Exploration to Structured Reasoning: Guideline and Refinement for LLMs

J. Chen, Z. Wang, M. Zou, Z. Li, Z. Zhou, S. Wang, Z. Xu

Findings of the Association for Computational Linguistics: EMNLP, 2025

SDPO: Importance-Sampled Direct Preference Optimization for Stable Diffusion Training

X. Yang, Z. Tan, J. Wang, Z. Zhou, H. Li*

arXiv preprint arXiv:2505.21893, 2025

Contact

Get in Touch

For collaboration, discussion, or paper-related questions, please reach out by email or via academic profiles.