I’m a PhD student in computer science at Princeton, advised by Sanjeev Arora.

My research focuses on how machine learning systems generalize beyond their training data. I’m particularly interested in self-improving models, synthetic data, and automated research.

Latest

Writing

All writing →

Research

I combine controlled synthetic problems, empirical scaling, and theory to study what models learn and what they can discover next.

Generalization beyond the training distribution

Understanding how learning systems extrapolate across length, hardness, and composition—and what their successes reveal about learned algorithms.

Self-improving learning systems

Studying synthetic data, weak-to-strong scaling, and feedback loops that let models create increasingly useful supervision.

Automating the research loop

Finding research problems where hypothesis generation becomes structured search, then designing agents that can explore that space in parallel.

Selected papers

Full CV →
2025NeurIPS · Spotlight

Extrapolation by Association: Length Generalization Transfer in Transformers

Ziyang Cai, Nayoung Lee, Avi Schwarzschild, Samet Oymak, Dimitris Papailiopoulos

2025ICML

Self-Improving Models Overcome Length and Hardness Generalization via Weak-to-Strong Scaling

Ziyang Cai*, Nayoung Lee*, Avi Schwarzschild, Kangwook Lee, Dimitris Papailiopoulos

2025ICML · Spotlight

Everything Everywhere All at Once: LLMs Can In-Context Learn Multiple Tasks in Superposition

Zheyang Xiong, Ziyang Cai, John Cooper, et al.

2022NeurIPS

Delving into Out-of-Distribution Detection with Vision-Language Representations

Yifei Ming, Ziyang Cai, Jiuxiang Gu, Yiyou Sun, Wei Li, Yixuan Li

Background

Before Princeton, I studied at UW–Madison and built automated synthetic-data pipelines for coding agents at Microsoft Research. I was previously advised by Dimitris Papailiopoulos, Junjie Hu, and Sharon Li.