Spotting Problems for Auto Research
How to recognize the parts of research that become combinatorial search—and where parallel research agents can make real progress.
My research focuses on how machine learning systems generalize beyond their training data. I’m particularly interested in self-improving models, synthetic data, and automated research.
Latest
How to recognize the parts of research that become combinatorial search—and where parallel research agents can make real progress.
An empirical study of long-horizon coding agents, optimization plateaus, and context-reset interventions.
A research note exploring the sample complexity of length and composition generalization in machine learning models.
I combine controlled synthetic problems, empirical scaling, and theory to study what models learn and what they can discover next.
Understanding how learning systems extrapolate across length, hardness, and composition—and what their successes reveal about learned algorithms.
Studying synthetic data, weak-to-strong scaling, and feedback loops that let models create increasingly useful supervision.
Finding research problems where hypothesis generation becomes structured search, then designing agents that can explore that space in parallel.
Ziyang Cai, Nayoung Lee, Avi Schwarzschild, Samet Oymak, Dimitris Papailiopoulos
Ziyang Cai*, Nayoung Lee*, Avi Schwarzschild, Kangwook Lee, Dimitris Papailiopoulos
Zheyang Xiong, Ziyang Cai, John Cooper, et al.
Yifei Ming, Ziyang Cai, Jiuxiang Gu, Yiyou Sun, Wei Li, Yixuan Li
Before Princeton, I studied at UW–Madison and built automated synthetic-data pipelines for coding agents at Microsoft Research. I was previously advised by Dimitris Papailiopoulos, Junjie Hu, and Sharon Li.