Machine learning and AI-assisted inference
From neural networks that weigh hidden planets to AI-accelerated models of disk kinematics and Gaia images
Many planet-forming disks have now been imaged at high resolution, and rings and gaps are ubiquitous among them, with spirals in some. If even some of these are carved by planets, they reveal a population of young planets that no other technique can currently detect. Turning this growing sample into planet demographics requires inference that is fast, uses the full information in an image, and propagates physical uncertainties honestly. My goal is not to replace physics with a black box, but to use machine learning to make rich, multi-physics models tractable and to confront them directly with data.
PGNets: weighing planets directly from disk images
For DSHARP I inferred planet masses from disk gaps using a large grid of planet–disk simulations and fitted scaling relations between gap width and depth and planet mass (Zhang et al., 2018). That approach works, but it collapses each image to a few azimuthally averaged numbers and throws away asymmetric features, while fine-tuned simulations of individual disks are far too slow for large samples.
In 2021 we developed PGNets (Planet Gap neural Networks), among the first applications of convolutional neural networks to infer planet masses directly from disk images (Zhang et al., 2022). Trained on synthetic ALMA continuum images from hydrodynamical simulations with dust and radiative transfer, PGNets:
- classifies planet mass across five classes from 11 Earth masses to 3 Jupiter masses with up to 92% accuracy (ResNet; 89% for a VGG-like network);
- regresses planet mass and disk turbulence simultaneously, with 1σ uncertainties of 0.16 dex in planet mass and 0.23 dex in the viscosity parameter α;
- rediscovers physics on its own: without being told, the networks recover the known degeneracy α ∝ Mp3 between planet mass and disk viscosity;
- looks at the right features: Grad-CAM activation maps confirm that the networks focus on the gaps when making predictions.
Once trained, PGNets returns a prediction instantly from any image and treats shallow, narrow gaps with the same effort as deep ones, making it well suited to large disk samples and to narrowing the parameter space before detailed simulations of individual disks. The code is public on GitHub, alongside the DSHARP fitting method.
Planet–disk inference also enables population-level comparisons: for compact disks in Taurus, we compared the inferred young planets with mature exoplanet populations (Zhang et al., 2023).
Where this is going
Which physics made this structure? Rings, spirals, and asymmetries can be produced by planets, but also by binaries, gravitational instability, magnetic processes, infall, or shadows. Current machine-learning pipelines often assume every substructure is planetary and neglect thermodynamics. I plan to develop hierarchical, simulation-based inference that first identifies the plausible physical model class and only then infers parameters within it, with uncertainties propagated throughout. Because my simulations include multi-frequency radiation transport, multiple dust species, and dust–gas coupling, they provide training sets far more realistic than those used in current pipelines.
Accelerating disk kinematics with modern GPU/CPU architectures. Discminer extracts disk geometry, temperature, and velocity structure from ALMA channel maps. As I add the non-axisymmetric temperature structures predicted by my radiation-hydrodynamical simulations, forward modeling and posterior exploration become very expensive. I plan to build machine-learning surrogate models and differentiable emulators that run efficiently on modern GPU and CPU architectures, to accelerate both the forward model and the likelihood. This will make Bayesian inference feasible for azimuthal temperature variations, multiple emitting surfaces, and non-Keplerian flows, tested against full data cubes rather than a few summary statistics.
Shadow motions from Gaia DR4. Gaia Data Release 4, expected in December 2026, will include a new residual image product that reveals scattered light from protoplanetary disks around bright stars. The residual image is built by stacking Gaia’s many scans of each star over the mission. A single scan has too little signal to show a disk on its own. But with a model of how shadows and the disk move over time, the individual scans can be fit together to infer how the disk varies. I am building software to do this for a large, homogeneous sample of resolved disks. The goal is to separate persistent disk structure from time-variable illumination, and to track shadow motions across a population rather than one disk at a time. Because shadows are cast by inner disks, their long-term variability measures how inner disks precess, which in turn constrains young planets too close to the star to be resolved directly.
From young planets to mature exoplanets. Roman’s microlensing survey will probe cold planets at the orbital separations where ALMA disks show gaps. By propagating physical-model uncertainty and observational selection effects, I aim to test how growth and migration transform the young planet population into mature planetary architectures.