Search papers, labs, and topics across Lattice.
This study leverages AstroPT, a transformer model trained on galaxy images, to investigate the emergence of concepts during training, providing a unique calibration testbed for understanding language models. By probing representations across various checkpoints and model configurations, the authors reveal that galaxy properties emerge in a fixed order that aligns with their known difficulty, with simpler properties being decodable earlier in training. This invariant ordering across training objectives and model capacities highlights the potential of astronomical data to enhance mechanistic interpretability in LLMs.
AstroPT reveals that galaxy properties emerge in a predictable order during training, offering a new lens to understand concept emergence in LLMs.
Interpretability research increasingly asks when concepts emerge during training and whether linear probes recover real structure, but in language models these claims are hard to validate because language offers little ground-truth ordering of concepts or relationships among them. We propose the use of astronomical ground truth through AstroPT, a transformer trained on millions of galaxy images, as a calibration testbed. AstroPT is an LLM-like model trained within a domain where the difficulty ordering of concepts and the relations among them are known in advance. Probing frozen representations across checkpoints, layers, model sizes, and objective choices, we find that galaxy properties emerge in a fixed order that tracks their known difficulty---quantities written almost directly into the pixels (band magnitude) become decodable early in training and shallow in the network, while multiband/spectra based and inferred quantities (such as redshift and specific star formation rate) emerge later and deeper. This order is invariant to our tested training objectives, and scales in magnitude but not in sequence with capacity. Our linear probe directions further recover the known physical structure among galaxy properties. Our findings suggest that astronomy offers a controlled sandbox for calibrating mechanistic interpretability methods we otherwise apply to LLMs blind.