Machine Learning Ph.D. Thesis Proposal - Jacob Springer

October 15, 2026  3:00PM—4:30PM

Location:
In Person and Virtual - ET - Traffic21 Classroom, Gates Hillman 6501 and Zoom

Speaker:
JACOB SPRINGER, Ph.D. Student, Machine Learning Department, Carnegie Mellon University
https://sprin.xyz/

Learning for Adaptation

Language models must learn from new data while preserving the capabilities and variety of behavior that make them useful. However, fine-tuning can lead to catastrophic forgetting, reduce semantic diversity, and cause learned behavior to generalize beyond its training context in undesirable directions. This thesis studies how to train models for future adaptation, with three objectives: mitigating catastrophic forgetting, preserving semantic diversity, and making generalization of learned behavior possible to understand, predict, and control.

We approach these objectives through optimization and data design during upstream training, including pretraining and mid-training. From an optimization perspective, we study how training methods and hyperparameters influence the ability to learn new tasks while preserving prior capabilities. From a data perspective, we study how the order and structure of training data can make capabilities more robust and retain variety in model completions. We evaluate these approaches through both preservation and learning, with the goal of developing models that can adapt to new objectives without giving up useful properties of their previous training.

We further study how learned behavior appears across different types of prompts and completions, with the goal of understanding where changes induced by fine-tuning will generalize and how this transfer can be controlled. Our proposed work will connect this question to robustness by testing whether upstream optimization and data order can shape how different context types respond to later updates. The goal is desirable selective transfer: new behavior should generalize where it is useful, while other behavior remains stable. Across these directions, we aim to guide the development of base models that retain useful capabilities and semantic diversity while allowing adaptation whose effects we can understand and control.

Thesis Committee:
Aditi Raghunathan (Chair)
Graham Neubig 
Andrew Ilyas
Alex Damian (Massachusetts Institute of Technology)

Additional Information

In Person and Zoom Participation. See announcement. 
 

For More Information:
lyonsmuth@cmu.edu


Add event to Google
Add event to iCal