Rather than asking only what a model predicts for a particular input, I am interested in how its behavior changes under systematic transformations of that input. The transformations a model ignores can reveal invariances it has learned, while transformations that produce strong changes can expose sensitivities in its learned representations.
Can transformations be used as questions we ask a model?
We are studying whether a model's response to carefully chosen transformations can provide an interpretable characterization of what the model considers relevant. This shifts interpretability from examining isolated activations toward probing behavior across structured families of inputs.
We are investigating transformation-based measurements together with neural tangent kernel ideas, with the goal of connecting changes in model behavior to a geometric characterization of sensitivity and invariance.
I am collaborating on the project and working on the conceptual and experimental framework for measuring model invariance and sensitivity across transformations.
This is ongoing work. The research questions and visualizations will evolve as we test which transformations provide useful and reliable probes of model behavior.