← Research

INTERPRETABILITY / TRANSFORMATIONS & NTK

Transformation-Based Interpretability

Understanding Model Behavior Through Transformations and Neural Tangent Kernels

Supervised by Seyed-Mohsen Moosavi-Dezfooli, Mahed Abroshan, and Mohammad Sabokrou

Can the transformations a model is invariant to—or sensitive to—reveal the structure of what it has learned?

Ongoing research

01 · Motivation

Rather than asking only what a model predicts for a particular input, I am interested in how its behavior changes under systematic transformations of that input. The transformations a model ignores can reveal invariances it has learned, while transformations that produce strong changes can expose sensitivities in its learned representations.

02 · Research question

Can transformations be used as questions we ask a model?

We are studying whether a model's response to carefully chosen transformations can provide an interpretable characterization of what the model considers relevant. This shifts interpretability from examining isolated activations toward probing behavior across structured families of inputs.

03 · Current approach

We are investigating transformation-based measurements together with neural tangent kernel ideas, with the goal of connecting changes in model behavior to a geometric characterization of sensitivity and invariance.

My role

I am collaborating on the project and working on the conceptual and experimental framework for measuring model invariance and sensitivity across transformations.

04 · Status

This is ongoing work. The research questions and visualizations will evolve as we test which transformations provide useful and reliable probes of model behavior.