Background. While Foundation Models like CLIP achieve unprecedented generalization, adapting them to specific downstream tasks via traditional fine-tuning is computationally expensive and fragments their generalized knowledge into isolated models. Task Arithmetic offers a powerful alternative by treating fine-tuning as a directional weight-space shift, or task vector. Manipulating these vectors through simple algebra enables zero-shot multi-task composition and domain adaptation without the need for retraining. Problem. A highly ambitious application of this framework is the zero-shot task analogy, which attempts to transfer learned capabilities from a source domain to a completely unobserved target domain. However, applying simple algebraic summation in a highly non-linear weight space inevitably causes high level of interference. This geometric misalignment entangles the analogical vector with source-specific artifacts. Although recent literature attempts to mitigate this by performing arithmetic in linearized regimes (e.g., the tangent space of the Neural Tangent Kernel), this thesis demonstrates that strict linearization artificially bottlenecks the model's expressivity, rendering it incapable of handling severe structural domain shifts. Furthermore, we identify the unconstrained subtraction of the source task vector tau_A as the primary cause of performance degradation, as it inadvertently erodes foundational, low-level feature extractors. Methodology. This thesis proposes solving the interference directly within the expressive non-linear space. To achieve this, we adapt GradFix, a gradient-guided masking mechanism rooted in model rebasin, to the novel context of zero-shot task analogies. Assuming domain shifts modify largely disjoint parameter subsets, we utilize a minimal target support set (few-shot oracle) to compute a binary mask via gradient sign-consensus. This discrete directional mask is applied asymmetrically (exclusively to the subtractive term) filtering out geometrically misaligned weights and safely decoupling domain-specific noise from transferable semantic knowledge prior to the algebraic projection. Results. Extensive empirical evaluations on the DomainNet benchmark validate both our theoretical diagnosis and algorithmic solution. Our targeted intervention effectively navigates the expressivity versus linearity trade-off, yielding zero-shot accuracy gains of up to +15\%. Furthermore, the asymmetric masking actively preserves the linear separability and the topological integrity of the latent manifold, successfully addressing the geometric limitations of task arithmetic under complex domain shifts.
Background. While Foundation Models like CLIP achieve unprecedented generalization, adapting them to specific downstream tasks via traditional fine-tuning is computationally expensive and fragments their generalized knowledge into isolated models. Task Arithmetic offers a powerful alternative by treating fine-tuning as a directional weight-space shift, or task vector. Manipulating these vectors through simple algebra enables zero-shot multi-task composition and domain adaptation without the need for retraining. Problem. A highly ambitious application of this framework is the zero-shot task analogy, which attempts to transfer learned capabilities from a source domain to a completely unobserved target domain. However, applying simple algebraic summation in a highly non-linear weight space inevitably causes high level of interference. This geometric misalignment entangles the analogical vector with source-specific artifacts. Although recent literature attempts to mitigate this by performing arithmetic in linearized regimes (e.g., the tangent space of the Neural Tangent Kernel), this thesis demonstrates that strict linearization artificially bottlenecks the model's expressivity, rendering it incapable of handling severe structural domain shifts. Furthermore, we identify the unconstrained subtraction of the source task vector tau_A as the primary cause of performance degradation, as it inadvertently erodes foundational, low-level feature extractors. Methodology. This thesis proposes solving the interference directly within the expressive non-linear space. To achieve this, we adapt GradFix, a gradient-guided masking mechanism rooted in model rebasin, to the novel context of zero-shot task analogies. Assuming domain shifts modify largely disjoint parameter subsets, we utilize a minimal target support set (few-shot oracle) to compute a binary mask via gradient sign-consensus. This discrete directional mask is applied asymmetrically (exclusively to the subtractive term) filtering out geometrically misaligned weights and safely decoupling domain-specific noise from transferable semantic knowledge prior to the algebraic projection. Results. Extensive empirical evaluations on the DomainNet benchmark validate both our theoretical diagnosis and algorithmic solution. Our targeted intervention effectively navigates the expressivity versus linearity trade-off, yielding zero-shot accuracy gains of up to +15\%. Furthermore, the asymmetric masking actively preserves the linear separability and the topological integrity of the latent manifold, successfully addressing the geometric limitations of task arithmetic under complex domain shifts.
Exploring Task Analogies in Deep Models: Mitigating Subtractive Interference in Weight Space
BOZZOLI, FABIO
2025/2026
Abstract
Background. While Foundation Models like CLIP achieve unprecedented generalization, adapting them to specific downstream tasks via traditional fine-tuning is computationally expensive and fragments their generalized knowledge into isolated models. Task Arithmetic offers a powerful alternative by treating fine-tuning as a directional weight-space shift, or task vector. Manipulating these vectors through simple algebra enables zero-shot multi-task composition and domain adaptation without the need for retraining. Problem. A highly ambitious application of this framework is the zero-shot task analogy, which attempts to transfer learned capabilities from a source domain to a completely unobserved target domain. However, applying simple algebraic summation in a highly non-linear weight space inevitably causes high level of interference. This geometric misalignment entangles the analogical vector with source-specific artifacts. Although recent literature attempts to mitigate this by performing arithmetic in linearized regimes (e.g., the tangent space of the Neural Tangent Kernel), this thesis demonstrates that strict linearization artificially bottlenecks the model's expressivity, rendering it incapable of handling severe structural domain shifts. Furthermore, we identify the unconstrained subtraction of the source task vector tau_A as the primary cause of performance degradation, as it inadvertently erodes foundational, low-level feature extractors. Methodology. This thesis proposes solving the interference directly within the expressive non-linear space. To achieve this, we adapt GradFix, a gradient-guided masking mechanism rooted in model rebasin, to the novel context of zero-shot task analogies. Assuming domain shifts modify largely disjoint parameter subsets, we utilize a minimal target support set (few-shot oracle) to compute a binary mask via gradient sign-consensus. This discrete directional mask is applied asymmetrically (exclusively to the subtractive term) filtering out geometrically misaligned weights and safely decoupling domain-specific noise from transferable semantic knowledge prior to the algebraic projection. Results. Extensive empirical evaluations on the DomainNet benchmark validate both our theoretical diagnosis and algorithmic solution. Our targeted intervention effectively navigates the expressivity versus linearity trade-off, yielding zero-shot accuracy gains of up to +15\%. Furthermore, the asymmetric masking actively preserves the linear separability and the topological integrity of the latent manifold, successfully addressing the geometric limitations of task arithmetic under complex domain shifts.| File | Dimensione | Formato | |
|---|---|---|---|
|
Bozzoli.Fabio.pdf
accesso aperto
Dimensione
7.95 MB
Formato
Adobe PDF
|
7.95 MB | Adobe PDF | Visualizza/Apri |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14251/7262