Brown University
Back to Results

Optimizing Machine Learning Models through Parameter-Efficiency

Description

Abstract:
Machine learning (ML) technologies have experienced unprecedented advancements, catalyzing theoretical breakthroughs and enabling wide-ranging applications across various domains. However, as these systems become more complex and are expected to perform a wider range of tasks, the challenges associated with efficient training, fine-tuning, and deployment become increasingly prominent. One key strategy to address these challenges is parameter efficiency, which involves minimizing the number of parameters that need adjustment during training or inference, thus reducing computational resources while maintaining or improving performance. This dissertation examines various approaches to enhance parameter efficiency in ML models through exploring novel architectures and unconventional computing paradigms. We present three main contributions aimed at enhancing parameter efficiency in ML models while maintaining or improving performance. The first contribution introduces MTLoRA, a fine-tuning framework that leverages low-rank adaptation matrices for parameter-efficient training of MTL models. By effectively disentangling the parameter space for diverse tasks, MTLoRA minimizes the number of parameters that need to be adjusted, training only a small fraction of the model while maintaining high adaptability and performance. Building on the principles of parameter-efficient architectures, the second contribution, CRLoMTL, addresses the challenges of multi-objective optimization in MTL related to sharing parameters between different tasks. CRLoMTL employs parameter-efficient methods to tackle the issue of conflicting gradients across the shared backbone. By utilizing low-rank matrix adaptations to handle gradient conflicts, CRLoMTL integrates a multi-pass training strategy to effectively capture task interactions and systematically minimize inter-task conflicts in the shared parameters, leading to improved convergence and overall model performance. We further extend these ideas by improving the parameter-efficiency of MTL models through an adaptive framework for input-dependent dynamic sparsification in multi-task learning. Our framework employs a lightweight policy network to recognize unnecessary computations based on input complexity, further optimizing model adaptability and efficiency. Moving beyond traditional silicon-based systems, the third contribution explores unconventional platforms to build parameter-efficient systems. Unconventional computing, such as chemical computing, offers unique computational paradigms with potential inherent parameter efficiency through analog encoding, parallelism, and bio-compatibility. We present a molecular computation system that uses acid-base reactions to perform calculations, efficiently encoding information in analog chemical form. By leveraging the natural complementarity of acids and bases, our approach demonstrates the feasibility of constructing neural network classifiers in a chemical medium. Our approach efficiently encodes information in analog chemical form, relying on the natural complementarity of acids and bases, demonstrating the feasibility of constructing neural network classifiers in a chemical medium. The model is distinguished by two key advantages: inherent parallelism and signal encoding efficiency. The model leverages simultaneous chemical reactions to model the parallel processing of the neural network computation. Furthermore, our model achieves parameter efficiency through analog encoding of signals as acid-base concentration levels, providing a more granular and expansive range of values for each computational unit compared to traditional binary encoding. Furthermore, we extend our chemical computation platform to leverage enzymatic reactions to model neural networks and digital gates through pH modulation, showcasing the potential for diversifying the computational strategies for building parameter-efficient architectures.
Notes:
Thesis (Ph. D.)--Brown University, 2024

Citation

Agiza, Ahmed Abdelrazek, "Optimizing Machine Learning Models through Parameter-Efficiency" (2024). Computer Science Theses and Dissertations. Brown Digital Repository. Brown University Library. https://repository.library.brown.edu/studio/item/bdr:tyrqr7f5/

Relations

Collection: