Gholami, K. (ECE) – Efficient Language Model Construction and Inference via Sparsity

While large language models can match or exceed human performance, they do so with memory and energy costs orders of magnitude greater than biological cognition. We investigate sparsity as a brain-inspired computational principle to address both. We first establish a framework for evaluating small language model construction methods, using the next-token logit distribution as a behavioral fingerprint. Then, we introduce a semi-structured correlation-aware weight sparsity (CWS) method that uses the full activation covariance to identify and prune correlated weights whose combined removal cost is lower than any individual score predicts. CWS, improves perplexity over existing criteria up to 70% sparsity. To extend this gain to extreme sparsity, we propose a hierarchical ADMM framework that optimizes pruning directly against cross-entropy and distillation loss, first layer-wise for efficiency and then globally for cross-layer coordination. This research establishes brain-inspired principles as a foundation for efficient language models that remain accurate even under extreme compression.
Event Host: Kimia Gholami, Ph.D. Student, Electrical & Computer Engineering
Advisor: Jason Eshraghian
Zoom: https://ucsc.zoom.us/j/9827512398?pwd=SGpDWGtVVG81dkgyTHhjbG81dEVUZz09&omn=98349793611
Passcode: 8398