Scott, J. (CSE) – Mechanistic Specialization Does Not Guarantee Performance: Evidence from Dual AttentionTransformers

Virtual Event

Dual Attention Transformers (DATs) extend decoder-only Transformers with a dedicated relational-attention stream, making them a natural architecture for abstract identity rules such asABA and ABB. Surprisingly, we find that comparably sized GPT-2 models outperform DATs on these tasks. We investigate this gap with two complementary mechanistic analyses. First, causal mediation analysis shows that DATs exhibit […]

Kembay, A. (ECE) – Sparse and Continual Foundations for Adaptive General Intelligence

Engineering 2 Engineering 2 1156 High Street, Santa Cruz
Hybrid Event

While the human brain learns continually, mastering new tasks without forgetting the old and adapting to unfamiliar ones from context alone, modern neural networks still lack both. To bridge the gap between biological adaptivity and modern AI, we have established foundational work on sparsity as a computational principle at three levels of neural computation, through […]