I'm a Principal Research Scientist at Foundation AI, Cisco.
TL;DR: Trust, Safety, Security & Privacy (TSSP) — Small Language Models (SLMs)
Almost everything a language model can do well, it does because it is large. That has been the pattern of the last few years, but it isn't an explanation, and it leaves a more basic question open: where does a model's capability actually come from? If we could answer that, we could ask whether size is the only way to get it.
There are three points at which capability can enter a small model, and I'm interested in all of them. The first is (1) pretraining — what the model is trained on to begin with, and how it is built. The second is (2) post-training: the data mixtures, training objectives and reinforcement learning methods that turn a small model into a competent one. What I find most interesting here is that these methods often don't transfer between scales, and something that barely helps a large model can be the thing that makes a small one work. The third is (3) compression, where the capability starts in a larger model and has to survive being carried down. Specifically, I work on quantization-aware training, which allows you to train for the capabilities you want to keep.
Each of these is a different answer to the same question. I work on them in the context of AI trust, safety and security — a field I care deeply about and have been rooted in for most of my research life.
Before this, I completed my Master's at LTI, CMU, where I was fortunate to be advised by Fernando Diaz, Maarten Sap, Niloofar Mireshghallah and Steven Wu. Prior to that, I worked as a Machine Learning Engineer on Adobe's Digital Experience team.
A nonprofit I founded in college to increase legal awareness among women. We found that a significant portion of women, including the educated ones, knew very little about their legal rights β that gap was the driving force behind the initiative. We're always looking for volunteers who can contribute to this cause; feel free to reach out here :)
I volunteered my time researching and developing secure, privacy-preserving AI tools, alongside a wonderful community of open-source contributors.