SUPRITI VIJAY
AI Trust, Safety & Security Researcher
Foundation AI, Cisco

I'm a Principal Research Scientist at Foundation AI, Cisco.

TL;DR: Trust, Safety, Security & Privacy (TSSP) — Small Language Models (SLMs)

Almost everything a language model can do well, it does because it is large. That has been the pattern of the last few years, but it isn't an explanation, and it leaves a more basic question open: where does a model's capability actually come from? If we could answer that, we could ask whether size is the only way to get it.

There are three points at which capability can enter a small model, and I'm interested in all of them. The first is (1) pretraining — what the model is trained on to begin with, and how it is built. The second is (2) post-training: the data mixtures, training objectives and reinforcement learning methods that turn a small model into a competent one. What I find most interesting here is that these methods often don't transfer between scales, and something that barely helps a large model can be the thing that makes a small one work. The third is (3) compression, where the capability starts in a larger model and has to survive being carried down. Specifically, I work on quantization-aware training, which allows you to train for the capabilities you want to keep.

Each of these is a different answer to the same question. I work on them in the context of AI trust, safety and security — a field I care deeply about and have been rooted in for most of my research life.

Before this, I completed my Master's at LTI, CMU, where I was fortunate to be advised by Fernando Diaz, Maarten Sap, Niloofar Mireshghallah and Steven Wu. Prior to that, I worked as a Machine Learning Engineer on Adobe's Digital Experience team.

You've made it this far πŸŽ‰. So let's take a glimpse of my more fun interests as well. Beyond research and various activities, I delight in 🎡 Bollywood music, and πŸŠπŸ»β€β™€οΈ swimming. Additionally, I have a keen interest in 🎭 theatre, a passion I cultivated during college, participating in various dramatic competitions throughout India.
🀝  Non-profit communities I'm passionate about
Felasa Initiative β€” Founder
A nonprofit I founded in college to increase legal awareness among women. We found that a significant portion of women, including the educated ones, knew very little about their legal rights β€” that gap was the driving force behind the initiative. We're always looking for volunteers who can contribute to this cause; feel free to reach out here :)
OpenMined β€” Research Lead & Collaborator
I volunteered my time researching and developing secure, privacy-preserving AI tools, alongside a wonderful community of open-source contributors.

Updates

Dec 2026: Does Safety Molt? was accepted to the WiML Workshop at NeurIPS 2026! πŸŽ‰
Oct 2026: Our paper, NetWorld: Learning Autonomous Cyber Defence Through World-Model Simulation, got accepted to CAMLIS 2026! Back at CAMLIS for a third year! πŸ›‘οΈ
Oct 2026: Attending COLM 2026 in SF! Happy to meet and talk about long-horizon agentic training, and how we can make cybersecurity evals more intrinsic and user-centric. β˜•
Sep 2026: Our team won the Strategic Impact Award at Cisco for our Foundation-Sec models! Insane!! πŸ†
Sep 2026: Giving a guest lecture at Stanford for MS&E 319: Efficient Generative Language Models! 🌲
Aug 2026: One month since we released Antares, and we've crossed 30K downloads and 350+ stars! YAYYY!! πŸŽ‰
Aug 2026: In Vegas for Black Hat 2026 to talk about Antares! Happy to chat about vulnerabilities and how AI is reshaping cybersecurity. 🎰
Jul 2026: NVIDIA and Hugging Face reposted Antares, and Cisco has now been added to open-source contributors! Huge milestone!! πŸ€—
Jul 2026: Antares has crossed 1M views on X! πŸš€
Jul 2026: I led the development and release of Antares, a family of small language models for vulnerability localization. Yayy! πŸš€
Jul 2026: Our paper, Characterizing Cultural Localization in AI-Generated Stories, was accepted to C3NLP at ACL 2026! I'll also be on a panel at the workshop. See you in San Diego! 🌴
May 2026: Attending CAIS '26 in San Jose to give an oral presentation on our paper, Does Safety Molt? Happy to talk about the emergent behaviours that show up in multi-agent settings! 🎀
Mar 2026: Attending RSA Conference 2026 in SF! Always up for a coffee and a chat about AI for security. β˜•
Jan 2026: Started full-time as an AI Research Scientist at Foundation AI, Cisco! πŸ‘©πŸ»β€πŸ’»
Dec 2025: Graduated early from LTI, CMU! πŸŽ“
Nov 2025: Went to Austin for Amazon's bug bounty program, where my team and I won Best Team and a $10,000 cash prize! πŸ†
Oct 2025: Presented my internship project at CAMLIS '25 in DC! πŸŽ‰
Aug 2025: Starting my last semester in CMU! πŸ“š
Aug 2025: Thrilled to share that we just released Foundation-Sec-8B-Instruct, our instruction-finetuned model - go check it out and download now! πŸ“ˆ
July 2025: My internship research project was accepted as a talk at CAMLIS ’25 - I’ll be presenting in DC. Super stoked for this! πŸŽ‰
July 2025: Cisco featured my first blog on zero-shot classification with Foundation-Sec-8B! Go check it out! πŸ’ƒπŸ»πŸ“
May 2025: I'm interning at Foundation AI, Cisco where I'll be primarily contributing to building foundation models in security! πŸ‘©πŸ»β€πŸ’»
Feb 2025: I'll be presenting my research on "Quantifying Political Neutrality in LLM-Generated News Summaries" at AAAI 2025. Come check it out! πŸ•΅πŸ»β€β™€οΈ
Aug 2024: Began my masters at LTI, CMU! Soo excited to see snow for the first time! β„οΈβ˜ƒοΈ
May 2024: My bachelor thesis has been accepted too, at NAACL 2024! It's Mexico TIME! πŸ‡²πŸ‡½
Mar 2024: Our paper, The Silent Curriculum has been accepted to the Workshop on Global AI Cultures at ICLR 2024. See you in Austria! πŸ’ƒπŸ»πŸŽ‰
Aug 2023: I started working at Adobe! πŸ‘©πŸ»β€πŸ’»
Jul 2023: Graduated from MIT, Manipal in Computer Science with a minor in Big Data Technology!
Jul 2023: I'll be presenting my research at the Eastern European Machine Learning School, Kosice! Excited to attend lectures by speakers from DeepMind, NVIDIA, University of Cambridge etc.
Jun 2023: I completed my bachelor thesis at the CoAStaL NLP Lab, under Dr. Daniel Hershcovich!
May 2023: I'll be volunteering at EACL virtually! Feel free to get in touch if you'll be attending the same.
Mar 2023: Guest Speaker for "Talk with Scholars Episode" held by IEEE Computer Society Kerala Chapter!
Feb 2023: I'll be attending AAAI 2023 as an undergraduate scholar where I'll be presenting my research in the Main Hall on Thursday. Come check it out!
Feb 2023: Our paper β€œF-BRIM: A Semi-Supervised Approach for Bias Mitigation with Activation-Weighted Neuron Regularization” is accepted to the Artificial Intelligence for Social Good Workshop at AAAI 2023! πŸŽ‰
Feb 2023: Guest Speaker for "How to get Internships - Job vs Research" at the Mentors Conclave hosted by LeanIn Manipal.
Jan 2023: Participating at the Research Week with Google India. See you in Bangalore!
Oct 2022: Guest Speaker for "A Guide to Research Internships" at Investigar, TechTatva - the official technical fest of MIT, Manipal. Join us at MV seminar hall!
Aug 2022: Invited once again as a guest by ACM-W, Manipal, to guide students in securing research internships! 🀩
Jul 2022: Won second runner's up at the Research & Reports Track in #ShowYourSkill (Coursera) where I developed an NLP-augmented Machine Learning Application for women's safety.
Mar 2022: Guest Speaker for "Paving Your Research Journey: A Guide for Undergraduate Students."
Mar 2022: Delivered a talk at the Search for Research Event held by ACM-W Manipal!
Dec 2021: Received the Mitacs Globalink Research Internship (GRI) Award!
Nov 2021: Won the Adobe India Women-In-Tech (WIT) 2022 Scholarship, securing one of only six scholarships awarded nationwide.
PRESS
Cisco’s AI shrinks haystacks to help security teams hunt needles
Coverage of Antares and how it points code reviewers to where vulnerabilities are likely to be.
InfoWorld, Jul 2026
ARTICLE
Cisco Makes the Case for Smaller AI in Enterprise Software Security as Antares SLMs Cut Token Costs
Coverage of Antares and the case for small language models in enterprise software security.
CX Today, Jul 2026
ARTICLE
ChatGPT Might Have a School Shooting Problem
Cites our FRACTURED-SORRY-Bench work on how splitting a harmful request into harmless-looking questions gets past model safeguards.
The Trace, Dec 2025
ARTICLE
PUBLICATIONS

Technical Reports

Antares: Foundation Models for Agentic Vulnerability Localization
Supriti Vijay, Aman Priyanshu, Didier Chapoteau, Arthur Goldblatt, ..., Amin Karbasi PDF● WEBSITE● MODELS
Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale
Aman Priyanshu, Supriti Vijay, Kimia Majd, Xuhong He, ..., Amin Karbasi PDF● ABS● WEBSITE
FAPO: Fully Automated Prompt Optimization of Multi-Step LLM Pipelines
Paul Kassianik, Baturay Saglam, Huaibo Zhao, Blaine Nelson, Supriti Vijay, Aman Priyanshu, Amin Karbasi PDF● ABS
Llama-3.1-FoundationAI-SecurityLLM-Reasoning-8B Technical Report
Zhuoran Yang, Ed Li, Jianliang He, Aman Priyanshu, Baturay Saglam, Paul Kassianik, ..., Supriti Vijay, ..., Amin Karbasi PDF● ABS● MODEL
Think Before You Retrieve: Learning Test-Time Adaptive Search with Small Language Models
Supriti Vijay, Aman Priyanshu, Anu Vellore, Baturay Saglam, Amin Karbasi PDF● ABS● TALK
Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
Sajana Weerawardhena, Paul Kassianik, Blaine Nelson, Baturay Saglam, Anu Vellore, Aman Priyanshu, Supriti Vijay, ..., Amin Karbasi PDF● MODEL

NLP + Privacy

Et Tu, Brute? Economic Misalignment in Personal AI Agents
Aman Priyanshu, Supriti Vijay, Brian Jabarian, Niloofar Mireshghallah PDF● ABS● X POST
Does Safety Molt? Evaluating LLM Safety in Multi-Agent Social Environments
Aman Priyanshu, Supriti Vijay, Esha Pahwa
Proceedings of the ACM Conference on AI and Agentic Systems (CAIS) 2026 & WiML Workshop, NeurIPS 2026 PDF● WEBSITE● TALK
Are Chatbots Ready for Privacy-Sensitive Applications? An Investigation into Input Regurgitation and Prompt-Induced Sanitization
Aman Priyanshu, Supriti Vijay, Ayush Kumar, Rakshit Naidu, Niloofer Mireshghallah PDF● ABS● CITE

NLP + Society

Characterizing Cultural Localization in AI-Generated Stories
Shaily Bhatt*, Supriti Vijay*, Jeremiah Milbauer, Fernando Diaz
Workshop on Cross-Cultural Considerations in NLP (C3NLP), ACL 2026 PDF● ABS● X POST
When Neutral Summaries are not that Neutral: Quantifying Political Neutrality in LLM-Generated News Summaries
Supriti Vijay, Aman Priyanshu, Ashique R. KhudaBukhsh
Student Abstract and Poster Program, AAAI'25 PDF
Can Abstract Meaning Representation Facilitate Fair Legal Judgement Predictions?
Supriti Vijay, Daniel Hershcovich
Workshop on Insights from Negative Results in NLP, NAACL’24 PDF● CODE● POSTER
The Silent Curriculum: How Does LLM Monoculture Shape Educational Content and Its Accessibility
Aman Priyanshu, Supriti Vijay
Workshop on Global AI Cultures, ICLR 2024 PDF● ABS● SLIDES
F-BRIM: A Semi-Supervised Approach for Bias Mitigation with Activation-Weighted Neuron Regularization
Supriti Vijay*, Aman Priyanshu*, Ashalatha Nayak
Proceedings of AAAI 2023 Undergraduate Consortium & AI4SG@AAAI 2023 PDF● POSTER
"Something Something Hota Hai!" An Explainable Approach towards Sentiment Analysis on Indian Code-Mixed Data
Aman Priyanshu*, Sudarshan Sivakumar*, Supriti Vijay*, Aleti Vardhan*, Nipuna Chhabra*
Workshop on Noisy User-generated Text@EMNLP 2021 PDF● ABS● POSTER
Detecting Gender Bias using Explainability
Gauri Gupta*, Supriti Vijay*, Krithika Ramesh*
Workshop on Widening Natural Language Processing@EMNLP 2021 TALK● SLIDES● PDF

NLP Research + Machine Learning

Counterfactual Explanation Policies in RL
Shripad V Deshmukh, Srivatsan. R, Supriti Vijay, Jayakumar Subramanian, Chirag Agarwal
Counterfactuals in Minds and Machines@ICML 2023 PDF● ABS● CITE
#maskUp: Selective Attribute Encryption for Sensitive Vocalization for English language on Social Media Platforms
Supriti Vijay*, Aman Priyanshu*
Eastern European Machine Learning School (EEML) 2023 PDF● ABS● CODE
AdaptKeyBERT: An Attention-Based approach towards Few-Shot & Zero-Shot Domain Adaptation of KeyBERT
Aman Priyanshu*, Supriti Vijay*
Eastern European Machine Learning School (EEML) 2023 PDF● ABS● CODE
NERDA-Con: Extending NER models for Continual Learning β€” Integrating Distinct Tasks and Updating Distribution Shifts
Supriti Vijay*, Aman Priyanshu*
Updatable Machine Learning Workshop@ICML 2022 TALK● PDF● ABS● CODE
RESEARCH EXPERIENCE
AI PhD Research Intern
Manager: Amin Karbasi
Foundation AI, Cisco Systems
May 2025 - Aug 2025
Research Assistant
Mentor: Dr Fernando Diaz
Starlight Lab, Carnegie Mellon University
Aug 2024 - Present
Aug 2023 - Aug 2024
Research Assistant
Mentor: Dr Daniel Hershcovich
CoAStaL NLP Lab, University of Copenhagen
Jan 2023 - Jun 2023
Research Intern
Mentors: Dr Chirag Agarwal, Shripad Deshmukh
Media and Data Science Lab, Adobe Inc.
May 2022 - Sep 2022
MITACS Globalink Research Intern
Mentor: Dr Amine Trabelsi
Natural Language Processing Lab, Lakehead University
Jul 2022 - Jan 2023
Undergraduate Research Assistant
Mentor: Dr Ashalatha Nayak
Manipal Institute Of Technology
Jan 2022 - Dec 2022
Research Intern
Mentors: Dr Saurabh Aggarwal, Dr. Akrati Saxena
Epidemic Detection Lab, A*STAR Singapore
Jul 2021 - Dec 2021
PROFESSIONAL EXPERIENCE
Adobe Experience Platform, Digital Experience Team
Manager: Mandeep Gandhi
Adobe Inc.
Aug 2023 - Aug 2024
Data Science Intern
Manager: Henry Ho
Asia Pacific Resources International Holdings (APRIL)
Mar 2023 - May 2023
TALKS
MS&E 319: Efficient Generative Language Models — Guest Lecture
A guest lecture for Stanford's course on efficient generative language models.
Stanford University, 2026
COURSE
Antares: Foundation Models for Agentic Vulnerability Localization
A talk on our family of compact models that localize vulnerabilities in real code repositories as autonomous agents.
Black Hat 2026
X POST● WEBSITE
Does Safety Molt? Evaluating LLM Safety in Multi-Agent Social Environments
A conference talk on how model safety behaviour shifts once models are placed in multi-agent social settings.
ACM Conference on AI and Agentic Systems (CAIS) 2026
VIDEO
Reason. Search. Retrieve. Repeat. Iterative Retrieval for Automating Vulnerable Code Discovery
A talk on teaching small language models to search adaptively at test time to find vulnerable code.
CAMLIS 2025
VIDEO● PAPER
Women in Stem Series: Leading Women in Cybersecurity
A spotlight interview where I share my journey in tech and challenges I faced along the way.
picoCTF - CMU Cybersecurity Competition
VIDEO● LINKEDIN POST
Talk With Scholars
An interview on the Adobe India Women-In-Tech Scholarship
IEEE CS Kerala Chapter
VIDEO● POSTER
How to get Internships - Job vs Research
An interview on how to secure internships, both industry and research
Mentors Conclave, LeanIn Manipal
POSTER
A Guide to Research Internships
An interview on how to secure internships, both industry and research
Investigar, TechTatva - Technical fest of MIT Manipal
POSTER
Securing Research Internships
An interview on how to secure internships, both industry and research
ACM-W, Manipal
POSTER
Paving Your Research Journey: A Guide for Undergraduate Students
An interview on how to secure internships, both industry and research
Research Society MIT
VIDEO● POSTER
Search for Research Event
An interview on how to secure internships, both industry and research
ACM-W Manipal
POSTER