My research sits at the intersection of AI, security, and software systems. I develop principled, practical techniques to make modern computing more reliable and secure. My earlier work improved Linux, which underpins computing for billions of people, by revealing where bugs concentrate and helping fix dozens of kernel security bugs and file-system bugs. Our 2015 paper coined the term machine unlearning and helped launch the field. In 2017, DeepXplore helped establish systematic testing for neural networks and influenced Google’s TensorFuzz. More recently, Radshield, our software-based radiation-protection system, was deployed for about two years on a commodity SoC aboard NASA’s Perseverance Mars rover, where it safeguarded a navigation algorithm used during autonomous driving. I earned my PhD and MS in Computer Science from Stanford University and my BS from Tsinghua University.
I'm looking for PhD students and postdocs, as well as MS and undergraduate interns. If you know how to build systems, tools, or models, we should talk. Just shoot me a human-written email.
Previously, I co-founded and led NimbleDroid, a Columbia spin-off that turned our research into automated mobile-app performance tools used by companies including Pinterest, Flipkart, Tinder, and The New York Times.
Selected Honors
- 2025 ACM Fellow
- 2025 ACM SIGOPS Mark Weiser Award
- 2025 IEEE S&P Test-of-Time Award
- 2012 Sloan Research Fellowship
- 2012 AFOSR Young Investigator Program Award
- 2011 NSF CAREER Award
Recent Papers
- TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory Minimization
- Understanding Automated Program Repair Agents Through the Lens of Traceability: An Empirical Study
- Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All
- kAgent: An execution-guided crash resolution agent for the Linux kernel
- zkFuzz: Foundation and Framework for Effective Fuzzing of Zero-Knowledge Circuits
- Detecting Privilege Escalation in Polyglot Microservices via Agentic Program Analysis
-
Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers
IEEE S&P Distinguished Paper Award
- Radshield: Software Radiation Protection for Commodity Hardware in Space
Selected Papers
These papers trace my research from foundational work on reliable storage and concurrency to today’s trustworthy AI systems.
See the complete list of 125 publications.
-
Diversity Helps Jailbreak Large Language Models
NAACL Oral Presentation
A state-of-the-art, fully automated red-teaming tool that turns prompt diversity into an efficient black-box jailbreak strategy, exposing 5× more safety failures across leading LLMs with 10× fewer queries.
-
Raidar: geneRative AI Detection viA Rewriting
RAIDAR introduced divergence under rewriting, an interpretable signal that generalizes across domains and requires only black-box LLM access. Ask another LLM to rewrite the text: AI-generated text changes less because it lies closer to the rewriter’s statistical norm. The idea later extended to AI-generated video and audio.
-
DeepXplore: Automated Whitebox Testing of Deep Learning Systems
CSAW 2018 Applied Research Second PlaceCACM Research HighlightSOSP Best Paper Award
DeepXplore was the first white-box testing tool to bring coverage-guided fuzzing and differential testing to neural networks. It introduced neuron coverage, uncovered thousands of flaws, helped launch a new research field, and influenced Google’s TensorFuzz.
-
Towards Making Systems Forget with Machine Unlearning
IEEE S&P Test-of-Time AwardICBS Frontiers of Science Award
This paper coined the term machine unlearning: the idea that models should efficiently forget selected training data and its influence. It helped launch a field that has grown from classical learning algorithms to deep neural networks and LLMs.
-
Making Parallel Programs Reliable with Stable Multithreading
StableMT introduced a new concurrency model built on a radical question: do parallel programs need exponentially many thread schedules? Reusing a small set of tested schedules across inputs makes production behavior more predictable, testable, and reliable.
-
Using Model Checking to Find Serious File System Errors
OSDI Best Paper Award
FiSC leveraged model checking to systematically explore crash states beyond conventional testing. It found serious bugs in every file system checked—32 across ext3, JFS, and ReiserFS—including failures that could irrecoverably destroy entire directories, even the file-system root; most were patched within a day.
Current Advisees
I'm fortunate to work or have worked with these brilliant people.
- Yun-Yun Tsai, PhD student
- Andreas Kellas, PhD student
- Raphael Jedidiah Sofaer, PhD student
- Harry Haoda Wang, PhD student
- Alex Mathai, PhD student
- Jinjun Peng, PhD student
- Hailie Mitchell, PhD student
- Hideaki Takahashi, PhD student
- Jihwan Kim, PhD student
- Chenxi Huang, PhD student
- Weiliang Zhao, PhD student
I co-advise some students in the SSL lab.
See current advisees and alumni.
Recent Teaching
- Fall 2026 E6998: Build an Agent Startup
- Fall 2025 W4152: Engineering Software-as-a-Service
- Fall 2025 E6113: Agent for Work
See the complete teaching history.
Support
We are grateful to the organizations that support our research and teaching. See acknowledgments.