Professor earns NSF CAREER Award to defend AI models from attackers

Weijie Zhao’s research aims to enhance machine learning safety, resilience, and accountability

Traci Westcott/RIT

Weijie Zhao, assistant professor of computer science, recently received an NSF CAREER Award to build machine learning models that are more secure and interpretable.

Artificial intelligence has become a powerful tool for critical systems in healthcare, finance, and national security. However, these complex machine learning systems can also be targets for attackers.

Weijie Zhao, an assistant professor of computer science at RIT, wants to shine a light on AI design and make sure that machine learning models are not embedded with hidden vulnerabilities or dangerous exploits. His research will help build machine learning models that are more secure and interpretable.

“As we use AI and agentic AI more, this really is a public concern,” said Zhao. “We want to make sure that there is no false information or misinformation and that decisions are not being manipulated.”

Zhao recently earned a prestigious National Science Foundation Faculty Early Career Development (CAREER) award and grant for his work. His five-year project is titled “Defending Machine Learning Models from Adversarial Threats via Unified Interpretability and Attribution.”

The main problem stems from the fact that complex machine learning behaviors are like a black box—while the outputs might be visible, the AI’s internal decision-making process is incomprehensible to humans. The result is a lot of unknowns about the AI system.

For example, an AI agent could be trained using a model that unintentionally has false information. Attackers could also build large language models that have a backdoor or watermark.

“Attackers don’t even need to steal your data,” explained Zhao. “They could make a model that influences decision generation in some way.”

Zhao noted that on OpenClaw, the popular free and open-source AI agent, threat actors have been caught using third-party extensions to essentially distribute malware. In February, researchers found 341 malicious ClawHub skills that were stealing data from users.

Making machine learning models transparent

The CAREER Award project seeks to enhance the safety, resilience, and accountability of machine learning systems that are deployed in high-stakes environments.

Through his research, Zhao hopes to bridge the gap between modern AI systems and classical machine learning frameworks that scientists already understand. With these tools, defenders would be able to trace the origin of failures and repair vulnerabilities. Zhao plans to:

  • Develop techniques to identify how adversarial inputs or training data components lead to harmful outputs.
  • Design fast strategies to remove harmful behavior without full retraining, utilizing surrogate models to semantically validate that repairs are localized and verifiable.
  • Build automated pipelines for training data auditing, provenance tracking, and security-aware valuation to detect data poisoning and instability.
  • Create an interface where users can explore suspicious outputs and interactively remediate vulnerabilities.

“Essentially, we want to find that malicious data in the inference time, correct it without having to retrain the model, and provide proof that it’s fixed,” said Zhao.

At RIT, Zhao is working with five computing and information sciences Ph.D. students. He hopes that future practitioners will use this defense framework to create more resilient, transparent, and trustworthy machine learning tools.

“This is very important, because right now, many developers are chasing the best AI performer,” said Zhao. “But they should also be chasing the security and guardrails for responsible AI systems.”

The prestigious CAREER Award program recognizes and supports junior faculty who exemplify the role of teacher-scholars through integrated research and educational activities. RIT has more than a dozen NSF CAREER award winners working at the university.