Google Funds UVA Research Into the Safety of Autonomous AI Agents
Google’s Gemini for Research initiative awarded Chirag Agarwal $20,000 in Google Cloud Credits last month to support AI research. Agarwal, an assistant professor of data science at the University of Virginia and head of the Aikyam Lab, will use the funds to investigate the safety of autonomous, multimodal agents — AI systems that can act on their own, processing text, images, and other data to complete tasks with little to no human oversight.
Agarwal says that evaluating the safety and interpretability of frontier AI systems, particularly as they transition into autonomous, multimodal agents, demands immense computational bandwidth.
According to Agarwal, AI safety testing is shifting from testing individual models to evaluating how hundreds of autonomous AI systems deployed across multiple platforms interact and make irreversible, real-world decisions. For example, if a patient goes to the doctor, an initial AI agent might interpret the patient’s symptoms and history, flagging specialists or tests that may be needed. That information might then go to a diagnostic agent, which analyzes labs and physician notes before proposing a diagnosis. Yet another agent may use that information to suggest a medication treatment plan. But if one AI agent makes a mistake early on, for example, misreading an ambiguous scan, that error will be perpetuated by the next agent and the one after that, resulting in a possible misdiagnosis that could cause real harm. Agarwal's research lab hopes to evaluate and identify these kinds of vulnerabilities.
“The Gemini award provides the crucial infrastructure needed to run these high-throughput, population-level scientific experiments,” Agarwal said. “It empowers the Aikyam Lab to move beyond surface-level analysis and pursue ‘actionable interpretability,’ driving tangible advancements in how we align and secure complex large reasoning models in multi-agent environments in critical real-world applications like financial trading, autonomous drug prescription, and healthcare.”
The Google funds will directly support Agarwal’s investigations into the reliability and vulnerabilities of multimodal agents. Specifically, he plans to use Gemini’s advanced processing capabilities to uncover, analyze, and map the thinking processes of frontier AI systems before they provide a response. He also plans to rigorously test if an AI model’s chain-of-thought (the reasoning process that the AI model provides) is a reliable way to catch dangerous behavior, or to inspect the latent representation space, to keep the system safe.
Agarwal’s research supports several major initiatives that intersect with the bleeding edge of AI research in large language models.
“We are advancing techniques to evaluate the causal validity of LLM reasoning pipelines,” he said. "That means testing whether an AI agent’s explanation for a high-stakes decision reflects how it arrived at that conclusion, instead of just assuming that its reasoning is correct.” His lab is currently building MEA, an agentic tool that automatically generates plain-language explanations of why a multimodal model behaved the way it did and mechanistic interpretability tools to capture the safety and vulnerabilities of multi-agent systems.
As AI systems become increasingly capable of acting independently, understanding not just what they do but why they do it will be critical. Agarwal's research aims to help ensure that the next generation of AI is not only more powerful, but also more transparent, reliable, and worthy of public trust.



