Researchers Advance AI Reasoning About Time Series Data with ICML-Selected Paper

Image
Tom Hartvigsen and Medhasweta San
Ph.D. candidate Medhasweta Sen (left) and assistant professor of data science Tom Hartvigsen (right)

A joint paper submitted by Ph.D. candidate Medhasweta Sen, assistant professor of data science Tom Hartvigsen, and several researchers at Capital One called BEDTime: A Unified Benchmark for Automatically Describing Time Series was selected for publication at the International Conference on Machine Learning (ICML), following a rigorous double-blind peer review. ICML is one of the most prestigious and fastest-growing artificial intelligence conferences in the world.

The research team's work involves the development of AI models that can reason about complex time series data in ways that resemble human reasoning, particularly through the use of language. Hartvigsen says that because today’s best AI models have not been directly trained on time series data, developing models with this ability is particularly challenging.

In their paper, the researchers lay the foundations for more complex reasoning by creating the BEDTime assessment tool to evaluate whether language models can describe different properties of time series using language. If this is possible, Hartvigsen says it clears the path for models to make inferences and draw conclusions about time series data.

He identifies three tasks that AI models should be able to perform to sufficiently describe time series: recognize time series, differentiate between them, and generate accurate descriptions from scratch. Although the first two tasks sound simple, Hartvigsen says they are surprisingly hard for some very capable AI models. Models perform worst at the third task, but Hartvigsen says there are promising signals that recent vision-language models can generate some meaningful descriptions.

“There’s an ongoing race to develop and deploy multimodal AI for time series analysis, and many models fall surprisingly short,” he said. “Our work indicates that the literature may be skipping over some very important tasks that suggest a serious lack of robustness in future deployed models.”

Hartvigsen would like to see the BEDTime assessment tool become a mandatory standard for evaluating multi-modal time series AI models. “I hope that this work inspires the community to pick up the torch and put in the hard work required to build out robust benchmarks to continue developing a rigorous science of benchmarking these generative models.”

The ICML conference will take place in Seoul, South Korea, from July 6-11, bringing together hundreds of international AI scientists, world-renowned AI pioneers, and top academic researchers who are advancing the field.

Author

Writer and Editorial Specialist