Soumyabrata Chaudhuri (Soumya)

I’m currently a MS in Computer Science (Thesis) student at UT Austin, where I work with Prof. Joydeep Ghosh. I’m broadly interested in multimodal large language models (MLLMs) and multi-agent systems. During my masters, my work has spanned multimodal reasoning, the evaluation and analysis of MLLMs, efficient VLM/LLMs, and also diffusion models (dLMs, dMLLMs, and image/video generative models). I also spent a summer at Amazon as an Applied Scientist Intern.

Before that, I completed my B.Tech (Honors) degree in Computer Science and Engineering at Indian Institute of Technology (IIT), Bhubaneswar, where my thesis received the Best Thesis Award from the Department of Computer Science. During my undergraduate studies, I conducted research across interrelated domains, including Machine Learning, Computer Vision, Natural Language Processing, and Multi-Modal Learning.

For my B.Tech thesis, I collaborated with Microsoft India to improve the travel-planning process using large language models, under the supervision of Dr. Shreya Ghosh, Dr. Abhik Jana, and Dr. Manish Gupta. I spent a wonderful summer interning at the University of Alberta’s Vision and Learning Lab (MITACS GRI), where I worked with Prof. Li Cheng on motion imitation learning and text-to-3D motion generation. I also had an enriching research internship at IIT Kharagpur, collaborating with Dr. Saumik Bhattacharya on multi-modal learning and action recognition in videos.

Email  /  GitHub  /  Google Scholar  /  Resume  /  LinkedIn  / 

profile photo

Affiliations / Collaborations

UT Austin
2025–present
Amazon
Summer 2026
Microsoft
2024–2025
Univ. of Alberta
Summer 2024
IIT Kharagpur
2023–2025
IIT Bhubaneswar
2021–2025

News

May – Aug 2026 Joined Amazon as an Applied Scientist Intern.
Apr 2026 TripTide is accepted at ACL 2026 (Findings). See you at San Diego!
Mar 2026 Received the Best B.Tech Thesis Award from the Department of Computer Science, IIT Bhubaneswar.
Aug 2025 Joined UT Austin as an MS in Computer Science student.
May 2025 TripCraft is accepted at ACL 2025 (Main Conference). Received an ACL award to subsidize the registration cost. Got a travel grant from IIT Bhubaneswar to present at Vienna.
Jul 2024 – Jun 2025 Collaborated with Microsoft on my B.Tech thesis, on travel planning with large language models.
May – Aug 2024 MITACS Intern at the Vision and Learning Lab, University of Alberta, with Prof. Li Cheng.
Apr 2024 Simba, the first attempt at using Mamba for skeletal action recognition in videos, is out on arXiv.
Dec 2023 ViLP is accepted at ACM ICVGIP 2023 as an Oral.
2023 – 2025 Research collaboration with IIT Kharagpur, working with Dr. Saumik Bhattacharya on multi-modal learning and action recognition in videos.

Research

'*' indicates equal contribution.

project image

TripTide: A Benchmark for Adaptive Travel Planning under Disruptions


Priyanshu Karmakar, Soumyabrata Chaudhuri, Shubhojit Mallick, Manish Gupta, Abhik Jana, Shreya Ghosh
ACL (Findings), 2026
paper / code /

We introduce TripTide, the first benchmark specifically designed to evaluate LLMs’ ability to adapt itineraries in the face of realistic disruptions. TripTide models key dimensions such as disruption severity levels and traveler tolerance profiles, enabling nuanced assessment of LLM responses to unexpected events like transit cancellations, weather-related closures, or overbooked attractions.

project image

TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning


Soumyabrata Chaudhuri, Pranav Purkar, Ritwik Raghav, Shubhojit Mallick, Manish Gupta, Abhik Jana, Shreya Ghosh
ACL (Main), 2025
paper / code /

We introduce TripCraft, a spatiotemporally coherent travel planning dataset for Large Language Models (LLMs) that integrates real world constraints, including public transit schedules, event availability, diverse attraction categories, and user personas for enhanced personalization.

project image

Simba: Mamba augmented U-ShiftGCN for Skeletal Action Recognition in Videos


Soumyabrata Chaudhuri, Saumik Bhattacharya
arxiv, 2024
paper / code /

Simba is the inaugural attempt at leveraging the Mamba model for skeleton action recognition task in videos.

project image

ViLP: Knowledge Exploration using Vision, Language and Pose Embeddings for Video Action Recognition


Soumyabrata Chaudhuri, Saumik Bhattacharya
ACM ICVGIP (Oral), 2023
paper / code /

ViLP explores cross-modal knowledge from the pre-trained vision-language model (e.g., CLIP) to introduce the novel combination of pose, visual information, and text attributes.





Design and source code from Leonid Keselman's website