Search papers, labs, and topics across Lattice.
This study integrates a three-agent workflow for active data collection, travel behavior modeling, and weather-sensitive demand prediction, utilizing a chatbot-administered survey to gather mode choices from student commuters under varying weather conditions. By employing a multinomial logit model alongside machine learning benchmarks, including logistic regression and random forest, the research evaluates the performance of nine large language models (LLMs) across different prompting strategies. The findings reveal that while traditional models achieved competitive accuracy, the best vision-based configuration of LLMs significantly improved predictions, highlighting the potential of multimodal approaches in behavioral modeling.
Conversational surveys combined with multimodal LLMs can enhance travel behavior predictions, achieving over 71% accuracy by leveraging visual context.
Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately. This study proposes a three-agent workflow integrating conversational data collection, structured data processing, and behavioral prediction. A chatbot-administered, image-augmented stated-preference survey collected mode choices from student commuters across five predefined weather scenarios, yielding 454 respondent-scenario observations. Weather-related associations were analyzed using a multinomial logit model, while logistic regression and random forest provided machine-learning benchmarks. Nine locally deployed large language models (LLMs), ranging from 2 to 35 billion parameters, were evaluated across four zero-shot prompt-and-context conditions and extended through persona, few-shot, and vision-based configurations. Random forest achieved 69.6% five-class accuracy, while the best text-only zero-shot LLM reached 69.9% without task-specific fitting. Habitual travel information produced the most consistent gains, Expert framing generally outperformed Role-Play, and persona information was most useful when habitual travel information was unavailable. Few-shot prompting improved prediction for several models, with gains stabilizing after a small number of examples. Using the same weather images shown to respondents, the best vision-based configuration reached 71.5% five-class accuracy, indicating that visual context may provide additional predictive information for selected models. Overall, the study shows how conversational surveys, structured data processing, conventional behavioral modeling, machine learning, and multimodal LLM prediction can be coordinated within an auditable multi-agent workflow.